Chapter 01
The duplication audit
Start here, before you buy anything. The single most common finding when we look at a company’s AI spend is that they are paying three vendors to do one job, and that at least one of those jobs is already included in a tool they were paying for anyway.
This happens for a structural reason, not because anyone was careless. Between 2024 and 2026 essentially every incumbent software vendor shipped AI features into products you already licensed. Your note-taking app got a summariser. Your CRM got a writer. Your design tool got generative fill. Meanwhile individual teams, moving faster than procurement, bought point solutions for exactly those jobs.
The audit is one page and takes an afternoon:
- List every AI tool anyone expenses. Include the $20/month personal subscriptions, in most companies that shadow spend is larger than the sanctioned spend.
- Next to each, write the job, not the product category. Not “AI writing” but “drafts first-pass replies to inbound support email.”
- Group by job. Any job with more than one tool against it is a candidate for consolidation.
- For each remaining tool, check whether something you already pay for now does that job. This is the step that finds the money.
- Only then look at what to buy.
Two numbers worth having before you start: 413 of the 716 tools in our index have a free tier, and the gap between the cheapest and dearest frontier model is roughly 385×. Both mean the same thing: the price you pay is much more a function of what you chose than what the task required.
Chapter 02
What the models cost
Almost every AI product is a wrapper around one of these. Knowing what the underlying model costs tells you what margin your vendor is taking, and whether building the thing yourself is remotely sensible.
Two numbers decide most of it. Context window is how much the model can consider at once. Price per million tokens is what it costs, quoted separately for input (what you send) and output (what it writes back). Output is always the expensive one, typically three to five times input, which is why the useful optimisation is almost always “make it write less,” not “make it read less.”
| Model | Developer | Context | In / 1M | Out / 1M |
|---|---|---|---|---|
| Qwen 3.7 Flash | Alibaba | - | $0.03 | $0.13 |
| DeepSeek V4 Flash | DeepSeek | 1M | $0.14 | $0.28 |
| DeepSeek V4 Pro | DeepSeek | 1M | $0.43 | $0.87 |
| GPT-5.6 Luna | OpenAI | 1M | $0.20 | $1.20 |
| Grok 4.20 | xAI | - | $1.25 | $2.50 |
| Claude Haiku 4.5 | Anthropic | 200K | $1 | $5 |
| Gemini 3.6 Flash | 1M | $1.50 | $7.50 | |
| GPT-5.6 Terra | OpenAI | 1M | $2 | $12 |
| Gemini 3.1 Pro | 2M | $2 | $12 | |
| Claude Sonnet 5 | Anthropic | 1M | $3 | $15 |
| Claude Sonnet 4.6 | Anthropic | 1M | $3 | $15 |
| GPT-5.4 | OpenAI | - | $2.50 | $15 |
| Claude Opus 5 | Anthropic | 1M | $5 | $25 |
| Claude Opus 4.8 | Anthropic | 1M | $5 | $25 |
| GPT-5.6 Sol | OpenAI | 1M | $5 | $30 |
| GPT-5.5 | OpenAI | 1M | $5 | $30 |
| Claude Fable 5 | Anthropic | 1M | $10 | $50 |
The arithmetic that matters. A moderately chatty production workload (100 calls a day, 10,000 tokens in and 2,000 out each time) costs roughly $2 a month on Qwen 3.7 Flash, and about $600 on Claude Fable 5. Same workload. The difference is not quality of answer on most tasks. It is whether the task actually needed the expensive model.
The mistake nearly everyone makes is buying capability they never measure. Start on a mid-tier model, build an evaluation you trust, and move up only when you can point at a specific failure the cheaper model produced. Teams routinely over-buy capability and under-buy evaluation, which is exactly backwards.
Watch for tiered pricing above a context threshold, several models roughly double their input rate past a certain point, which turns a long-document workload into a surprise invoice.
Chapter 03
Buy, wrap, or build
There are three ways to add AI to a workflow, and most bad decisions come from not knowing which one you are looking at.
Buy
A finished product. Someone else owns the model, the prompt, the interface and the maintenance. Costs the most per seat and the least in time. Right for anything that is not your differentiator, which is nearly everything.
Wrap
You call a model API and build a thin layer of your own: your prompts, your data, your interface. Cheap per unit, and the layer is genuinely yours. Right when the workflow is specific to how your business works and no product fits it. The trap is underestimating maintenance: a wrapper is not finished when it works, it is finished when it still works after the model changes underneath it.
Build
Training or fine-tuning your own model. Right in a genuinely small number of cases: you have proprietary data nobody else has, the task is narrow and repetitive enough that a small model beats a general one, or per-unit API economics genuinely break at your volume. If you cannot state which of those three applies, you are not in this category.
The eighteen-month test. Price all three over eighteen months including the engineering time, not over one month including only the licence. Buying usually wins on a one-month view and loses on a three-year one; building loses on both more often than its advocates expect. Wrapping wins more often than people expect, because the expensive part, the model, is rented rather than owned.
Chapter 04
Category by category
The top five in each category by trust score, with what they cost. These are starting points for a shortlist, not verdicts, the right answer depends on the job you wrote down in chapter one.
Coding
See all →| 01 | Claude Code · Anthropic | Included in Claude Pro $20/month, Max from $100/month, Team and Enterprise; or pay-as-you-go API credits |
| 02 | GitHub Copilot · GitHub/Microsoft | Free; Pro $10/month, Pro+ $39/month, Max $100/month |
| 03 | Cursor · Anysphere | $20/month Pro |
| 04 | Warp · Warp | Free - $15/month |
| 05 | Snyk Code · Snyk | Free - Enterprise |
Writing
See all →| 01 | Grammarly · Grammarly, Inc. | Free - $12/month |
| 02 | Grammarly AI · Grammarly | Free - $30/month |
| 03 | Copyleaks · Copyleaks | Free - $9.99/month |
| 04 | Plus AI · Plus | $10-$21/month |
| 05 | DeepL Write · DeepL SE | Free - $10.99/month |
Chatbot
See all →Image
See all →| 01 | Midjourney · Midjourney | $10-60/month |
| 02 | DALL-E 3 · OpenAI | $0.04-0.08 per image (API) |
| 03 | Freepik · Freepik | Free - $12/month |
| 04 | Gigapixel AI · Topaz Labs | $99.99 one-time |
| 05 | Adobe Firefly · Adobe | Included with Creative Cloud ($22.99+/month) |
Video
See all →| 01 | Vyond · Vyond | $49-$92/month |
| 02 | Runway Gen-3 · Runway | $12-76/month |
| 03 | Sora · OpenAI | Included with ChatGPT Plus ($20/month) |
| 04 | DeepBrain AI · DeepBrain AI | $30-$225/month |
| 05 | Wondershare Filmora · Wondershare | Free - $49.99-$79.99/year |
Audio
See all →| 01 | Whisper · OpenAI | Free (open source) / $0.006/min API |
| 02 | ElevenLabs · ElevenLabs | Free - $99/month |
| 03 | Speechmatics · Speechmatics | Free tier + usage-based |
| 04 | Sonix · Sonix | $22/month |
| 05 | Alitu · Alitu | $38-$48/month |
Agents
See all →| 01 | Liner · Liner | Free - $20/month |
| 02 | Warmly · Warmly | Free - $700/month |
| 03 | Persana AI · Persana AI | Free - $85/month |
| 04 | MindStudio · MindStudio | Free - $20/month |
| 05 | CrewAI · CrewAI | Free (open source) |
ML & infrastructure
See all →| 01 | Hugging Face · Hugging Face | Free - $9/month |
| 02 | Coefficient · Coefficient | Free - $49/month |
| 03 | Seldon · Seldon | Free (open source) |
| 04 | Slidebean · Slidebean | $29-$149/month |
| 05 | FinChat · Fiscal AI | Free - $50/month |
Chapter 05
The procurement questions
Eleven questions to ask before signing. Most of them are about what happens to your data, and what happens when you want to leave. Send them in writing and keep the reply.
- 01
Does our data train your models?
The answer should be no, or off by default. If it's buried in a settings page, assume the default is the one that matters.
- 02
Where is our data stored, and under whose jurisdiction?
Matters more than people think the moment a regulator, an enterprise customer, or a lawyer asks.
- 03
What happens to it when we leave?
Export format, retention period, deletion guarantee. Ask for it in writing.
- 04
What model is underneath, and can it change without telling us?
Most products are a wrapper. A silent model swap can change your output quality overnight.
- 05
Is pricing per seat, per usage, or per outcome?
Per-seat is predictable. Usage-based tracks value but can surprise you. Agent products increasingly bill per run: model it at 10× your pilot volume.
- 06
What's the real cost at 10× our current volume?
Ask them to price it. The gap between the pilot and the rollout is where budgets die.
- 07
What's the rate limit, and what happens when we hit it?
Degraded, queued, or refused are three very different answers.
- 08
Who is accountable when it's confidently wrong?
Especially in legal, medical and financial contexts. 'The model said so' is not a defence.
- 09
Can we get our prompts and configuration out?
Prompt and workflow lock-in is the switching cost nobody prices at purchase.
- 10
What's your uptime record, and where is it published?
If there's no public status page, that is the answer.
- 11
Who else our size uses this, and can we talk to them?
A vendor who can't produce a comparable reference is telling you something.
Chapter 06
Assembling the stack
Three reference stacks. Not recommendations to copy. They are shapes, showing where the money goes at each size and what gets added when.
The solo operator
$40 to 80 / month| One good assistant | The frontier chatbot you like best. This is the workhorse. | $20 to 30 |
| A coding assistant | Only if you write code. If you do, this is the highest-return line on the list. | $20 |
| Everything else | Free tiers. Seriously, at this size, paid image and video generation is usually a want, not a need. | $0 |
The failure mode here is subscription creep: six $20 tools nobody cancels. Review quarterly and be ruthless.
The ten-person team
$400 to 900 / month| Assistant seats | Team plan rather than individual subscriptions. You get admin, shared context and one invoice. | $25 to 30 × seats |
| Coding assistant seats | For the engineers only. Don't buy it for everyone. | $20 to 40 × devs |
| One category tool | Whichever job in chapter one had the largest hours-per-week against it. | $50 to 200 |
| API budget | For the wrapper you will inevitably build. Start at $100 and watch it. | $100 |
This is the size at which the duplication audit pays for itself, usually several times over. Do it before adding the category tool, not after.
Mid-market
$3,000 to 15,000 / month| Enterprise assistant | With SSO, admin controls and a data-processing agreement. The controls are the product. | $30 to 60 × seats |
| Engineering tooling | Assistants, review, and increasingly agents that open pull requests. | $40 to 100 × devs |
| Function-specific platforms | Support, marketing, or analytics. Bought by the department that owns the outcome. | $500 to 5,000 |
| Model API + infrastructure | Direct API spend, plus observability and evaluation. See the infrastructure category. | $1,000+ |
| Governance | Someone's actual job, not a policy document nobody reads. | Headcount |
At this size the binding constraint stops being cost and becomes governance: who is allowed to send what data where. Budget for the answer before a regulator or an enterprise customer asks the question.
One last thing
Every figure here is drawn live from our index, which means this document is accurate on the day you read it rather than the day it was written. It also means it will disagree with itself over time. That is the intention.
If something in it is wrong, and in a market moving this fast, something will be, tell us and we will fix it. If you want a second opinion on a stack you are putting together, reply to the email this came with. We read everything.