reflection 17 min read

When the best AI model stops being worth it

23 August 2026


Anthropic’s frontier model costs twice as much as the one just behind it and gets businesses barely a tenth of their spend. That isn’t a verdict on the model. It’s a market working out how much intelligence a task actually needs.

I read a Financial Times report on businesses resisting Anthropic’s most powerful model and had an odd reaction: of course they are. Not because the model is bad, and not because AI progress has stalled. Because I have access to the same choice in my own work, and I’ve struggled to find a practical reason to make the most powerful model my default.

Anthropic’s Fable 5 is billed as its most capable generally available model, built for ambitious, long-running, asynchronous work, at $10 per million input tokens and $50 per million output tokens. Opus 5, its next model down, costs half as much, and Anthropic itself says Opus comes within about half a percentage point of Fable’s peak performance on one of its coding benchmarks. Pay twice as much for the last sliver of capability — there will be jobs where that’s worth it, but I’m not convinced most of my jobs are among them.

Businesses appear to be reaching the same conclusion by a more rigorous route. Ramp analysed model spending across the businesses on its platform and found Fable accounting for only 6% of Anthropic tokens purchased and 11.4% of dollars spent on Anthropic models. Ramp’s economist Ara Kharazian put the implication plainly: this may be an upper bound on what businesses are prepared to pay for additional AI performance. More performance exists. Customers aren’t automatically buying it.

Practical capability saturation

That’s more consequential than another leaderboard entry. It looks like evidence that for a large category of professional work, we’re approaching something close to practical capability saturation — not maximum intelligence, and not the end of AI progress, but the point at which another increment of model capability stops materially changing the work.

My own experience is a sample of one and proves nothing about the market by itself, but it’s consistent with what the market is apparently doing. I use AI across research, argument development, writing, editing, document analysis, interrogating a growing repository of past work, building websites and writing code, and I move between models depending on what I’m trying to accomplish. There are real differences between them — one is better at a particular coding job, another is easier to think an argument through with, one interface surfaces existing context more usefully — and I’ve written about that trade-off before in Team Claude or Team ChatGPT?.

But those differences increasingly sit on top of a much larger shared capability. For most of what I hand these systems, the live question is no longer whether a model can do the job. It’s which model gives me the most convenient or reliable route to doing it.

There’s a basic economic reason the earlier jumps in generative AI don’t repeat forever. Going from a model that can’t build a useful application to one that can, or from unreliable prose to something that can help construct and interrogate a serious argument, or from being unable to work usefully across a large codebase to being able to — those are enormous increases in practical value, because they cross a threshold: something that effectively wasn’t possible becomes possible.

Eventually you reach a different part of the curve instead, where one model does the work very well, the next does it extremely well, and the one after that does it extremely well with marginally fewer errors on some benchmark at substantially higher cost. Those gains can be technically real while barely registering for the person sitting in front of the computer.

The difference between 40% and 80% task success is transformative. The difference between 95% and 96% might matter enormously inside a highly automated system running millions of tasks a day, and be almost invisible to a single knowledge worker supervising the output. Those are not the same market, and pricing built for the first one doesn’t automatically survive contact with the second.

Redundancy is cheap, and that changes the purchase

There’s a second reason the economics have shifted, and it has less to do with any one model’s ceiling and more to do with how easy it now is to walk away from it. If a model can’t do something, I don’t need to buy vastly more access to a single supermodel — I can move the problem. For roughly another twenty pounds a month I get access to another highly capable general-purpose system, so if one is struggling with a task I send it to the other.

That turns the purchasing question from how do I get unlimited access to the most intelligent model available into something closer to normal software procurement: what’s the cheapest reliable combination of models that covers the work I actually do. It’s a far less flattering question for the idea that every new frontier release automatically deserves a large pricing premium.

Once cheap redundancy is on the table, routing starts to beat brand loyalty. Suppose three models can all handle 95% of my ordinary work, and one costs considerably more because it can also handle a narrow class of unusually hard problems. I could route everything through the expensive model by default. Or I could use a cheaper model as the default and escalate only the small number of jobs that actually need the extra capability.

For a lot of professional use I suspect the second pattern wins, which makes frontier intelligence an escalation layer rather than a default — the same logic organisations already apply to expertise generally. You don’t send every support ticket to your most senior engineer, put every legal question in front of your most expensive barrister, or ask a CFO to sign off a twelve-pound software renewal; work gets routed by difficulty because expertise is expensive, and AI spending looks to be settling into the same pattern.

It’s a reversal of how the first few years of generative AI actually behaved, where a new frontier model shipped and everyone moved to it as a matter of course, because the biggest benchmark number was treated as the default product. Ramp’s data — Fable’s usage share small despite being described as the most performant model available — reads less like disappointment with the model and more like a market learning to route around that reflex.

What saturation doesn’t mean

Saturation isn’t the same as a ceiling. I don’t think we’re near the limit of what AI can do, and that would be a large claim on thin evidence. There are whole areas where current systems still feel primitive — reliable autonomous agents that need real supervision, checking and recovery rather than a clean handoff; video work where the outputs have improved a great deal but keeping something consistent across shots into a repeatable professional workflow still feels awkward; complex multimodal work with similar problems; and a broad category of tasks where the model is arguably capable enough already but the interface between the human and that capability is still the bottleneck.

The next improvement I actually want in most of my own work usually isn’t a model that reasons another few percent harder. It’s something easier to direct — that knows when to act and when not to, that can operate reliably across tools, that can hold a complicated task for hours without needing to be restarted or redirected, that can show its work.

Those are large unsolved problems, so saturation here means something narrower: for some existing workflows, the marginal value of more raw intelligence is falling, which is enough on its own to change how the market buys.

There’s a real counterargument here, too. Anthropic positions Fable around multi-day autonomous tasks and genuinely ambitious coding and knowledge work, and if that capability lets a company automate something it previously couldn’t touch at all, paying double per token is trivial against the labour value created. Someone has to build the capability that eventually becomes ordinary — today’s expensive breakthrough is next year’s default — and businesses declining to route everything through the most expensive model is not the same claim as there being no business case for building it.

The mistake would be turning “most tasks don’t need the frontier model” into “the frontier model has no case,” which doesn’t follow. If anything, the better routing gets, the more frontier models may end up reserved specifically for the problems that require frontier capability — which could make them both more valuable and rarer in everyday use at the same time. Those aren’t contradictory positions.

Where the value actually moves

This is where the argument connects to what I’ve been building through Editorial Intelligence. If capable models are scarce, model access is the advantage. Once capable models become abundant, the advantage moves to whatever the model doesn’t arrive with: evidence, institutional knowledge, accumulated context, standards, past decisions, understanding of an audience, the judgement to recognise when a plausible-sounding answer is wrong, the workflow connecting all of it, and increasingly the decision about which model should handle which job.

That’s the case for model-agnostic workflows, and I’ve made a version of it before in the workflow should outlive the model: my Editorial Intelligence repository isn’t valuable because Claude can read it, or because ChatGPT can read it — it’s valuable because it exists independently of either, so the context survives a vendor change, a price change or a model I currently prefer being overtaken by something else. Until now I’d mostly made that case around resilience.

There’s a second argument for portability that matters just as much and is more directly commercial: if your knowledge and workflows belong to you rather than to whichever tool happens to hold them, you can send that context to whichever model gives the best price-performance for a given job. That’s much harder once the useful part of the system is trapped inside one vendor’s interface.

There’s a paradox in that. Better models might actually increase the relative importance of context rather than shrink it — the opposite of what you’d expect. Give two companies exactly the same extraordinarily capable model. Ask one to write a thought-leadership article about AI. Give the other years of customer research, published positions, expert interviews, performance data, past editorial decisions, brand constraints and a structured argument to investigate before it writes anything. They technically have the same AI. They don’t have the same capability, because the difference sits in the system surrounding the model rather than the model itself — and as models converge, that surrounding system is doing more of the work, not less.

I’ve made a related case in context is capital: access to any particular model depreciates fast, sometimes within a year, while context captured properly keeps being useful to whatever model comes next. That’s also why I’m unconvinced by AI strategies built mainly around access to one frontier model. Access is temporary. Context compounds.

Routing is becoming part of the job

One implication I hadn’t fully named until now: professional AI use is starting to look like routing rather than prompting. Prompting dominated the first wave of generative AI because everything happened inside a single conversation and the skill was asking one model well. Once there are multiple capable models, agents, tools and repositories in play, the more important decision becomes where a given piece of work should go — an inexpensive fast model for straightforward transformations, a stronger reasoning model once a problem gets genuinely difficult, a different model deliberately used as an adversarial reviewer, a specialised tool for code, escalation only when the work actually warrants it.

I’ve already been doing a crude, manual version of this without naming it as such: one model develops an argument, a different one is asked to try to break it, and I decide what survives both — the reasoning behind that split is in the second model is not a proofreader. The useful part was never discovering which model is objectively smarter. It was assigning different roles to different systems, and routing is the general principle sitting underneath that specific habit.

The more interchangeable capable models become, the more that principle does the work that picking a favourite used to do — and eventually I probably shouldn’t be doing all of that routing by hand either. A routine request can go to a cheap model, a complex analysis to a stronger one, a high-stakes final check to an independent one, and whatever neither can solve escalates to frontier capability. That’s a much better description of an operating system than of a chatbot.

The benchmark that matters

Years of AI coverage have trained me to watch model releases like processor launches — which model is on top, who’s overtaken whom, how big is the context window, what’s the benchmark score. I still find that interesting. But increasingly the question I actually ask is narrower: does this change anything I can actually do? If the answer is no, the benchmark improvement can matter enormously to researchers while barely touching my own workflow, and businesses eventually ask the same thing about every technology — not is this the most advanced product available, but what additional outcome am I buying. That question gets harder for frontier AI to answer every time the cheaper tier improves.

Capability is becoming abundant. Context, judgement and orchestration aren’t — and that’s a claim I wouldn’t have made about generative AI a few years ago, when basic capability was itself the scarce, breakthrough thing. Now several extraordinarily capable systems are available to me for the price of a few software subscriptions. If one fails a task, I can move it. If another gets better, I can switch to it. If the expensive one genuinely unlocks something the others can’t do, I can escalate to it.

None of that makes the underlying models unimportant. It changes what’s worth building around them. I don’t want Editorial Intelligence to work because Claude happens to be good this quarter, or because ChatGPT is — I want a durable layer of evidence, context, decisions, standards and workflow, with whichever models are actually useful attached to it.

Today that might mean Claude here and ChatGPT there. Next year the names attached to those roles may change. That’s fine, because if the system survives the substitution, the system is where the value was all along.

That may be what the surprisingly thin demand for the world’s most powerful AI model is starting to tell us: the next stage of professional AI use isn’t about buying access to more intelligence at any price. It’s about getting better at recognising how much intelligence a given problem actually needs.

Topics

aiai-workflowseditorial-intelligencepractical-aithought-leadershipportable-intelligence

What to explore next

See how the ideas in this Field Note connect to the frameworks, diagnostics and workflows in Editorial Intelligence OS.

Explore the EI OS →

Keep in touch with Editorial Intelligence

Occasional updates on new research, findings and ways to take part.

No spam. Unsubscribe in one click.

Your address is used only to send these updates. Read the privacy policy.