Why one AI model was never going to be enough for writing
24 August 2026
For a while, I thought the question was which AI model was best at writing. ChatGPT or Claude. One model for drafting, another for editing. Try the latest release, compare the prose, decide which one sounds least like AI.
I’m increasingly convinced that’s the wrong question — not because the models are becoming identical (they aren’t, and I still find different models useful for different things), but because there may be no such thing as the best model for writing. There’s no single definition of good writing for it to be the best at.
I’m not sure there ever will be.
Writing isn’t one task
We talk about writing as though it’s a single operation: give a model a brief, get some words back, judge the words. Professional writing isn’t really like that. Before anything gets published, somebody might have to find the interesting thing, understand what the organisation actually knows, research it, decide what the evidence supports, choose an angle, work out what the audience needs, fit that into a larger narrative, satisfy stakeholders, draft it, challenge it, simplify it, preserve the necessary complexity, make it sound like the organisation, make it sound like a person, adapt it for search, adapt it for AI search, turn it into a LinkedIn post, an email, a video script — and decide whether any of that should exist at all.
Calling all of that “writing” hides most of the work. I’ve made a version of this argument before, about what the finished article hides — the words were never the whole value makes the fuller case. What I want to add here is what it does to the model question specifically: one model might produce a beautifully structured first draft but be too agreeable when I want the argument attacked. Another might be excellent at finding holes but make the resulting prose worse. A cheaper, faster model might be perfectly adequate for turning an existing argument into another format, while a more capable reasoning model would be wasted on that job but valuable earlier, when I’m still working out whether the argument survives at all.
That’s increasingly why I think prompting is becoming routing. The interesting question isn’t always what should I ask the AI. It’s where should this piece of work go next.
There isn’t even one audience for “good writing”
The model problem gets harder because writing doesn’t have an objective finishing line. A calculation can be correct. Code can run. A factual claim can be verified. Writing can be grammatically perfect, factually accurate and completely wrong for the job.
A subject expert may want qualification and complexity. A time-poor executive may want the same argument reduced to three sentences. A brand team may want consistency; a writer may want distinctiveness. A search engine may reward explicit structure; someone scrolling LinkedIn wants the point immediately. One stakeholder thinks the draft is too informal, another thinks it sounds corporate, a third hates anything that reads remotely AI-generated, and a fourth doesn’t care how it was produced as long as it tells them something useful. Increasingly there’s another audience sitting alongside all of them: machines retrieving, summarising and citing the work.
There isn’t one universally correct version waiting to be discovered. There are choices — and a model can only ever be evaluated against one of them at a time.
Can you tell whether AI wrote it? I’m less sure that matters
We’ve become strangely obsessed with detecting AI writing — the em dash, the neat groups of three, the slightly frictionless rhythm. Some of that criticism is useful; models have defaults, and publishing them untouched produces bland writing. But there’s a bigger question behind the detection game: if you correctly identify that AI produced a sentence, what have you actually established? Not where the idea came from, not whether the evidence is original, not whether anyone exercised judgement over what got published. You’ve established something about how the sentences were produced — which used to feel synonymous with authorship because sentence production was expensive. Now it isn’t.
I’ve previously broken “wrote” into four different questions — where the idea originated, what evidence supports it, who decided it was worth saying, and who produced the sentences — in the words were never the whole value. AI detection mostly interrogates the last one. Increasingly, I think it’s the least interesting.
Human writing can be slop too
This changes how I think about AI slop, too. It obviously exists, and AI makes the economics of producing it dramatically worse. But “AI-generated” and “slop” aren’t synonyms — the opposite of AI slop isn’t human writing, it’s original insight, and I’ve argued that case in full elsewhere: we spent years before generative AI producing search-driven filler that a human researched, typed and edited, and it could still add almost nothing.
That distinction matters here specifically because otherwise we optimise for the wrong thing in a model — training it to disguise its stylistic fingerprints while leaving the underlying work just as empty. I’d rather read something obviously AI-assisted that contains an observation, a dataset or an argument I couldn’t get elsewhere than beautifully human-sounding prose that tells me nothing.
The scarce thing moves upstream
This is where the model question gets more interesting. If I can generate twenty competent versions of an article in a minute, my problem is no longer producing enough words. My problem is deciding which version deserves to exist — and that moves value upstream, into what the organisation knows, what it can prove, what hasn’t already been said, and what belongs in the argument it’s building; and downstream, into what’s credible, distinctive, right for the channel, and what should be cut or refused outright. Those are editorial questions, not typing questions, and I’ve made the wider version of this case — that AI didn’t replace editorial judgement, it made it more valuable — before. AI didn’t suddenly make that upstream work valuable. It made much more of it affordable, which is a significant difference.
Context may matter more than the model
There’s another reason I’m losing interest in finding the perfect writing model: the model is only part of the system. Give an excellent model a generic prompt and it has generic material to work with. Give it accumulated research, interviews, customer evidence, previous arguments, brand decisions, rejected ideas and a clear narrative, and you’ve changed the problem entirely. Context is capital makes the fuller case for why that accumulated material is the durable asset and the model diffuses too quickly to be one.
What I’ve started doing differently because of it is running a test before I publish anything: what is in this piece that only I could have supplied? Not an AI-generated interpretation of a public argument, but an observation, a conversation, some work I’ve actually done, a number from material I have access to, something I noticed. If I can’t find that, changing models isn’t going to rescue the article.
This is why I want models to disagree
It also explains why I don’t want one perfect model. If ChatGPT drafts something and Claude independently tells me it’s excellent, that’s reassuring — it isn’t necessarily useful. Sometimes I want the second model to find the reason the first one is wrong, or to challenge an argument I like, or to leave the expensive reasoning model alone entirely and let a fast one handle an uninteresting transformation. I’ve set out the fuller case for that split: the point of a multi-model workflow isn’t an AI beauty contest, it’s introducing different jobs, different constraints and, when it’s useful, disagreement. You wouldn’t hire six copies of the same editor and ask all of them whether they liked the headline.
So where does the writer go?
There’s an appealing conclusion available here: all this complexity means writers are safe. I don’t think that’s defensible. AI can remove a lot of writing labour, and some writing jobs existed partly because producing competent sentences at scale was expensive — that economic protection is disappearing. But the whole job doesn’t disappear with it. The valuable part shifts: the writer becomes more researcher, interrogator, editor, narrative builder, audience advocate, system designer and decision-maker. Maybe “writer” eventually becomes an increasingly inadequate description of the job, in the same way the job itself is moving upstream more broadly.
The strange thing is that better writing models could accelerate that shift rather than end it. The closer every major model gets to producing excellent prose, the less interesting the question which AI writes best becomes, and the more interesting question becomes what surrounds it: the knowledge it can access, the evidence behind the claim, the narrative the organisation is building, the audience being served, the other models or humans allowed to challenge it, and the person who ultimately takes responsibility for deciding — this is worth saying, this is the version we’re going to say, and this is why.
There may never be one perfect AI writer. There was never one perfect human writer either. Writing has always involved too many competing objectives, audiences and judgements for that. The difference now is that we can make many more versions, much faster. Which makes producing words easier, and choosing what deserves to survive harder.
What to explore next
See how the ideas in this Field Note connect to the frameworks, diagnostics and workflows in Editorial Intelligence OS.
Explore the EI OS →Keep in touch with Editorial Intelligence
Occasional updates on new research, findings and ways to take part.
Almost there — check your inbox.
A confirmation email is on its way. Your address is only added to the list once you click the link in it, so if it does not arrive, nothing has been signed up.