Use one AI to build the argument. Use another to break it.
21 August 2026
Asking a second model to improve a draft is hiring another copy editor. Asking it to check the draft against everything you’ve already said, and tell you whether it should exist at all, is something else.
I originate the idea. ChatGPT develops it — talks it through, connects it to everything else I’ve written, argues with me until it holds together or falls apart. What survives that goes into the repository. Claude Code meets it there, reads it more coldly, checks it against the corpus, and tries to break it. Then I decide what publishes.
I didn’t design that pattern in advance. I noticed it after the fact, the way you notice a habit once somebody points it out.
Which AI is best is the wrong question
Most discussion about using two frontier models treats it as a comparison exercise. Which one writes better. Which one reasons better. Which one is worth the second subscription.
Those aren’t bad questions. They’re just a smaller question dressed up as the interesting one. The more useful version isn’t which AI is best — it’s which role should each AI perform in this piece of work.
Traditional editorial teams didn’t separate writing, editing, fact-checking and commissioning because any one of those people couldn’t do the others. They separated them because an independent pass catches what the person who produced the work cannot see in it. Generative AI makes it tempting to collapse all of that back into one step: generate, publish. That’s efficient, and it quietly asks the system that produced an argument to also approve it.
Splitting the roles across two models doesn’t reproduce human independence — they can share training patterns, biases, the same blind spots. But asking one model to develop an argument and a different one to challenge it still produces more friction than asking either to do both, and friction is exactly what a generate-and-publish workflow removes too easily.
What the brief actually needs to say
The distinction that matters is in the instruction, not the model.
Can you make this better? is a proofreading brief regardless of which AI receives it. It improves sentences and leaves the argument’s assumptions untouched.
Check this against the repository. Where does it contradict something I’ve already published? Is it duplicating an existing piece? What’s the evidence for this claim, and what am I overstating? Should this exist at all? is a different job. It requires the second pass to actually read the corpus rather than the draft in isolation, which is the reason this only works because the accumulated context — the frameworks, the previous arguments, the positions already taken — lives somewhere a second system can get to it. Why every company needs a GitHub for knowledge is the infrastructure underneath this piece: without a durable, structured record of what’s already been said, a second model can only critique the sentence in front of it, not the argument’s place in a larger body of work.
An example where this actually caught something
The clearest evidence I have is a case where it worked and then immediately showed its own limits.
I’d developed three Field Notes with ChatGPT and didn’t like them without being able to say why. Claude’s diagnosis was specific: the ideas were mine, but the rhythm had quietly become the model’s — stacked fragments, bold phrases standing in for argument, lists where connected reasoning should have been. I took that back to ChatGPT, which updated its own instructions and produced something new — that came back as dense, hedged and over-qualified, a caricature of Claude’s register instead of a correction to its own. The full account is here, including the part I keep coming back to: I’d passed on a diagnosis without saying how far it should reach, and a model has no independent sense of proportion to correct that for itself.
That’s not a story about one model being better than the other. It’s a story about the second pass finding something real — a pattern invisible to the process that produced it — and then needing a human to scope the correction, because neither model could tell “this paragraph is too fragmented” from “abandon this register entirely.” The friction did its job. It didn’t do the whole job.
Then it found something that wasn’t about prose
The cadence catch was useful, but modest — the kind of thing a good copyeditor does. A more recent case was a different order of catch.
Before Gameweek 1 of a small side experiment I’ve been running — Editorial Intelligence vs FPL — I gave the second pass the same brief on a decision instead of a draft: attack the reasoning, don’t help pick a better answer. It found that I’d ruled a piece of evidence too weak to justify one call, then quietly relied on the same evidence to justify a different call going the other way. Not a tone problem. A hole in the argument, sitting in my own decision log, that I hadn’t seen because I already believed the conclusion.
That’s the actual test this note set for itself — whether the second pass keeps finding things worth finding, or settles into confirming what the first pass already produced. So far it’s finding bigger things, not smaller ones: a style habit, then a genuine reasoning error. That direction matters more than either individual result.
This isn’t just how I happen to work
Microsoft has built something close to this directly into Copilot — a mode where one model drafts and a second is specifically tasked with verifying it, sold as a check rather than a second opinion you have to go looking for. There’s research behind the same instinct: debate-driven verification between models, where one agent cross-examines another’s claims, measurably improves how reliably factual errors get caught.
Two models agreeing isn’t evidence of anything
The obvious objection is that none of this proves the second model is right, or that disagreement between them means anything beyond style preference. It doesn’t. Two models converging on a view is not evidence the view is true — they may simply share the same tendencies. There’s research behind this caveat too: in a debate between models, the more confident-sounding argument can win even when it isn’t the correct one, because confidence and correctness aren’t the same signal. Human judgement stays the deciding layer regardless of how many passes came before it.
The other honest caveat is that I can bias both models through how I frame the task, and a determined author can walk either one toward whatever answer they want. A second opinion you can quietly steer isn’t independence. The value here is closer to a discipline than a guarantee: it’s harder to skip an editorial stage you’ve deliberately built two different tools to occupy than one you could just as easily do yourself in five minutes.
Nor should this become mandatory process. A typo fix doesn’t need a two-model editorial board, and treating every small correction as if it did would turn a useful discipline into performance. It earns its place on the work where an unsupported assumption, an unnecessary framework, or a claim that overstates its evidence would actually cost something.
Having the capability is not the same as using it
Sage, my full-time employer, gives people Microsoft Copilot — not to be confused with Sage’s own product of the same name — and that same draft/verify pattern is sitting right there inside it. So is most software a lot of people already use.
The capability isn’t rare any more. What’s rare is somebody actually invoking it — every time, not just the time they remember, and especially the time it would be more convenient not to. A feature that exists but isn’t triggered does nothing at all; it’s functionally identical to a feature that was never built. Nothing about a tool sitting inside a Microsoft tenant forces anyone to route a draft through a critique pass before acting on it. That decision stays entirely human, which is exactly why it’s the one that actually matters.
This is a case for more editorial scrutiny, not less
There’s a common argument that using AI to help write something is evidence of not caring much about the writing. Sometimes that’s true — generate-and-publish is perfectly capable of producing a lot of indifferent material quickly.
But manually typing every sentence was never a reliable measure of care either. A workflow that deliberately restores commissioning, argument development, independent challenge and human sign-off — even with AI doing some of the work at each stage — can involve more editorial process than one person drafting alone with no time for an edit. The question that actually matters isn’t whether AI touched the sentences. It’s what process produced the argument, and whether anything independent had the chance to find its weaknesses before an audience did.
It’s also a smaller instance of a bigger claim I’ve made elsewhere: editorial judgement doesn’t have to disappear after the edit — it can become infrastructure for the next piece of work. A standing assignment to attack a draft, rather than a one-off request, is that same principle applied to a single role instead of a whole team.
Holding this loosely
I’m treating this as an emerging working practice, not a settled method — the same way I’d want anyone reporting a result from a sample of one to treat it. The test I’m actually running is whether the second pass keeps finding things worth finding across the next several pieces, or whether it settles into confirming what the first pass already produced. If it’s the second, the discipline has become theatre and should be dropped rather than defended.
Which is itself the argument, working as intended: a case for keeping a stage that AI makes tempting to skip should not get to skip its own evidence requirement either.
What to explore next
See how the ideas in this Field Note connect to the frameworks, diagnostics and workflows in Editorial Intelligence OS.
Explore the EI OS →Keep in touch with Editorial Intelligence
Occasional updates on new research, findings and ways to take part.
Almost there — check your inbox.
A confirmation email is on its way. Your address is only added to the list once you click the link in it, so if it does not arrive, nothing has been signed up.