reflection 24 min read

Editorial Intelligence vs FPL, part three: I gave AI all the context. It still got it wrong.

24 August 2026


Part of Editorial Intelligence vs FPL, a running experiment in evidence, judgement, football and occasionally making terrible decisions.

I said before the season started that I wasn’t going to judge this experiment purely on Fantasy Premier League points.

That is now extremely convenient.

Editorial xG finished Gameweek 1 with 38 points.

The average was 48.

João Pedro scored 11 of them.

My bench scored 17.

There are probably better advertisements for artificial intelligence.

But there may not be many better demonstrations of what I actually wanted to test.

This experiment was never about asking ChatGPT to pick a Fantasy Football team.

I wanted to see what happened when I applied something closer to an Editorial Intelligence system to a problem with lots of competing information.

We looked at fixtures, prices, expected minutes, roles, pre-season performances, expert opinion and changes in player ownership.

ChatGPT interpreted the evidence and made recommendations.

Claude was brought in to attack some of ChatGPT’s reasoning.

I questioned recommendations myself.

And, importantly, I wrote the decisions down before the matches happened.

The loop was:

information → signals → interpretation → decision → outcome → learning

There was quite a lot of intelligence involved.

Unfortunately, there weren’t many fantasy football points.

The AI made a very convincing mistake

The clearest example is Gabriel versus Riccardo Calafiori.

The broad idea was fine. I wanted an Arsenal defender. The difficult bit was deciding which one.

Gabriel cost £8m. Calafiori cost £5.5m.

Calafiori had shown attacking promise and his ownership was rising sharply. I noticed that. I questioned whether I should pick him.

ChatGPT argued strongly for Gabriel.

The reasoning wasn’t stupid. Gabriel had a longer track record of Fantasy returns. His position appeared secure. He carried set-piece threat. There were legitimate questions about Calafiori’s role and minutes.

I want to be precise about what that reasoning actually said, because I have it in writing rather than relying on memory. Part two recorded it before a ball was kicked: “Calafiori became the fashionable cheaper alternative after scoring in the Community Shield, but ownership movement is evidence of what other managers believe, not proof that Gabriel has become a worse asset.”

So I followed the recommendation.

Arsenal beat Coventry 3–0. Clean sheet. Gabriel took his share of it: five points.

Calafiori took a bigger share. He set up Kai Havertz’s opener, and finished with nine.

The four-point difference is almost the least interesting part.

Calafiori also cost £2.5m less. That money could have improved another position.

The model hadn’t hallucinated anything. It had relevant evidence. Its reasoning was coherent. It simply gave the variables the wrong weight. Historical performance and perceived security mattered too much. Price and opportunity cost mattered too little.

And I think we made another mistake.

I pointed out that lots of Fantasy managers were suddenly selecting Calafiori. ChatGPT essentially told me popularity wasn’t a reason to pick somebody.

That’s true. But it wasn’t quite the right question.

A crowd isn’t automatically correct. But when a lot of independent people suddenly move in the same direction, something may have changed.

The useful response isn’t: follow them.

It also isn’t: ignore them.

It’s: why are they doing that?

That may be the first useful rule to come out of this experiment. Consensus isn’t truth. But a sharp movement in consensus is a signal worth investigating — not a signal worth obeying, and not one worth waving away either.

Then I spent ages debating Bruno Fernandes versus Haaland

This may be the funniest decision of the week.

After Manchester City’s poor Community Shield result, my instinct was to remove Erling Haaland. ChatGPT argued against it. One game wasn’t enough evidence to dismantle a squad built around a £15.5m player. The reigning FPL champion, independently, made essentially the same case. Fair enough. We kept him.

But we reduced our confidence enough to captain Bruno Fernandes instead.

Claude then spotted something interesting. If the Community Shield wasn’t enough evidence to remove Haaland, were we being inconsistent by allowing it to influence our captaincy decision?

That forced us to improve the argument. Owning Haaland was a multi-week structural decision. Captaincy was a one-week decision. Bruno appeared to have the more attractive fixture. So Bruno got the armband, on cleaner logic than the one we started with.

Manchester United lost 2–0 away at Hull.

Bruno scored two points. Haaland scored two points.

After all that analysis — after finding and fixing an actual hole in our own reasoning — the decision made absolutely no difference to the score.

There is probably a lesson in that too. Sometimes the thing you spend the most time optimising turns out not to be the thing that determines the result. Good reasoning bought us a better argument. It didn’t buy us a better outcome. Those are different things, and this is as clean a demonstration as I could ask for.

The Manchester United bet was worse

The bigger issue wasn’t really Bruno versus Haaland.

It was how much of my team rested on one broader assumption.

I had Bruno. I had Mbeumo. I had Maguire. Three players relying, to different degrees, on the idea that Manchester United had an attractive opening fixture.

Hull City — newly promoted, playing their first home Premier League match in years — beat them 2–0. Both goals came from set pieces. Maguire, whose whole job that day was defending exactly that scenario, couldn’t stop it.

Their returns: Bruno two, before captaincy. Mbeumo two. Maguire one.

That wasn’t a marginal miss.

Here’s the more useful way to say what actually went wrong, though. It isn’t that the fixture read was crazy — newly promoted sides beating established ones on opening day happens often enough that treating Hull as “safe” wasn’t an unreasonable call in isolation. The real mistake was concentration. Three players, same ninety minutes, same assumption. If the read is right eighty per cent of the time, stacking three assets on it means the twenty-per-cent version costs you three times what it should. That’s true whether or not the original judgement about Hull was defensible, and it’s a cleaner lesson than “we got the fixture wrong.”

I need to be careful here too. One game doesn’t suddenly prove Manchester United assets were irrational selections. Otherwise I’d simply be making the opposite version of the mistake I was trying to avoid before the season: treating every new result as overwhelming evidence.

That’s where this experiment gets more interesting than: AI picked a bad Fantasy team.

I need to distinguish between a bad outcome and a bad decision.

My bench scored 17 points

This deserves its own section because it is objectively funny, and because one bench player is directly responsible for the disaster above.

Egan scored eight. Thomas scored three. Slater scored six.

That’s seventeen. The entire starting XI finished with 38.

Here’s the number that actually stings. My £15.5m forward and my captain — the two decisions I spent the most time reasoning about all week — combined for four points before the armband did its doubling. My bench defender, who cost £4m and was picked mainly to be forgotten about, beat each of them individually on his own.

There’s a sharper irony underneath that one, too. Slater plays for Hull. He set up both goals that beat Bruno, Mbeumo and Maguire. The same ninety minutes that wrecked a third of my starting XI is the exact match that fed my bench. I didn’t get outscored by my bench through bad luck spread randomly across the gameweek. I got outscored by the specific player who did the damage to my own team.

That does not mean I should obviously have started him. Fantasy Football makes hindsight extremely seductive — once the points exist, every decision looks obvious, and it wasn’t obvious beforehand. The bench was deliberately cheap because the strategy was to concentrate money in the starting XI. That’s still a defensible strategy. Sometimes cheap players score points on the wrong week, for you or against you. That’s football.

So the bench result is mostly variance, dressed up as irony.

Compare that with Gabriel versus Calafiori. There, I can actually identify a weakness in the reasoning that existed before the outcome — the price logic, the ownership-signal dismissal. Here, I mostly can’t. That’s an important distinction. Bad result doesn’t automatically mean bad decision. But equally: good reasoning doesn’t guarantee good judgement.

One pick escapes both problems, and deserves saying plainly. Igor Jesus was already flagged, in writing, before a ball was kicked, as “the weakest starting assumption, consciously accepted” in the whole XI — the risk was named risk of him losing his place. He started. He had two good headed chances, one well saved, one nodded over. Forest actually outshot Leeds on expected goals and lost 0–1 anyway. The specific risk we named didn’t happen. A different kind of bad luck did. That’s not a failure of the process. It’s what the process is supposed to sound like when it works and the result still goes wrong.

So did AI actually work?

As a Fantasy Football selector in Gameweek 1? Not very well. Thirty-eight points against an average of 48 isn’t much of a victory lap.

But something else did work.

Because I recorded the reasoning before the games, I can inspect the failure — including the parts of it I’d have sworn, from memory alone, weren’t written down. Without that, I’d probably do what everyone does after a bad Fantasy week. I should have picked him. Obviously that player was a mistake. Why did I captain him?

Instead I can ask: what did we actually know at the time? What assumptions did we make? Which evidence did we overweight? Which signals did we dismiss? Was the reasoning poor, or did a reasonable decision simply produce a bad outcome?

That turns AI from an answer machine into something closer to a decision instrument.

And that is where this stopped being mainly about Fantasy Football for me.

Fantasy Football should actually be easier than editing

Think about the environment. Fantasy Football has explicit rules. Players have prices. Fixtures are known. There are statistics. There is a fixed budget. And eventually somebody gives you a number telling you how your decisions performed.

Writing doesn’t even give us that.

Recently I was looking at an Economist-related LinkedIn post. I thought it had a strong argument. It suited LinkedIn. It generated good engagement. Two commenters said it looked AI-written.

Who’s correct? Potentially all of them. The engagement doesn’t prove the writing was good. Two negative comments don’t prove it was bad. And “looks AI” isn’t an objective quality metric anyway — it’s an audience response. That makes it useful information. It doesn’t make it a verdict.

And suddenly we’re back at Calafiori.

The fact lots of people selected Calafiori didn’t prove he was the right player. But it should have made me investigate why they were doing it. Two people saying a post “looks AI” doesn’t prove the post has failed. But it may tell the writer something about how part of the audience is interpreting the style.

The mistake in both situations is either blindly following the signal or dismissing it. The useful question is: why are people reacting this way, and how much should that reaction matter for what I’m trying to achieve?

Here’s where the comparison actually breaks, and it’s worth saying so rather than pretending it holds all the way through. Football has a scoreboard. Eventually, points settle the Calafiori question — not this week, but over enough weeks, reality forces a correction. Writing mostly doesn’t have that. Nothing ever definitively proves a piece was authentically good or not. A wrong read on a footballer gets fixed. A wrong read on a piece of writing can sit there indefinitely, unresolved, with nothing forcing anyone to revisit it. That’s not a flaw in the comparison. It’s the actual reason editorial judgement is the harder problem.

Now imagine trying to automate that

This is where I’m becoming sceptical about the idea that sufficiently powerful AI eventually removes the need for editorial judgement.

Not editorial production. AI is obviously going to automate enormous amounts of that. Research assistance. Summaries. Drafting. Rewriting. Formatting. Repurposing. Variations. Probably lots more.

It can also increasingly automate parts of editorial analysis. Compare these two headlines. Identify weaknesses in this argument. Check whether this paragraph contradicts the brief. Show me how different audiences might interpret this. Find gaps in the evidence. All useful.

But editorial judgement sits somewhere above that.

Which audience matters most? Whose objection should change the piece? What trade-off are we willing to make? Does greater accuracy justify greater complexity? Does stronger brand consistency justify something less distinctive? Does an executive’s preference outweigh what the audience appears to respond to? How much should “looks AI” matter if the piece is actually accomplishing its objective?

There isn’t necessarily one correct answer waiting for a sufficiently intelligent model to discover. There are too many audiences, too many styles, too many objectives and too many people with different opinions about what good looks like.

And every organisation adds another layer. The writer has an opinion. The editor has one. The subject expert has one. Marketing has one. Brand has one. Legal has one. The executive has one. The customer has one. Then the person commenting on LinkedIn has another.

Even if an AI could model every one of those perspectives perfectly, somebody would still have to decide which perspective matters most — and be the one who answers for it if that call turns out to be wrong. That’s the part I don’t think gets automated away by a smarter model. Not the modelling of the perspectives. The standing to choose between them, and the accountability that comes with the choice. That isn’t purely an intelligence problem. It’s a legitimacy problem, and legitimacy isn’t something a model can grant itself.

AI doesn’t need to be stupid to be wrong

Most discussion of AI failure still gravitates toward obvious mistakes. Hallucinated facts. Fabricated sources. Generic prose. Odd phrasing. Those matter.

But I’m increasingly interested in a harder problem.

The evidence is real. The reasoning is coherent. The recommendation is persuasive. And the judgement is still wrong.

That is what happened with Gabriel and Calafiori. The scary bit isn’t that ChatGPT said something ridiculous. It didn’t. It gave me a perfectly plausible explanation for making the wrong decision. And I accepted it.

That is much closer to the problem organisations are going to face as these systems improve. AI doesn’t have to hallucinate to mislead you. It can simply be extremely convincing while assigning the wrong importance to one part of a complicated problem.

Humans do this too, of course. That’s important. This isn’t an argument that humans possess some magical judgement AI could never approach. Humans are inconsistent, biased and regularly wrong. I just scored 38 points with a team I helped select.

But replacing one imperfect judge with another imperfect judge isn’t the interesting opportunity. The opportunity is using AI to make the whole process better.

Maybe the goal isn’t automated judgement

The conclusion I’m arriving at is slightly different from where I started.

I don’t need AI to make judgement infallible. I want it to make judgement better informed, more explicit, easier to challenge, easier to inspect afterwards.

Show me the evidence. Show me the alternative. Tell me which assumptions you’re making. Tell me which variables you’re weighting heavily. Tell me what evidence would change the recommendation. Then make the human decision visible. And once reality answers back, learn from it.

That is much closer to what I mean by Editorial Intelligence. Not: AI decides. More like:

information → competing interpretations → judgement → decision → outcome → learning

AI can be involved throughout that system without pretending it has removed uncertainty from it.

Some people will look at the amount of thought going into a fantasy football team and understandably think this is overkill. It is. That’s also why I like it as an experiment. The point isn’t that the team deserves this much analysis. The point is that the stakes are low enough for me to watch how AI reasons, where it fails, where it challenges me, and how its judgement coexists with my own without either side pretending to be infallible.

Which brings me to Gameweek 2

The worst thing I could do now is react to 38 points by tearing everything apart. That would undermine the entire experiment.

And this is where AI may turn out to have been more useful than the score suggests.

My normal instinct after a week like this would be to act. Chase the players who just scored. Promote the bench. Consider several transfers. Maybe convince myself a chip could repair the damage. The bad result creates a feeling that something must be changed immediately.

The AI-assisted process is currently doing the opposite. It is forcing me to ask which changes are actually justified by new evidence.

So my plan for Gameweek 2 is deliberately boring: keep the squad largely intact, avoid a reactive rebuild, and give the original assumptions another data point. Gabriel versus Calafiori stays under review rather than becoming an automatic switch because Calafiori scored nine. Igor Jesus stays in the squad rather than being sold because he blanked. The Manchester United concentration gets another test rather than being dismantled after one bad afternoon.

I’m not going to publish another pre-Gameweek squad post this week either. The more interesting experiment is to see what happens when a bad first result is followed by restraint rather than activity, then review it properly afterwards.

That creates a different test of AI from Gameweek 1. The first week showed the danger of trusting a persuasive model recommendation too readily. Gameweek 2 may show the opposite benefit: AI as a brake on human overreaction.

Not: trust the AI.

More like: use AI to challenge the action you feel most compelled to take.

The proper season is starting to create stories

There is another thing I hadn’t really anticipated before the season began.

Pre-season gave me signals and hypotheses. Once the competitive football actually started, those hypotheses began turning into stories.

Palmer looked good, and suddenly there is a plausible Palmer is back story. Chelsea looked dangerous enough for João Pedro to rescue my week, and suddenly Chelsea attack feels like something I may want more of. Bruno blanked inside a terrible United result and suddenly the story becomes United are a problem. Haaland blanked after I had already spent a week questioning whether the £15.5m structure was worth it, so that story gets another chapter too.

Hull are suddenly more than the “easy fixture” they were in a model. They are the team that wrecked three of my starters while Slater sat on my bench collecting six points from the same match.

And Arsenal’s 3–0 start creates a potentially bigger question than whether I should have owned Gabriel or Calafiori: what if Arsenal are simply going to dominate? If that story keeps developing, I may eventually need to think about Saka, Ødegaard and where the best Arsenal exposure actually sits rather than treating this as a defender-selection problem.

None of those stories is established yet.

That’s important.

The proper season has only just started, so the narratives are only just beginning to take shape.

“Palmer is back.” “United are broken.” “Arsenal will dominate.” “João Pedro is essential.” “Haaland isn’t worth the money.” They are all attractive stories after one week. They may become true. They may be almost completely wrong.

And that gives me another Editorial Intelligence problem to test.

Narrative isn’t just something you add to the end of analysis to make the article interesting. Narrative is one of the ways humans organise evidence in the first place.

A player becomes the safe premium. The comeback. The bargain everyone else spotted. The expensive mistake. The one who saved the week. A team becomes dominant, unreliable, underestimated or broken.

Those stories are useful because they make a complicated environment legible. They are dangerous because once the story feels right, every new piece of evidence can start getting fitted into it.

So I want to start tracking the narratives as explicitly as I track the numbers.

What story am I telling about this player or team? What evidence created it? What doesn’t fit? What’s the strongest alternative story? What would strengthen it? What would make me abandon it? And, crucially, has the story actually become strong enough to justify an action?

That changes the loop slightly:

signals → emerging narrative → adversarial test → decision threshold → action or hold → outcome → narrative update

That feels surprisingly close to editorial work.

Businesses do this constantly. They construct stories about what customers want, why a campaign worked, what a competitor is doing, what a product stands for, why an audience is changing. AI can help gather the evidence and strengthen those narratives. It can also make a weak narrative sound extremely convincing.

Perhaps another useful role for it is to keep asking what doesn’t fit the story yet.

The live stories I’m watching

Palmer versus Bruno. Palmer looked good. Bruno and United looked terrible. But United have another favourable fixture. The question is whether the Chelsea and United stories have genuinely diverged enough to justify moving premium midfield budget, not whether Palmer had the nicer opening afternoon.

João Pedro and Chelsea. His 11 points reinforce the case that already existed before GW1, but one excellent return is still only one chapter. If Chelsea’s attacking structure and his role stay strong, the story becomes more durable.

Haaland. The premium-anchor story survived the Community Shield because one game wasn’t enough evidence. A two-point GW1 does not suddenly make the opposite story true either. The question is whether the £15.5m structure continues to earn its opportunity cost over several weeks.

Hull and United. Hull were supposed to be the easy fixture. They became the antagonist of the entire week. United now get another favourable opponent, which gives the original thesis a fair second test rather than an automatic funeral.

Arsenal. The Calafiori mistake matters, but a larger storyline may be emerging: Arsenal might be strong enough that the real question becomes how much of the team should be exposed to them. Saka and Ødegaard now belong on the serious watchlist alongside the defender/value debate. If the dominance thesis persists, the right response may be a budget-allocation decision rather than simply swapping one Arsenal player for another.

If the largely unchanged team rebounds, that does not automatically prove the original squad was good. If it fails again, that gives me a second evidence point and a much stronger case for structural changes. Either way, the next review becomes more useful because I will know I didn’t simply chase the previous week’s points.

Gameweek 1 gives me evidence. It doesn’t give me certainty.

There are decisions I still need to review. The Gabriel versus Calafiori choice clearly deserves another look, tested against the actual written reasoning rather than a reconstructed memory of it. The Manchester United exposure needs challenging, specifically on concentration, not just on whether Hull was a fair read. The cheap bench may deserve slightly more respect. João Pedro, who scored inside the first minute against Chelsea and delivered exactly what his pre-season case predicted, has certainly made his.

Here’s what would actually change my mind, rather than just confirm what I already think happened. On Calafiori: not one more good week, but a real gap over the next three or four gameweeks at that price difference — one week is Slater’s afternoon, not a pattern. On the bigger question — whether AI can do this — the thing to watch for isn’t another wrong pick. It’s whether the human sign-off in this process ever starts feeling automatic. The day I stop finding anything to disagree with is the day to worry, not the day to relax.

The question for Gameweek 2 isn’t: who scored last week?

It’s: what did Gameweek 1 actually teach me that should change a decision?

Which is inconveniently close to the question I’m increasingly asking about AI at work.

Fantasy Football 1. Editorial Intelligence 0. For now.

Topics

editorial-intelligenceevidenceaipractical-aiknowledge-systems

What to explore next

See how the ideas in this Field Note connect to the frameworks, diagnostics and workflows in Editorial Intelligence OS.

Explore the EI OS →

Keep in touch with Editorial Intelligence

Occasional updates on new research, findings and ways to take part.

No spam. Unsubscribe in one click.

Your address is used only to send these updates. Read the privacy policy.