Human-AI Communication

The AI Reviewer Only Sees What You Hand It: Three Ways a Verdict Fails

The AI Reviewer Only Sees What You Hand It: Three Ways a Verdict Fails

An AI reviewer is any model you ask to check a draft before it goes out, and its verdict is only as good as the briefing it receives, because it judges the draft against whatever you put in front of it. A manager who pastes a memo into a chatbot and asks whether it holds up is using one. So am I, on everything I publish. In September mine passed a draft it should have stopped, and three weeks later flagged real facts as wrong. Both errors came from what I left out of the brief.

The reviewer that passed the wrong draft

I run two checks on everything that goes out under my name. The first is a scanner, a script that looks for the patterns readers cite when they call writing machine-made and returns a score and an exit code. The second is an editor: a separate AI model that starts with a fresh context, can read files but cannot change them, and returns PASS or FAIL with line-level fixes. If you have ever asked a chatbot whether a draft holds up, you have run the second kind. The scanner catches words and punctuation. The editor is there for what a script cannot see, such as flat rhythm, a claim with no evidence behind it, or a contrast dressed up to sound like insight.

On 4 September I drafted a newsletter edition from a journal entry about an advisory call. The scanner scored it 1 out of 100, which is a clean result. The editor returned PASS on both versions, the LinkedIn one and the X one. I held it anyway, because I found two problems myself when I went looking through my own files.

The first was freshness. The same journal entry had already become a LinkedIn post on 31 August, with the same audience, the same argument and the same example. The newsletter would have put that story in front of a heavily overlapping readership four days later, at greater length.

The second was worse. On 31 August an earlier review had ruled that a particular set of details from the story, taken together, pointed to the industry of the company involved, and every one of them had been stripped from that post. My 4 September draft put the whole set back and added one more.

Both checks were right about the text they were shown. The editor had never been told about the 31 August post or the ruling that came with it. I had handed it a draft, a complete rulebook, and no record of what had already run or already been ruled on.

Three weeks later the same gap produced the opposite error. On 25 September the editor reviewing an X article raised three hard fails. It said a $2.5 billion figure was the size of a startup's funding round when it was the valuation, and it flagged two details as invented. All three findings were wrong: the round was $250 million at a $2.5 billion valuation, and both details were in the news coverage I was working from. My prompt to the editor had left those sources out. The next day, on a rewrite, it failed a line saying an assistant summarises my email. That use case was in the first note I wrote for the piece, and the prompt had passed along only the second.

There is a third failure, quieter than the other two. On 14 September the editor caught three real problems in a blog draft, including a sentence that credited a report with a claim the report does not make. Its suggested replacement for that sentence was itself a binary contrast of the kind my rules ban. On 24 September it asked me to add a concrete example from a team to support one paragraph. The only way to supply that example was to invent an observation and write it in my own first person, so I cut the paragraph down instead.

None of this makes the editor useless. On 22 September it caught a date list in a draft that read 14, 19, 20 and 21 September when the log said 13, 14, 20 and 21, and 19 September was a Saturday with no scheduled run at all. The scanner passed that draft at every stage, wrong date included, because a scanner reads text and has no way to check a claim against a file.

The editor also catches me when my fixes are cosmetic. On 24 September the scanner passed a newsletter draft that contained three copies of a construction my rules ban, the kind that sets up a claim only to knock it down. The editor failed it. I rewrote the closing pair, swapped the vocabulary, added real numbers and kept the matched two-sentence shape. The editor failed it again and called the fix camouflage. It was right. I deleted the second sentence and replaced it with a plain statement of what I would do next. Two days earlier, on the blog draft, it had caught the same banned pattern surviving a paraphrase. A scanner cannot see that kind of failure, because the words change and the shape stays.

The briefing packet

Every one of those failures traces back to what went into the brief. I built the fix out of them one incident at a time, and it has four parts. I would ask any team using a model as a checker to use the same brief.

   DRAFT
     |
     +-- 1. SOURCES ---- every fact the draft leans on
     +-- 2. HISTORY ---- what already ran, what was ruled
     +-- 3. RULES ------ the house standard, in writing
     |
   VERDICT
     |
     +-- 4. ADJUDICATION -- each finding checked against the files
     |
   SHIP / FIX / HOLD

1. Sources: every fact the draft relies on. The editor cannot tell a sourced detail from an invented one unless the source is in front of it. Paste the article, the transcript, the research notes or the original request, in full. When I left out the news coverage, the editor did exactly what I had asked it to do, which was to challenge anything it could not trace, and it could not trace three true facts. If the source is a conversation, pass every message, because the detail you need is often in the first one.

2. History: what has already run and what was already decided. This is the part I skipped on 4 September, because it does not feel like part of the draft. A reviewer with a fresh context has no memory of last week's post, last month's ruling or the reason a paragraph was cut. My 4 September draft was clean against every rule and wrong against its own history. The brief should list what went out on the same topic in the last 14 days and quote any prior ruling that applies, word for word.

3. Rules: the standard, written down. This is where I started, and on its own it was not enough. Put the rules in a document rather than in your head, and point the editor at it by path. Mine is a written directive covering banned phrasing, sentence rhythm and the requirement that every claim carry a number, a name or a source. The rules are the smallest of the three inputs by volume, and they do the least work when the other two are missing.

4. Adjudication: every finding checked before it is applied. A verdict is a list of claims about the draft, and each one can be wrong. Before I change a line, I check the finding against the source files. That is how the three false fails on 25 September were caught before I applied them. Applying them would have meant deleting true facts from the draft before it went out. The same step catches suggested rewrites that reintroduce the pattern they were meant to remove, and requests for evidence that does not exist.

The order matters. Sources and history go in before the review, rules sit alongside it, and adjudication happens after it. Adjudication is only quick when the reviewer saw the sources.

Why a model checking a model makes this harder

An AI editor gives no sign of what it was not shown. My 4 September PASS was blind to history and my 25 September FAIL was blind to sources, and neither verdict mentioned it.

A second reason is that a fresh context is the whole appeal. I run the editor in a clean session so that it does not share my assumptions about the draft. That same property means it does not share my files, my memory of last week or the reason a paragraph was cut. The brief is the only way any of that reaches it.

The third reason is volume. My editor is called on every blog post, newsletter edition, thread and batch of X posts, and no call remembers the one before it. The history has to be written into each brief. When I skipped that on 4 September, the cost was a repeated story and a set of details that an earlier review had already taken out.

Three things to do this week

  1. Add a sources block to your review prompt. Take the prompt you use to have a model check your writing, and add a section that pastes in every source the draft relies on, in full. Run it on one piece this week and count how many findings disappear.
  2. Write a 14-day history note. Before your next review, list everything you published on the same topic in the last two weeks, with dates, and any rule you set about what may or may not be said. Paste it into the brief and see whether any verdict changes.
  3. Check three findings before you apply them. On your next review, pick three findings and trace each one back to a file before touching the draft. Mark each as right, wrong or right-but-bad-fix. Keep the third category separate, because a finding can be right while its suggested fix is wrong.

What to read next

While you are here, the back catalogue has more on this:

Footer

Anees Merchant writes one essay every Tuesday about enterprise AI, agentic systems, and the human side of the work. He is the author of Merchants of AI, a TEDx speaker, and a doctoral researcher in Human-AI Communication at the Swiss School of Business and Management.

The newsletter version, with extra commentary, goes out separately on Mondays. Subscribe at /newsletter.

See you Tuesday.

FAQ

Common Questions

What is an AI reviewer?

An AI reviewer is a separate model that reads a draft against written rules before it is published and returns a verdict with specific fixes. It usually runs with a fresh context so it does not share the writer's assumptions. It covers what a word-level scanner cannot, such as unsupported claims, flat sentence rhythm and misattributed facts. Its verdict depends on what it is given: the draft, the sources, the history and the rules.

How does an AI reviewer differ from an AI-writing scanner?

A scanner is deterministic. It matches known patterns in the text, such as punctuation, phrasing and formatting, and returns a score and an exit code. A reviewer is a model that makes judgment calls about evidence, rhythm and meaning. The scanner cannot check a claim against a source file, and the reviewer can only check claims against sources it has been shown. Running both, scanner first, covers more ground than either alone.

Why does an AI reviewer pass drafts it should fail?

An AI reviewer passes a flawed draft when the flaw lives outside the text it was shown. If a story already ran last week, or an earlier review ruled that certain details must stay out, a reviewer with a fresh context has no way to know. The draft can be clean against every written rule and still be wrong against its own history. The fix is to include recent published work and prior rulings in the brief.

When should you override an AI reviewer's finding?

Override a finding when you have checked it against the source files and it is wrong. Common cases are facts the reviewer called invented because the source was not in its prompt, suggested rewrites that reintroduce the pattern being fixed, and requests for supporting examples that would require inventing evidence. Record the reason when you override, so the next review can be briefed with it.

What is the first step to briefing an AI reviewer properly?

The first step is to paste every source the draft relies on into the review prompt, in full. That includes articles, transcripts, research notes and the original request, with every message rather than only the latest one. Without sources, a careful reviewer will flag true facts as unsupported, and a writer who applies those findings without checking them will delete accurate material from the draft.