AI Leadership for Executives

The Review Ceiling, and What Executives Should Staff Before Buying More AI

The Review Ceiling, and What Executives Should Staff Before Buying More AI

A review ceiling is the point where an organisation produces more AI output than any qualified person can judge, because generation capacity scaled with spend and judgment capacity did not. Most executives have already crossed that line. Very few have named it, which is why their AI programme reports more work and delivers the same result.

The ceiling is not a technology problem and no vendor will sell you a fix for it. It is a staffing decision that sits with the person who signs the budget, and it is usually made by accident.

The quarter the drafts tripled and the calendar did not move

Earlier this quarter I sat with the leadership team of a large consumer business that had run a clean AI pilot. Nobody had oversold it. They took their quarterly performance review, a piece of work that had reliably eaten two days of an analyst's week, and rebuilt it around agents. Pull the numbers, clean them, build the charts, write the first commentary. A usable version now lands in about twenty minutes.

So I asked the obvious question. What happened to the two days?

Nothing happened to the two days. The analyst now produces four versions of that review instead of one, because she can, and all four go to a director who has exactly the eight hours he had last year. The reviews queue in his inbox. That queue is invisible, because an inbox is not a project plan and nobody reports on it, so the programme dashboard shows a ninety-five percent cut in production time and is not lying. It is measuring the half of the system that got faster.

That business is not unusual. KPMG surveyed 2,145 senior leaders this year. Seventy-six percent said AI is delivering value. Among the leaders who had full visibility into what AI was costing them, fifteen percent could show a measured return. Among those without that visibility, three percent could. Read the two halves together. Output is real and rising. Almost nobody can prove what the output was worth, and the instrument they are missing sits on the human side of the system rather than the model side.

I see the same shape in my own work, at a much smaller scale and with nobody to blame. I run more than twenty automations across content, research, briefings, and scheduling. Last month I counted two things for one week: how many finished pieces those systems produced, and how many I read closely enough to genuinely approve or reject. The production number was flattering. The read number was eleven. Everything past eleven either went out lightly checked or sat in a folder I promised myself I would get to. That folder is my review ceiling. I built it myself, on a budget of zero, with a great deal of enthusiasm.

The hiring market has worked this out ahead of most boards. Look at what companies are actually paying up for right now. Forward deployed engineers. People who can translate a business problem into a specification. Change and adoption leads. Every one of those is a judgment role. When your competitor rents the same reasoning you do, at the same price, from the same three providers, the model stops being the difference. The difference becomes the people who can decide what should be built and tell whether what came back is any good.

That capacity does not arrive with headcount, which is the part most hiring plans miss. I have spent the last few weeks in one-on-one conversations with early career professionals, and the pattern repeated often enough that I stopped treating it as bad luck. Ask what changed in AI over the past fortnight and the answer arrives fast and confident. Ask one follow-up question and it falls apart. They have the headline and none of the substance underneath it. That is not a skills gap in the usual sense, because the material is free and sitting in public. It is an attention gap.

This matters for the ceiling because a reviewer who only has headlines cannot review. They can approve, which is a different act performed at the same speed and with none of the protection. Add ten people like that to a team and your review ratio looks better on paper while the actual judgment in the system stays exactly where it was. Reviewers are made deliberately or not at all.

The review ceiling, and the three capacities underneath it

The ceiling is simple to compute and almost nobody computes it.

Count the units of AI generated work your team produces in a week. A unit is anything a human is supposed to sign off on before it counts: a draft, an analysis, a recommendation, a code change, a customer response. Then count the units a qualified person can judge properly in that same week. Not skim. Judge. Divide the first number by the second. That is your review ratio.

Below one, you have slack and the system is healthy. At one, you are sitting on the ceiling. Above one, you are accumulating work that nobody has judged, and that pile is a liability rather than an asset. Unreviewed output ships with defects you pay for later, or it never ships at all and you paid for it twice.

Raising the ceiling means funding three capacities. They are separate, they fail separately, and most AI budgets fund none of them.

1. Specification capacity. How many people in the room can write down what finished looks like before the work starts. Not a goal, not a brief, an acceptance test. If nobody can specify, then review collapses into taste, and taste does not scale past one reviewer with a strong opinion. This is the cheapest of the three capacities to build and the one skipped most often, because it looks like paperwork and feels like a delay. It is the reason a director spends forty minutes on a draft that a written standard would have rejected in four.

2. Review capacity. How many units a qualified person can judge in a week, multiplied by how many such people you have. Almost every organisation I ask has one or two real reviewers per domain and has never counted them. Two things raise this number. Tier the work, so routine output gets a defined sample check rather than a full read, and reserve deep review for what carries real consequence. Then widen the bench, which means training reviewers deliberately instead of waiting for people to become senior enough by accident.

3. Adoption capacity. How much changed work the organisation can absorb. Output that lands in a workflow nobody has updated is not value. It is inventory. This capacity is where the change and adoption roles earn their pay, and it is the one executives are most tempted to treat as a communications exercise rather than a staffing line. A team can review perfectly and still get nothing out of the programme, because the approved work arrives into a process, a system of record, and a set of habits that were all designed around the old volume.

The fix people reach for first is to point AI at the review step and call the problem solved. Automated review has a real place, and I use it. It works on anything with a written standard: format, completeness, whether the numbers reconcile, whether the claim appears in the source. It does not work on the judgment that made the standard, and it is worst exactly where the stakes are highest, on the unfamiliar case where the right answer is not in any rule you have written down yet. So automate the mechanical layer of review by all means. It buys back real hours. Just do not book those hours as a solved problem, because what you removed was checking and what you still need is deciding.

Notice also that the three capacities have to be built in order. Review without a specification is opinion. Adoption without review is risk moved downstream at speed. Teams that skip to the third and run a large enablement programme on top of unchecked output usually spend the following quarter rebuilding trust in the tool, which is a much more expensive project than writing the standard would have been.

   Generation capacity        scales with spend
   ------------------------------------------------
   Specification  ->  Review  ->  Adoption
        |               |            |
        +-------+-------+------+-----+
                |              |
      review ratio below 1   review ratio above 1
        work ships           work piles up
        and gets used        and gets counted

Put those three next to a purchase order and the budget conversation changes. The next line item on most AI programmes is another seat licence or a bigger context window, and both of those buy more generation. If your review ratio is already above one, more generation makes the number worse. The money belongs on the right hand side of that diagram.

There is a board-level version of this too. Ask for the review ratio as a standing programme metric, alongside adoption and cost. It is a number your team can produce in an afternoon, and it is far more honest about the state of an AI programme than a time-saved percentage, because time saved on one task tells you nothing about whether the organisation got faster.

Raising review capacity is not a hedge against AI. It is the only way to safely use more of it. Every organisation that has successfully pushed automation deep into its operations did it by getting very good at judging output quickly, not by trusting output blindly. The ceiling is what stands between a pilot that impressed a boardroom and a system that quietly runs a business.

Three things to do this week

  1. Compute your review ratio for one team. Pick a single team, count the AI generated units they produced last week, and count what a qualified reviewer actually judged. You will have a number by Thursday. In my experience the first count comes in somewhere between two and five, and the shock of seeing it is what gets the rest of this work moving.
  2. Write one acceptance standard. Take your highest volume AI assisted output and write down, on one page, what makes it acceptable. Circulate it to everyone who produces and everyone who reviews. This single page will cut review time on that output more than any tooling change you make this quarter.
  3. Name your reviewers out loud. For each domain where AI output flows, name the people qualified to approve it and write down how many hours a week they have for that. If the name is one person and the hours are four, you have found your ceiling, and you now have a hiring or training decision instead of a mystery.

What to read next

While you are here, the back catalogue has more on running AI programmes that survive contact with an organisation:

FAQ

Common Questions

What is a review ceiling in an AI programme?

A review ceiling is the point where an organisation produces more AI generated output than any qualified person can properly judge. Generation capacity rises with spend, because more licences and better models produce more work. Judgment capacity does not move unless someone deliberately staffs it. Once output passes what your reviewers can assess, the extra work stops being an asset. It either ships without a real check, and you pay for the defects later, or it queues in an inbox nobody reports on. The ceiling is a staffing condition, not a technology limit.

How do you calculate a review ratio?

Count the units of AI generated work a team produces in one week, where a unit is any single item a human is meant to sign off on before it counts. Then count how many of those units a qualified person can judge properly in the same week, not skim. Divide the first number by the second. Below one you have slack. At one you are on the ceiling. Above one you are accumulating unjudged work. Most teams computing this for the first time land somewhere between two and five.

How does review capacity differ from adoption capacity?

Review capacity is about judgment, meaning how many units of AI output a qualified person can assess per week and how many such people you have. Adoption capacity is about absorption, meaning how much changed work the organisation can actually take into its daily operations. They fail differently. Weak review capacity produces a backlog of unchecked output and hidden quality risk. Weak adoption capacity produces approved work that nobody uses. A programme can pass one and fail the other, which is why they need separate owners and separate budget lines.

Why do AI pilots report large time savings without changing business results?

Because the dashboard measures the half of the system that got faster. Production time on a single task falls sharply and gets reported. The review, approval, and adoption steps downstream keep the same capacity they always had, so the work queues where nobody is looking, usually in a manager's inbox. The time saved is real and the business result is unchanged, which is how a programme can be honest and misleading at the same time. Tracking a review ratio alongside time saved exposes the gap in one number.

What is the first step to raise an organisation's review capacity?

Write one acceptance standard for your highest volume AI assisted output. One page describing what makes that output acceptable, shared with everyone who produces it and everyone who reviews it. This works before hiring because most review time is spent rediscovering the standard rather than applying it. Once the standard is written, routine output can be sample checked against it and deep review can be reserved for what carries real consequence. After that, name your qualified reviewers and the hours they genuinely have, then hire or train against that gap.