Enterprise Agentic AI

The 50-Case Test That Decides If Your Enterprise Agent Ships

The 50-Case Test That Decides If Your Enterprise Agent Ships

An enterprise agent is ready for production when it survives a test set, not when it wins a demo, because a demo proves the agent can succeed and a test set proves it fails safely. The demo is the sales pitch. The test set is the job. Almost every agent I have watched stall inside a large organization stalled in the gap between those two, and the handful that reached production shared a shape I want to walk through.

Here is the part nobody says in the room. A demo is designed to work. The person running it picks the inputs, avoids the ugly edge cases, and reruns the flow until it looks clean. That is not dishonest. It is what a demo is for. The problem starts when a leadership team watches a flawless demo and treats it as evidence that the agent is ready for real work. It is not evidence of that. It is evidence that the agent can handle the exact cases someone chose to show you. The cases nobody chose to show you are the ones that decide whether the thing lives or dies once it touches production.

The trap catches sophisticated teams too. I have watched groups who would never accept a vendor's benchmark at face value walk straight into accepting their own demo at face value. The reason is human, not technical. A working demo feels like proof because you watched it happen with your own eyes. But you watched a rehearsal. The gap between a rehearsal and opening night is exactly the gap between the twelve cases in the demo and the twelve thousand cases waiting in the queue. Nobody rehearses the cases that make them look bad, which is why those cases are missing from the very moment you use to decide.

I have spent the last two years putting agentic systems into large organizations across banking, retail, and telecom. The model is almost never the reason a project fails. Reliability is. And the reason reliability keeps surprising teams is that they measured the wrong thing before launch. They measured whether the agent could win a scripted moment. They never measured how it behaves on the messy, half-finished cases that make up a real backlog. A model that scores well on a public benchmark can still act wrong on your particular refund policy, your particular customer, your particular exception that only your team knows about. The only way to find that out is to feed it your own history and watch.

The demo that dazzled and the backlog that told the truth

Last quarter I sat in a review with an operations team at a large enterprise. They had built an agent to triage a high-volume support queue, and the demo was genuinely impressive. Clean intake, correct routing, a tidy summary at the end. The room was ready to sign off and push it live to the whole queue that week.

I asked one question. Can we run it against fifty cases you have already closed, pulled at random from last month, including the ones your own people found hard?

We did. The result was not the demo. The agent handled thirty-eight of the fifty end to end. It flagged eight as uncertain and escalated them to a human, which is exactly what you want it to do when it is not sure. And it got four wrong while believing it was right. Those four are the whole story. Two were refund cases where the agent applied the wrong policy with total confidence. If those had gone out in production, the first anyone would have heard about it was an angry customer or a compliance note, not a dashboard.

The team's first instinct was disappointment. Seventy-six percent handled sounded worse than the demo had promised. My read was the opposite. For the first time, we knew where the agent broke, why it broke, and how badly. Those four wrong cases became the specification for the next two weeks of work. We added a policy check the agent had to pass before it could act on a refund, and we widened the escalation rule so anything touching money it was not certain about went to a person. On the next run, the misses dropped to zero and the handled rate climbed as trust let us give it more room.

That agent shipped. The demo did not ship it. The fifty cases did. And the two cases it originally got wrong closed the deal with the business owner faster than any feature list, because for the first time she could see exactly what she was signing up for.

Why the test set closes the deal, not just the doubt

There is a counterintuitive thing here worth sitting with. Showing a buyer where your agent fails makes them trust it more, not less. I have watched this happen across every serious deployment. A feature list invites suspicion, because the buyer knows a feature list is written to sell. A table that says "handled forty-two, flagged six, missed two, and here are the two and the fix" invites the opposite reaction. It reads as someone telling the truth about their own product, and truth is the rarest thing in an enterprise sales cycle. The doubt that was slowing the deal comes from not knowing the failure modes. Name the failure modes and the doubt has nowhere left to hide.

This matters because the buyer for an agent is not the buyer for software. Software was sold to a technology team out of a software budget, and the question was whether it worked. An agent is sold to whoever owns the function, out of a labor budget, and their question is different. They are not asking whether it works in general. They are asking whether they can put their name on the outcomes it produces while they are not watching. No demo answers that question. A backlog test answers it directly, in their own cases, in numbers they recognise from their own operation. That is why the test set does more selling than any slide.

The same discipline is what lets you go from one agent to many. Teams that skip the test set and launch on a demo end up with a single fragile agent that everyone is afraid to touch. Teams that build a test set end up with a repeatable pattern. Once you have scored one workflow into three buckets and driven the misses to zero, you know the exact shape of the work: pull the cases, score the buckets, set the ceiling, raise the automation. You run that same loop on the next workflow, and the next. A modest agent that runs reliably every day and can be measured is worth more than a brilliant one that breaks on Tuesday and cannot be trusted, because only the measurable one can be copied.

The framework: read the three buckets, not the one number

When you run an agent against real cases instead of a demo, every single case lands in one of three buckets. The discipline is knowing which bucket matters.

  1. Handled. The agent did the work end to end and got it right. This is your automation rate, and it is the number everyone quotes. On its own it means very little, because a high handled rate hides how the agent behaves when it is unsure.
  2. Flagged. The agent knew it was not confident and escalated to a human. This is your safety valve. A high flag rate is not a weakness. An agent that says "I am not sure, look at this one" is doing the single most valuable thing an enterprise agent can do. Flagging is a feature, not a failure.
  3. Missed. The agent acted wrong while believing it was right. This is the only bucket that can hurt you, because a miss is silent. Nobody escalates it. It goes out into the world looking like success and comes back as a customer complaint, a wrong payment, or a regulator's question.
              50 real backlog cases
                       |
        +--------------+--------------+
        |              |              |
     Handled        Flagged        Missed
      (38)            (8)            (4)
        |              |              |
   Automation       Safety         Danger
      rate           valve          zone
        |              |              |
   quote this     welcome this   fix this first

Here is the shift most teams never make. They optimise the Handled number and ignore the Missed number, when the Missed number is the one that decides whether the agent is allowed near production. An agent that handles forty-two, flags six, and misses two is more shippable than an agent that handles forty-eight, flags zero, and misses two. The second one looks better on a slide. It is far more dangerous, because it never tells you when it is out of its depth. It is confident right up until it is confidently wrong, and you cannot see the wrong ones coming.

So the order of operations flips. You do not set an automation target and hope the errors stay low. You set a miss ceiling first, the number of silent wrong actions you can tolerate in production, and often that number is zero for anything involving money, safety, or a legal commitment. Then you push the automation rate as high as it will go underneath that ceiling. Reliability is not a percentage you brag about. It is a ceiling you refuse to cross.

This is also why the first agent in any organization should own one narrow workflow, not a role. A scoped workflow has a knowable set of cases, which means you can actually build a test set for it. "Run sales" has no test set. "Re-engage customers who went quiet ninety days ago" has fifty real examples you can pull this afternoon. Scope is what makes the test set possible, and the test set is what makes production possible.

Three things to do this week

You do not need new tooling or a bigger budget to run this. You need a backlog and some honesty.

  1. Build a test set from your own closed cases. Pull fifty real examples of the workflow you want an agent to own, from your actual history, including the ones your team found hard. Do not clean them up. The mess is the point.
  2. Score every case into three buckets, not one accuracy number. Mark each result Handled, Flagged, or Missed. A single accuracy percentage hides the difference between an agent that escalates when unsure and one that guesses. That difference is everything.
  3. Set a miss ceiling before you set an automation target. Decide how many silent wrong actions you can live with in production. For anything touching money or compliance, start at zero. Then raise automation only as far as that ceiling allows.

What to Read Next

While you are here, the back catalogue has more on this:

FAQ

Common Questions

What is the backlog test for enterprise AI agents?

The backlog test is a way to decide whether an enterprise AI agent is ready for production by running it against fifty real, previously closed cases from your own history rather than a scripted demo. Each case lands in one of three buckets: Handled, where the agent worked correctly end to end; Flagged, where it escalated because it was unsure; and Missed, where it acted wrong while believing it was right. The test shows where the agent breaks and how badly, which a demo is designed to hide. Agents that pass a backlog test reach production. Agents judged only on a demo tend to stall.

How does a test set differ from a demo when evaluating an AI agent?

A demo is chosen to succeed. The person running it picks clean inputs and reruns the flow until it looks polished, so it proves only that the agent can handle the cases someone selected. A test set is chosen to expose failure. It uses real cases from your backlog, including the hard ones nobody would put in a demo, and it measures how the agent behaves when the input is messy. A demo answers whether the agent can win a moment. A test set answers whether it fails safely, which is the question production actually cares about.

Why do enterprise agents that look good in demos fail in production?

They fail because a demo measures scripted success and production delivers unscripted cases. A leadership team watches a clean demo and treats it as proof the agent is ready, but the demo only covered inputs someone chose in advance. Production sends the edge cases and the half-finished requests that never appear in a demo. Without a test set built from real backlog cases before launch, the team has no idea where the agent silently acts wrong. The gap between the demo and the workflow is where most pilots quietly die.

When should a team run a backlog test on an AI agent?

Run the backlog test before launch, not after. It is the last checkpoint between a promising demo and a production commitment. The right moment is once you have scoped the agent to a single workflow, because a narrow workflow has a knowable set of cases you can actually pull and score. Run it again after every change to the agent's behaviour, so you can see whether a fix that raised the handled rate also raised the miss rate. Teams that run it only after something breaks in production are reading the test too late.

What is the first step to building a test set for an AI agent?

The first step is to pull fifty real, already-closed cases of the exact workflow you want the agent to own, taken from your own history and including the ones your team found hard. Do not clean them or simplify them, because the mess is what production will send. Then score each case as Handled, Flagged, or Missed rather than as a single accuracy number. That scoring reveals the difference between an agent that escalates when unsure and one that guesses confidently, which is the difference that decides whether it is safe to ship.