An enterprise agent is ready for production when it survives a test set, not when it wins a demo, because a demo proves the agent can succeed and a test set proves it fails safely. The demo is the sales pitch. The test set is the job. Almost every agent I have watched stall inside a large organization stalled in the gap between those two, and the handful that reached production shared a shape I want to walk through.
Here is the part nobody says in the room. A demo is designed to work. The person running it picks the inputs, avoids the ugly edge cases, and reruns the flow until it looks clean. That is not dishonest. It is what a demo is for. The problem starts when a leadership team watches a flawless demo and treats it as evidence that the agent is ready for real work. It is not evidence of that. It is evidence that the agent can handle the exact cases someone chose to show you. The cases nobody chose to show you are the ones that decide whether the thing lives or dies once it touches production.
The trap catches sophisticated teams too. I have watched groups who would never accept a vendor's benchmark at face value walk straight into accepting their own demo at face value. The reason is human, not technical. A working demo feels like proof because you watched it happen with your own eyes. But you watched a rehearsal. The gap between a rehearsal and opening night is exactly the gap between the twelve cases in the demo and the twelve thousand cases waiting in the queue. Nobody rehearses the cases that make them look bad, which is why those cases are missing from the very moment you use to decide.
I have spent the last two years putting agentic systems into large organizations across banking, retail, and telecom. The model is almost never the reason a project fails. Reliability is. And the reason reliability keeps surprising teams is that they measured the wrong thing before launch. They measured whether the agent could win a scripted moment. They never measured how it behaves on the messy, half-finished cases that make up a real backlog. A model that scores well on a public benchmark can still act wrong on your particular refund policy, your particular customer, your particular exception that only your team knows about. The only way to find that out is to feed it your own history and watch.
The demo that dazzled and the backlog that told the truth
Last quarter I sat in a review with an operations team at a large enterprise. They had built an agent to triage a high-volume support queue, and the demo was genuinely impressive. Clean intake, correct routing, a tidy summary at the end. The room was ready to sign off and push it live to the whole queue that week.
I asked one question. Can we run it against fifty cases you have already closed, pulled at random from last month, including the ones your own people found hard?
We did. The result was not the demo. The agent handled thirty-eight of the fifty end to end. It flagged eight as uncertain and escalated them to a human, which is exactly what you want it to do when it is not sure. And it got four wrong while believing it was right. Those four are the whole story. Two were refund cases where the agent applied the wrong policy with total confidence. If those had gone out in production, the first anyone would have heard about it was an angry customer or a compliance note, not a dashboard.
The team's first instinct was disappointment. Seventy-six percent handled sounded worse than the demo had promised. My read was the opposite. For the first time, we knew where the agent broke, why it broke, and how badly. Those four wrong cases became the specification for the next two weeks of work. We added a policy check the agent had to pass before it could act on a refund, and we widened the escalation rule so anything touching money it was not certain about went to a person. On the next run, the misses dropped to zero and the handled rate climbed as trust let us give it more room.
That agent shipped. The demo did not ship it. The fifty cases did. And the two cases it originally got wrong closed the deal with the business owner faster than any feature list, because for the first time she could see exactly what she was signing up for.
Why the test set closes the deal, not just the doubt
There is a counterintuitive thing here worth sitting with. Showing a buyer where your agent fails makes them trust it more, not less. I have watched this happen across every serious deployment. A feature list invites suspicion, because the buyer knows a feature list is written to sell. A table that says "handled forty-two, flagged six, missed two, and here are the two and the fix" invites the opposite reaction. It reads as someone telling the truth about their own product, and truth is the rarest thing in an enterprise sales cycle. The doubt that was slowing the deal comes from not knowing the failure modes. Name the failure modes and the doubt has nowhere left to hide.
This matters because the buyer for an agent is not the buyer for software. Software was sold to a technology team out of a software budget, and the question was whether it worked. An agent is sold to whoever owns the function, out of a labor budget, and their question is different. They are not asking whether it works in general. They are asking whether they can put their name on the outcomes it produces while they are not watching. No demo answers that question. A backlog test answers it directly, in their own cases, in numbers they recognise from their own operation. That is why the test set does more selling than any slide.
The same discipline is what lets you go from one agent to many. Teams that skip the test set and launch on a demo end up with a single fragile agent that everyone is afraid to touch. Teams that build a test set end up with a repeatable pattern. Once you have scored one workflow into three buckets and driven the misses to zero, you know the exact shape of the work: pull the cases, score the buckets, set the ceiling, raise the automation. You run that same loop on the next workflow, and the next. A modest agent that runs reliably every day and can be measured is worth more than a brilliant one that breaks on Tuesday and cannot be trusted, because only the measurable one can be copied.
The framework: read the three buckets, not the one number
When you run an agent against real cases instead of a demo, every single case lands in one of three buckets. The discipline is knowing which bucket matters.
- Handled. The agent did the work end to end and got it right. This is your automation rate, and it is the number everyone quotes. On its own it means very little, because a high handled rate hides how the agent behaves when it is unsure.
- Flagged. The agent knew it was not confident and escalated to a human. This is your safety valve. A high flag rate is not a weakness. An agent that says "I am not sure, look at this one" is doing the single most valuable thing an enterprise agent can do. Flagging is a feature, not a failure.
- Missed. The agent acted wrong while believing it was right. This is the only bucket that can hurt you, because a miss is silent. Nobody escalates it. It goes out into the world looking like success and comes back as a customer complaint, a wrong payment, or a regulator's question.
50 real backlog cases
|
+--------------+--------------+
| | |
Handled Flagged Missed
(38) (8) (4)
| | |
Automation Safety Danger
rate valve zone
| | |
quote this welcome this fix this firstHere is the shift most teams never make. They optimise the Handled number and ignore the Missed number, when the Missed number is the one that decides whether the agent is allowed near production. An agent that handles forty-two, flags six, and misses two is more shippable than an agent that handles forty-eight, flags zero, and misses two. The second one looks better on a slide. It is far more dangerous, because it never tells you when it is out of its depth. It is confident right up until it is confidently wrong, and you cannot see the wrong ones coming.
So the order of operations flips. You do not set an automation target and hope the errors stay low. You set a miss ceiling first, the number of silent wrong actions you can tolerate in production, and often that number is zero for anything involving money, safety, or a legal commitment. Then you push the automation rate as high as it will go underneath that ceiling. Reliability is not a percentage you brag about. It is a ceiling you refuse to cross.
This is also why the first agent in any organization should own one narrow workflow, not a role. A scoped workflow has a knowable set of cases, which means you can actually build a test set for it. "Run sales" has no test set. "Re-engage customers who went quiet ninety days ago" has fifty real examples you can pull this afternoon. Scope is what makes the test set possible, and the test set is what makes production possible.
Three things to do this week
You do not need new tooling or a bigger budget to run this. You need a backlog and some honesty.
- Build a test set from your own closed cases. Pull fifty real examples of the workflow you want an agent to own, from your actual history, including the ones your team found hard. Do not clean them up. The mess is the point.
- Score every case into three buckets, not one accuracy number. Mark each result Handled, Flagged, or Missed. A single accuracy percentage hides the difference between an agent that escalates when unsure and one that guesses. That difference is everything.
- Set a miss ceiling before you set an automation target. Decide how many silent wrong actions you can live with in production. For anything touching money or compliance, start at zero. Then raise automation only as far as that ceiling allows.
What to Read Next
While you are here, the back catalogue has more on this: