Enterprise Agentic AI

Every Agent Demo Sells the Ceiling. Adoption Is Decided by the Floor.

Every Agent Demo Sells the Ceiling. Adoption Is Decided by the Floor.

The adoption floor is the full cost of getting one new person from access to first useful output, and it decides whether an agent spreads inside a company or dies quietly, because nobody abandons a tool that has already worked for them once. Every demo sells the ceiling, which is the best thing a system can do on its best day. Almost nobody measures the floor. Four numbers make it visible.

The ceiling is what vendors compete on and what evaluation committees score. The floor is what decides whether the second person in your building ever tries the platform at all. Those are different questions, and most agent programmes only ask one of them.

Look at how a normal platform evaluation is run and the reason becomes obvious. A vendor sends two engineers who have lived inside the product for a year. They arrive with an environment already configured, credentials already issued, and a use case chosen because it is the one the system handles best. Nothing in that hour resembles what a new employee will experience on a Tuesday with a half hour gap between meetings. The score you produce is a measurement of the vendor's best people on their best day, and then you hand it to a procurement committee as though it were a prediction about your own staff. It is a fair test of the ceiling. It says nothing at all about the floor.

Two platforms I gave up on, and one I did not

Over the past few months I ran the same experiment three times, on three different agent platforms. Same intent each time. Build something small and personal, find out what the system can really do, keep going if it earns the hours.

Two of them I abandoned. Not because they were weak. Both had real capability, and I could see it sitting there behind the setup. The problem was everything in front of the capability. Configuration files that assumed knowledge I did not have. Dependencies that failed with error messages written for the person who built the tool. A first run that technically produced something, and produced nothing I would call useful. Both took the better part of a weekend of evenings before anything worked, and by then the question had quietly changed from what can I build with this into whether I still wanted to.

The third one I had running in under an hour.

What surprised me was not the speed. It was what the speed did to my behaviour. On the two hard platforms I arrived with a list of things to try, and I worked the list, because I had paid so much to get in that I was going to extract value from the plan I walked in with. On the easy one I abandoned my list inside a day. I kept following tangents instead. A thought would arrive, and rather than estimating whether the thought was worth the setup, I just tried it. Most of those tangents were dead ends. Three of them turned into the best things I built that month.

That is the part I have not seen written down anywhere. Setup cost does not only slow you down. It changes what you are willing to attempt. Above a certain friction level, every idea gets silently priced before it gets tested, and the ideas that lose that pricing exercise are the speculative ones, the ones where you cannot say in advance what you will find. Curiosity has a cost threshold. Below it you follow the tangent. Above it you decide the tangent is not worth the setup and you never learn what was there.

I have some standing to say this, because I built the counterexample the slow way. I am a finance graduate. I did not know how to use git when I started. Over roughly six months I put together a personal assistant system that now runs more than twenty automations across content, research, briefings and scheduling, and almost none of that came from a plan. It came from evenings where trying an idea was cheap enough that I did not stop to justify it first. The parts of that system I use daily were nearly all tangents. The parts I designed deliberately in advance are the ones gathering dust.

Let me be honest about the limits of that story. Three weeks is a honeymoon, not a verdict, which is why I am not naming any of the three platforms here. A first run experience is the easiest thing in software to get right and the least predictive of whether a system holds up at month six. What I am confident about is the mechanism, not the scoreboard. Two systems with genuine capability lost me on the floor, and their ceilings never got a vote.

Now put that at enterprise scale. I have watched pilots where the platform was fine and the programme still went nowhere, and the post mortem always reaches for the same three words. Training. Adoption. Change management. Look earlier than that. The people who were supposed to use the platform had to raise a ticket for access, wait two days, read a nine page setup document, and then produce something they could not tell was any good. Most of them tried once. A few tried twice. Almost none of them reached the point where the system did something they could feel. The programme did not fail at the ceiling. It failed in the first forty minutes, one user at a time, repeatedly, and nothing in the reporting captured it.

The adoption floor, in four numbers

Four numbers describe your floor. None of them appear on a standard vendor evaluation form. All four can be collected in a week, by one person, with a stopwatch and a notebook.

One. Time to first useful output. Measure from the moment a new person is given access to the moment they are holding something they would have wanted anyway. Not a hello world. Not a demo script somebody else wrote. Something they can actually use. Measure it on a real colleague, in real conditions, rather than accepting an estimate from the team that built the thing, because that team has the environment already configured and has forgotten what they know. My working threshold is one hour. Past an hour, the people who continue are the ones who were already committed, which means you are measuring enthusiasm rather than the system.

Two. Failure surface. Count the setup steps that can fail without telling the user what to do next. Every silent failure is a place where a reasonable person concludes that the tool is broken and that they are stupid, in that order, and quietly stops. A platform with eleven setup steps and honest error messages is friendlier than one with four steps and a single cryptic stack trace. Count the steps, then count the ones that fail badly, and treat the second number as the real one.

Three. Dead end cost. Ask what it costs somebody to try an idea and be wrong. If a failed experiment costs an afternoon, people run one experiment a month and choose it carefully. If it costs ten minutes, they run twenty and choose badly, which is exactly the point. Cheap dead ends are the mechanism by which teams find things nobody specified in advance. This number governs how speculative your users are permitted to be, and speculation is where the surprising wins come from.

Four. Distance to the second user. Can person two start from what person one built, or do they start at zero? This is the number that decides whether the floor gets paid once or gets paid every single time. Most agent deployments pay it every time, because the first user's work lives in their own environment and there is no route from their setup to anybody else's. A programme with a low floor for its first user and no route to the second is a hobby with a budget code attached.

                   CEILING        what the demo sells
                      ^
                      |           scored in the bake-off
   -------------------|---------------------------------
                      |           nobody scores this
                      v
   Time to first  ->  Failure   ->  Dead end  ->  Distance to
   useful output      surface       cost          second user
        |                |             |               |
   who bothers       who quits    how bold they    whether it
   to try            in week one  are after that   leaves one desk

Two of these are the vendor's to fix and two of them are yours, which is worth being clear about before anybody starts negotiating. Time to first useful output and failure surface belong mostly to the product, and you can put both into a contract as acceptance criteria rather than discovering them after signature. Dead end cost and distance to the second user are almost entirely yours. They are set by your access process, your environment, your permissions, and whether anyone owns the job of moving a working setup from one desk to another. Blaming a vendor for the second pair is common and it never fixes anything.

The four compound in that order. Time to first useful output decides who bothers to try. Failure surface decides who quits in the first week. Dead end cost decides how adventurous the survivors are willing to be. Distance to the second user decides whether any of it ever leaves the desk it started on. A programme can score well on any one of them and still go nowhere, which is why the floor is a sequence and not a score.

There is a second order effect worth naming. When the floor drops far enough, the work that separates a good practitioner from an average one stops being throughput and finish, because the system supplies both. What is left is judgment. Which tangent was worth following. Which of those twenty cheap experiments produced something real. I sat with a leadership team recently who showed me an organisation chart carrying three seniority levels of what was, on paper, the same role, and asked what separated them in terms of actual tasks and responsibilities. Nobody in the room could answer. A ladder built on volume and polish has nothing holding it up once the floor drops, and most job architectures are still written that way.

Three things to do this week

  1. Stopwatch one real onboarding. Take a colleague who has never touched your agent platform, give them access on a Monday morning, and time them until they produce something they would have wanted anyway. Do not help. Write the number down. Expect it to be four to ten times whatever your platform team told you.
  2. Write your failure surface on one page. List every setup step between access and first output, then mark the ones that can fail without saying what to do next. That marked list is your actual onboarding document, and it is usually shorter and more useful than the nine page version.
  3. Price a dead end out loud. Ask your team what it costs them to try something with the platform and be wrong. If the honest answer is half a day, you have your explanation for why nobody experiments, and you have a target worth more than the next licence upgrade.

What to read next

While you are here, the back catalogue has more on this:

FAQ

Common Questions

What is the adoption floor in an AI agent programme?

The adoption floor is the full cost of getting one new person from access granted to first useful output. It covers setup steps, configuration, permissions, waiting time, and the effort of producing something the person would have wanted anyway. The floor is separate from what the platform can do at its best, and it is the number that decides whether a tool spreads past the person who championed it. Nobody abandons a system that has already worked for them once, so the floor is really a measure of how many people ever reach that first moment.

How does the adoption floor differ from the capability ceiling?

The ceiling is the best thing a system can do on its best day, demonstrated by people who already know it well in an environment already configured for it. That is what vendors compete on and what evaluation committees score. The floor is what a new user pays before the system does anything for them at all. A platform can hold a high ceiling and a punishing floor at the same time, which produces the common outcome of a pilot that wins the bake-off and then never spreads. They are measured differently and they fail differently.

Why do capable agent platforms fail to spread inside a company?

Because the failure happens one user at a time, early, and nothing in standard reporting captures it. A prospective user raises a ticket, waits, reads a long setup document, hits a step that fails without saying what to do next, and quietly stops. The programme dashboard records licences issued rather than first useful outputs produced, so the drop-off is invisible. Post mortems then reach for training and change management, which are downstream explanations for a problem that occurred in the first forty minutes of contact with the tool.

When should a team measure time to first useful output?

Before signing a platform contract, and then again every quarter with a genuinely new user. Measure it on a real colleague under real conditions rather than accepting an estimate from the team that built or installed the system, because that team has the environment configured and has forgotten what they already know. A useful working threshold is one hour. Past an hour, the people who keep going tend to be the ones who were already committed, so the measurement stops describing the system and starts describing the enthusiasm of the volunteer.

What is the first step to lowering an agent platform's adoption floor?

Watch one real onboarding with a stopwatch and take notes, without helping. Give a colleague who has never used the platform access on a normal working morning and time them until they produce something they would have wanted anyway. The list of places they stalled is your fix list, ordered by cost, and it is almost always shorter than the setup document you already have. Fix the steps that fail without explaining what to do next first, because those are the ones that make reasonable people conclude the tool is broken.