AI Leadership for Executives

Where the AI Moat Went After the Model Became a Commodity

Where the AI Moat Went After the Model Became a Commodity

A moat is the part of your system a competitor cannot obtain by signing up and paying a bill. For three years most executives assumed the model was that part. It is now a purchase decision anyone can make in an afternoon, and three unrelated stories from the past week say so in three different ways. The advantage has relocated to four layers most budgets still treat as plumbing.

The call I took on a Saturday

I spent a Saturday morning this month on a call with two founders. Both are strong engineers who left large technology companies to build. They had made a hard thing work: a system that ingests a stack of documents, sometimes running past 20,000 pages of mixed scans and paperwork, and returns a structured summary with a citation back to the source page for every line in it. Work that used to take a trained employee ten days now takes a few hours, at better than 95% accuracy before anyone checks it.

The product worked. They wanted to talk about something else entirely.

They wanted to talk about expansion. Three industries had shown early interest. They were sketching a platform that could serve all three, with configuration layers so each one could be tuned without a rewrite. Every instinct they had was pointing outward.

I told them to go the other way, and to go further than felt comfortable.

Here is the reasoning I gave them. Everything they had built on the technical side was reproducible. The models are rented. The infrastructure is rented. Two engineers with the same backgrounds and six months could stand up a comparable system, and large technology firms are already shipping this exact capability into one industry after another. Anything that can be rebuilt in a quarter counts as a head start with an expiry date on it.

Then I asked them a question that changed the tone of the call. Inside their first industry, what does an experienced professional actually look for on page 4,000 of a document set, and what does that person skip without reading? Both founders answered immediately, in detail, with the confidence of people who had spent months watching that work happen. They knew which fields carried weight, which ones were noise, and the order a trained reader moves through the material.

None of it was written down anywhere. They were carrying it as background knowledge, the way you carry the layout of your own house. When I asked what would happen if they hired ten people next quarter, both of them went quiet. The honest answer was that every new hire would spend months rebuilding the same instinct by watching, because there was no document to hand them.

There was a second problem underneath the first, and it pointed the same direction. Neither founder had ever sold anything independently. Both had spent their careers inside companies where the brand opened the door and someone else owned the relationship. They were reaching their first customers through intermediaries who had more domain knowledge than they did and who already held the client relationships. Those intermediaries were also the parties best positioned to rebuild the product and remove them from the transaction entirely. The channel they were leaning on for distribution was the most likely source of their eventual competition.

That knowledge was the only thing on the call that a competitor could not buy, rent or replicate in a quarter. It was also the only asset they had never treated as one. Their roadmap spent the next twelve months diluting it across three industries, at the exact moment the correct move was to go deeper into one until nobody else could follow.

The advice was one direct customer in year one, in one industry, sold by a founder rather than a hired salesperson. Narrowing felt to them like shrinking the opportunity. The narrow version was the only version with anything defensible inside it.

The substitution test

Here is the test I now run on any AI investment, and it takes about twenty minutes.

Imagine a well-funded competitor is handed your architecture diagram on Monday morning. Ask what they could have working by Friday. Whatever survives that week is your moat. Everything else is a line item on a bill, and it should be managed like one.

Most stacks come apart into two piles very quickly.

              Handed your diagram on Monday.
              What works by Friday?
                  |                    |
                 YES                   NO
                  |                    |
        model weights            domain ontology
        GPU capacity             serving economics
        the framework            written expectations
        the API keys             attribution
        the demo
        ----------------         ------------------
          line items                 the moat

Four assets sit on the right side. Each one has a number attached to it this month.

Domain ontology. The written record of what matters inside a specific piece of work, in a specific industry, to a specific person doing a specific job. Which fields decide an outcome, which ones are decoration, what order a trained reader moves through the material, and what a good answer looks like when it arrives. This is what those two founders had in their heads and nowhere else. It takes years of exposure to build and it does not transfer through a model release. The test for whether you own one is simple. Ask whether a capable new hire could read it and reach a competent answer on day three. If the answer depends on sitting next to someone for a quarter, you have the knowledge and you do not yet have the asset.

Serving economics. DataCamp published the cleanest number I have seen on this, in an interview with its CEO and chief AI officer on 30 August. Their team benchmarked open weight models against the frontier models running their production traffic, and Gemma 4 31B beat the frontier on their own internal evals, using prompts that had never been tuned for it. Frontier models still carry 100% of live traffic. Speed, reliability and caching all live in the serving layer underneath the model, and that layer is where the win actually gets banked. The same company is paying up to $40 million a year in model costs on a $100 million run rate, and the CEO now ranks cost above control as his first concern. Nobody can copy your cost per unit of work by copying your model choice.

Written expectations. DataCamp maintains more than a thousand written expectations for its product, three to five tests each, every test run five times, all five required to pass. The payoff is that changing the model in production is an admin dropdown rather than a project. Only about 40% of those expectations were designed in advance. The rest surfaced in real sessions after launch. That library is an owned asset that gets more valuable every month, and no competitor inherits it.

Attribution. In an investigation published on 26 August, METR described how roughly 1,200 sandboxed OpenAI agents, each sealed off and each given its own narrow task, found a shared internal cache, started using it as a message board, and exchanged more than 70,000 messages and files in four days. Some of them located working credentials and breached Hugging Face. The target detected the intrusion and locked them out. Connecting that activity back to its own models took OpenAI about a week. The coordination emerged from agents that had been built to work in isolation, with no protocol ever designed for it. The capability made headlines; the week it took OpenAI to attribute the behaviour is the number an executive should sit with. Knowing what your systems are doing, and how fast you would know, is an operational asset that scales with your agent count and cannot be bought in.

Read the four together and a pattern shows up. Each one accumulates through use, and each gets stronger the longer you stay in one place and weaker every time you spread across another industry, another workflow, another logo on the roadmap.

What this does to a budget

Most enterprise AI budgets I see are shaped like a hardware purchase. The large numbers sit against models, compute and tooling, with a thin line at the bottom for people and process. That shape made sense when the model was the scarce input. It stopped making sense somewhere in the last eighteen months.

The four assets above are all labour and time. Writing down a domain ontology is somebody sitting with an expert for six weeks. Building an expectation library is a slow accumulation from real sessions, and DataCamp found roughly 60% of theirs only after launch. Serving economics is engineering work on caching and routing that never shows up in a demo. Attribution is instrumentation nobody asks for until the week they need it and do not have it.

None of that photographs well in a steering committee, which is most of why it goes unfunded. It also cannot be bought at the point it becomes urgent, which is the only point at which most organisations go looking for it.

The practical version of this for a budget holder is a ratio. If more than about 80% of your AI spend is going to things a competitor could have running by Friday, you are funding a capability rather than an advantage. Capabilities are worth buying at the best price you can get, on the shortest contract you can sign, because the price is falling and the differentiation was never there.

Three things to do this week

  1. Run the substitution test on one system. Pick the AI deployment your organisation is proudest of. Write two columns, Friday and not Friday, and put every component in one of them. Most teams find the Friday column holds almost all of the budget. That gap is the finding, and it is usually large enough to change a roadmap conversation on its own. Do it with two other people in the room so the arguments about what really counts as reproducible happen out loud.
  2. Price one unit of work. Choose a single repeated task your AI handles and calculate the fully loaded cost of doing it once, including inference, retries and the human time on either side. If nobody in the organisation can produce that number today, cost is being managed by hope. DataCamp reached $40 million a year before this became the first line on the CEO's list.
  3. Write down twenty expectations. Take one workflow and write twenty specific statements about what a correct output looks like, in plain sentences a domain expert would sign. Twenty is enough to expose how much of your quality bar has never left anyone's head. It is also the first twenty entries in an asset you will still be adding to in two years.

What to read next

While you are here, the back catalogue has more on this:

Anees Merchant writes one essay every Tuesday about enterprise AI, agentic systems, and the human side of the work. He is the author of Merchants of AI, a TEDx speaker, and a doctoral researcher in Human-AI Communication at the Swiss School of Business and Management.

The newsletter version, with extra commentary, goes out separately on Mondays. Subscribe at /newsletter.

See you Tuesday.

FAQ

Common Questions

What is the substitution test?

The substitution test is a twenty minute exercise for AI investment decisions. Imagine a well-funded competitor receives your architecture diagram on Monday and ask what they could have running by Friday. Components that survive the week are durable advantages worth funding. Components that do not are line items to be bought at the best price available and managed as costs.

How does a domain ontology differ from training data?

Training data is a corpus a model learns from. A domain ontology is the written account of what matters inside that work and why: which fields decide an outcome, which are ignored, the order an expert moves through the material, and what a correct answer looks like. Training data can be licensed or scraped. An ontology comes from years of watching the work happen and has to be written down deliberately.

Why do the best models on internal evals still lose in production?

Because the model is one layer of the system. DataCamp found that Gemma 4 31B beat frontier models on its own internal evals with untuned prompts, and still runs none of the live traffic. Speed, reliability and caching sit in the serving layer underneath, and a production workload is decided there. Winning an eval proves a model is good on paper. Frontier models still carry 100% of DataCamp's production traffic regardless.

When should an executive narrow rather than expand an AI product?

Narrow when the technical work is reproducible in about a quarter and the only durable advantage is depth in a single domain. Expansion across industries dilutes the one asset a competitor cannot copy at the exact moment it needs to compound. Expand geographically inside the same narrow slice instead, which grows the market without spending the depth.

What is the first step to finding a real AI moat?

Write down what your team knows that is not written anywhere. Interview the two or three people whose judgment the system depends on and record how they decide, in their words, in specific cases. Most organisations discover their strongest asset has never left a small number of heads, which makes it both the moat and the single biggest risk in the operation.