Human-AI Communication

The Limitation Inventory: Four Gaps That Cap What AI Does for You

The Limitation Inventory: Four Gaps That Cap What AI Does for You

A limitation inventory is a written list of what you cannot do, cannot name, and cannot picture, because those three gaps decide what you are able to ask an AI system for in the first place. The model answers the request. The request carries your limits into it. Four gaps show up in almost every weak prompt I read, and writing them down is the only reliable way I have found to close them.

Most people upgrade the model and wait for the output to improve. Then they conclude the technology was oversold. What actually happened is that a better system received the same small request and did a slightly better job of the small thing it was asked to do.

The question I could not answer about my own work

I sat down on Sunday to do something ordinary. Take a piece of work I had already finished, hand it to a model, and see whether it could push the thinking further. It could not. Not because it refused, and not because the answer was wrong. The answer was fine. It was fine in the same shape as the thing I had already written, which is the tell.

So I ran the obvious control. I gave the same brief to a colleague who works in a completely different field, watched what he asked for instead, and the difference was not subtle. He asked for the same analysis structured as a supply chain problem. I would never have thought to ask for that, because supply chains are not a place I have spent any time. The model was capable of both answers on Saturday. Only one of us could request the second one.

That reframed the whole thing for me. I had been treating the model as the variable and myself as the constant. It is the other way round. The system's capability is roughly fixed on any given day and shared by everyone with an account. The request is the part that differs, and the request is downstream of what I can picture, what I can name, and what I have been exposed to.

There is a result reported this month that puts a number on the same shape of failure. I am repeating it as reported, because I have not read the underlying paper myself and the figures come to me second hand. Research agents were given six days and a three thousand dollar budget to produce two papers. On execution they were strong. Hundreds of experiments run. Crashing hardware debugged without help. Camera-ready documents produced end to end with nobody touching them. Both papers were rejected, which is not the interesting part. The interesting part is that both runs finished with more than half the money unspent.

Read that number carefully. They did not run out of budget. They ran out of ideas about what to try next. When the reviews came back flagging problems, the agents narrowed their claims and added caveats. What they never did was step back and conclude that the direction itself was wrong and something else was worth attempting. The capability to execute was there. The capability to ask a different question was not, and the leftover budget is the receipt.

I recognised that immediately, because it is the same failure I had just watched in myself, and it is the same failure I watch in review meetings every month. The blank prompt box has the same problem in reverse. It looks generous. It behaves like a bill. It hands the entire burden of knowing what a system can do to the person who knows the system least, which is exactly why, on the vendor's own published seat numbers, the most widely deployed office assistant in the world has been paid for on roughly three percent of the seats it could run on, while a narrower tool from the same company that simply writes the next line of code for you holds its subscribers. One of them asks you to imagine. The other one does not need you to.

The enterprise version of this arrives on my desk pre-diagnosed. Almost every AI request I receive has already decided what it wants. We need a chatbot. We need this approval workflow automated. Nobody arrives with the problem, because arriving with the problem would mean sitting in the discomfort of not yet knowing the shape of the answer, and a request is easier to write than a question. So the request that lands is the first solution the requester could picture, and the whole programme gets built on top of it. Six months later the thing works exactly as specified and nothing measurable has moved, and the review calls it an adoption problem.

I have stopped treating those requests as requirements. The only useful first meeting is the one where nobody is allowed to name a tool. What backs up, where does work sit waiting, what does that waiting cost per month. Half the time the answer is not an AI problem at all, and the other half it is a different AI problem than the one on the ticket. That meeting is uncomfortable for everyone in it, and it is the highest return hour in the entire programme, because it is the only point where the picture is still open.

The four gaps

Four gaps sit between what a system can do and what you actually get from it. None of them are about the model. All four are visible on a single page once you write them down.

One. The naming gap. You cannot ask for what you cannot name. Every field has a vocabulary that compresses a whole method into two words, and if those two words are not in your head, the entire method is unavailable to you regardless of how good the system is. Ask for "a better summary" and you get a better summary. Ask for a pre-mortem, a red team pass, a base rate check, or a sensitivity analysis and you get four different pieces of work. The naming gap is the cheapest of the four to close and the one nobody audits, because not knowing a word feels like a small thing rather than a hard limit on your available requests.

Two. The picture gap. You can usually picture only one shape for the output, and you ask for that shape without noticing you chose it. A brief becomes a document because documents are what you have seen. It could have been a decision table, a set of three rival explanations scored against evidence, a one-page argument written against your own position, or a simulation of how the plan fails. The picture gap is the most expensive of the four, because it caps the output before a single word of the prompt is written and it leaves no trace. You never see what you did not request.

Three. The standard gap. You cannot improve what you cannot judge. If you do not know what excellent looks like in the thing you asked for, you will accept the first competent draft, and competent is exactly what these systems produce on the first attempt. This is why people with deep expertise get dramatically more out of the same tools. They are not writing better prompts. They are rejecting more first drafts, and each rejection is a real instruction the rest of us never send.

Four. The field gap. Every analogy you reach for comes from somewhere you have been. Work inside one industry for fifteen years and the model of how things work is that industry's model, so the questions you generate are that industry's questions. The non-obvious move almost always arrives from a field with no visible connection to your own. My colleague did not out-think me on Sunday. He had simply stood somewhere I had not stood.

                 THE SYSTEM                THE REQUEST
              (roughly fixed,            (yours, and the
             shared by everyone)          only variable)
                     |                          |
                     |            +-------------+-------------+
                     |            |      |          |         |
                     v          naming picture  standard    field
              capability          gap    gap       gap       gap
                     |            |      |          |         |
                     +------------+------+----+-----+---------+
                                              |
                                        WHAT YOU GET

Fix them in that order, because they do not cost the same. Naming is a weekend of reading and it is done. Picture takes a few weeks of deliberate practice and pays the largest return. Standard is slow, because judgment only comes from seeing enough examples to rank them, and no shortcut has ever worked for me there. Field is measured in years, which is precisely why it has to start now rather than after the other three are finished. People invert this constantly. They spend months on prompt technique, which sits underneath all four gaps, and never spend one hour on the vocabulary that would have opened a hundred new requests immediately.

The inventory is the artefact that closes them. One page, four headings, written honestly, in your own words, about your own work. Under naming, the vocabulary you have been nodding along to without owning. Under picture, the output shapes you default to and the ones you have never once asked for. Under standard, the areas where you genuinely cannot tell good from adequate. Under field, the domains you have never entered.

Mine is uncomfortable to read, which is roughly the test of whether it is any good. One line under standard says that I cannot reliably tell a well-built financial model from a plausible one, despite a finance degree, because I have not built one in years and the tells have moved. That single admission changed how I ask. I no longer request the model. I ask for the criteria an experienced analyst would use to reject it, and only then do I ask for the thing itself, so I have something to judge with when it arrives. One honest sentence produced a permanently better request.

That page is not self-improvement filing. It is a working document you point a system at. Once the gaps are written down, you can ask for exactly the thing you could not previously request, which is the whole point. Here is the shape I always default to, give me three other shapes. I cannot judge this domain, so show me the criteria an expert would score it against before you produce anything. Those are prompts you cannot write until you have admitted what is missing.

Two honest caveats. The inventory is a description of your limits at one moment, and it goes stale, so a page written in January is a poor guide by June. And it only works if nobody else is going to read it, because the moment you write it for an audience you start writing gaps you can defend rather than gaps you actually have.

Three things to do this week

  1. Write the one-page inventory. Block sixty minutes, take four headings, and be specific enough to be uncomfortable. Vague entries like "I should read more" produce nothing. "I cannot tell a good pricing model from a plausible one" produces a real prompt on Monday morning.
  2. Run the second-shape test on your next three requests. Before sending, write down the output shape you were about to ask for, then name two others and request one of them instead. Keep a note of which shape you actually used. Inside a week you will see your own default, and the default is the picture gap made visible.
  3. Spend two hours in a field you have no business being in. Pick something with no obvious connection to your work, go deep enough to learn one specific mechanism, and write one sentence on where that mechanism might apply in your own domain. One transfer per week compounds faster than any prompt technique on offer.

What to read next

While you are here, the back catalogue has more on this:

FAQ

Common Questions

What is a limitation inventory?

A limitation inventory is a one-page written list of what you cannot name, cannot picture, cannot judge, and have never been exposed to, organised under those four headings and written about your own actual work. It exists because the quality of what you get from an AI system is set by the request, and the request is built out of your vocabulary, your default output shapes, your ability to tell good from adequate, and the fields you have spent time in. Written down, each gap converts into a prompt you could not previously have composed.

How is a limitation inventory different from a prompt template library?

A prompt library gives you better wording for requests you already know how to make. An inventory changes which requests are available to you at all. The two solve different problems. Templates raise the floor on execution and are useful for repeated tasks with a known output shape. An inventory works one level up, on the choice of what to ask for, which is where most weak output actually originates. Someone with a large template library and an unexamined picture gap will produce the same narrow work faster and more consistently.

Why does output stay the same after upgrading to a better model?

Because the model was rarely the limiting factor. Capability on any given day is roughly fixed and shared by everyone with an account, while the request is the part that varies between users. A stronger system given the same small request returns a slightly better version of the same small thing. The visible improvement people expect from an upgrade comes from asking for something they had not previously thought to ask for, which is a change in the person rather than a change in the software.

When should a professional write a limitation inventory?

Twice a year is enough, plus once immediately after any move into a new role, sector, or problem type, because a move changes which gaps matter. It is also worth writing one at the point where AI has clearly stopped adding value to your work, since that plateau is usually a signal that you have automated everything you already knew how to do and have not yet expanded what you know how to ask for. Sixty minutes is sufficient. Longer usually means the entries are drifting toward generalities.

What is the first step to widening what you can ask an AI system for?

Take one request you sent this week, write down the output shape you asked for, then name two other shapes that would have answered the same underlying question. A document could have been a decision table, three rival explanations scored against evidence, a written case against your own position, or a description of how the plan fails. Send one of the alternatives and compare. That single exercise makes your default shape visible, and the default is the gap that caps output before a prompt is even written.