Usable AI output is a constraint problem because the model is already capable enough for most of the work you hand it. The gap between a result you ship and a result you delete is how much you took off the table before the model started writing.
I spent a while learning that the expensive way.
Something specific changed, and it is worth naming. For years the honest answer to a disappointing result was that the model could not do it. Ask for a competent first draft of a technical brief in 2023 and you got something you had to rewrite from the first line. That era ended quietly. The models I use daily now clear the bar on capability for almost everything I put in front of them, and the failures that remain look nothing like incompetence. They look like an answer that is correct, complete, well organised and unusable, because it opened with two paragraphs of preamble, or answered a slightly different question with total confidence in a register borrowed from a consultancy brochure. You cannot fix those by asking for more.
The week my ban list beat my prompt
I run an automated pipeline that drafts one long essay a week. The first version worked exactly as badly as you would expect. It produced fluent English at the right length and on the right topic, in a voice that belonged to nobody. Every draft was publishable and none of it was mine.
So I did what everyone does. I fed it more context. I added examples of my own writing. I lengthened the brief until it ran to several pages. I moved to a stronger model. The drafts got smoother while staying exactly as far from my voice as they had started.
Then I inverted the whole thing. Instead of describing what I wanted, I wrote down what the output was forbidden to contain. A list of banned words, the ones that show up in machine prose and almost never in speech. Zero em dashes, because I do not use them and their presence is a tell. A required shape for the opening sentence. An exact count of questions at the end. A word band with a hard floor and a hard ceiling. Then I wired a script that reads the finished draft and refuses to publish when any of those conditions fail, printing the specific violations rather than a general complaint.
The version of that pipeline that finally produced work I would sign has the shortest brief and the longest list of prohibitions of any version I have built. The brief shrank by roughly two thirds. Quality moved when the option space shrank, and only then.
What the checks catch is instructive. Most weeks the script passes on the first run. When it fails, it almost never fails on substance. It fails because a stray dash slipped in, or the closing section ran to six questions instead of five, or the opening sentence went for atmosphere instead of stating the thing. Small and invisible to me on a screen at eleven at night, and every one of them a signal a reader would register without being able to name. A person reviewing the same draft would wave all of it through. The script does not get tired and does not extend the benefit of the doubt, which is the entire point of putting the rule somewhere other than my own attention.
The same month, a smaller failure taught me the same lesson from another angle. An image step in one of my content pipelines was downloading generated files and reporting success. Six of the seven files that week were truncated mid transfer. They kept valid headers and correct dimensions, so every cheap check passed, and they rendered half blank once published. The fix was a verifier that walks the file to its end marker and refuses to return success on a partial download, sitting in front of a model and a retry loop that were both working correctly. One refusal, placed at the single point every workflow goes through.
Both fixes are the same move. Both remove permission to be wrong in a specific way, without adding any capability at all.
The four constraints
Watch what people actually adopt around AI systems and a pattern shows up quickly. The tools that stick are not adding intelligence to the model. They are fencing it. There are four fences, and most useful tooling is one of them wearing a product name.
THE FOUR CONSTRAINTS
SHAPE fix the form of the answer before generation
SURFACE fix what the model can see before generation
SPEND fix what being wrong costs during generation
SET fix the options on the table before you ask
capability is assumed. the fence is the product.One. Shape. Fix the form of the answer before it is generated. The clearest example I saw recently is an agent skill that forces writing into the structure consultants are trained on: give the answer first, then the reasons, then the evidence. It teaches the model nothing about the subject. It removes the model's freedom to warm up for three paragraphs before saying anything. A related tool does the same job for diagrams, replacing the identical flowchart every assistant produces with a house style that has rules about hierarchy and callouts. Both of them constrain form, and form is where most AI output falls apart. The content was fine. The arrangement made it unreadable.
Two. Surface. Fix what the model can see. Anyone who has watched an agent lose the plot between sessions has felt this. The instinct is to add memory. Organising what already exists works better, because it makes retrieval boring and predictable, whether that means a file tree an agent can search or a graph built over scattered company records. Worth being honest about the ceiling here, because the people building these systems are honest about it. Structure imposed automatically buys you a little. Structure imposed by a person who spends months on the taxonomy buys you the rest. Point an agent at an unsorted drive and you have reproduced the search box you were already failing to use.
Three. Spend. Fix what being wrong costs. Model routers make this concrete. One pattern sends every request to the cheapest model in your stack first, checks the result against the request, and escalates only when the answer does not hold up. Small local model, then a cheap hosted one, then the expensive one, and most requests never reach the top of the ladder. The cost saving is visible immediately. What matters more is that you have priced failure deliberately instead of paying premium rates for every trivial call, which is the default behaviour and the reason for the bill nobody forecast.
Four. Set. Fix the options on the table. This is the least technical and the most neglected. Given a task, the reflex is to hand the whole problem over. Find me a weather provider. A directory of free public APIs is one of the most starred repositories on the internet, and its value is that you can read the list yourself, form an opinion, and then ask the model to argue with you about which option fits. Any model can find an API on its own. That was never the difficulty. You get a better answer because you narrowed the field before asking, and you can tell whether the answer is good, which matters more.
Four constraints, and every one of them is a subtraction. Shape removes formats. Surface removes context. Spend removes spending. Set removes options. All four make the output usable rather than the model smarter, and usable is the only property you were ever paying for.
The reason this stays counterintuitive is that constraints feel like admissions of weakness. Writing a ban list feels like conceding that the tool cannot be trusted. Capping a router at a cheap model feels like settling. It reads as a downgrade right up to the moment you compare outputs, and then it stops reading that way permanently.
How to tell a real constraint from a decorative one
Most teams already believe in constraints. Very few have any. The difference is testable, and three questions settle it.
Can it refuse? A constraint that cannot block anything is advice. Instructions inside a prompt are advice. Style guidance in a shared document is advice. A check that reads the finished artefact and stops the process is a constraint, and the gap between the two is the gap between a preference and a rule.
Does it name the reason? A step that fails with a generic error trains everyone to rerun it and hope. I lost a full day once to a button that reported only the word Failed, when the underlying service had been returning a precise and helpful explanation the whole time, thrown away by my own error handling. A refusal that says which rule broke turns a mystery into a fix. A refusal without a reason sends the next person hunting in the wrong place, which is worse than no check at all because it carries the authority of a check.
Does it sit at the choke point? A rule enforced in three places out of four is not enforced. Find the single step every path goes through, whether that is the publish call, the file download, the model request, and put the check there. Scattering the same validation across every caller guarantees that the one you forget is the one that ships.
Run those three questions across your own workflow and the results are usually uncomfortable. Almost everything people describe as their AI guardrails turns out to be advice written down neatly.
Three things to do this week
- Write the ban list before the brief. Take the one recurring AI task you dislike the output of and spend fifteen minutes writing what the answer must never contain. Words, formats, openings, lengths. Run the task again with a shorter brief and that list attached. Expect the result to be worse on polish and better on fit.
- Put one refusal in one pipeline. Find a step that currently reports success without checking anything, which in most workflows is a download, a file write, or an API call whose response nobody reads. Add a check that fails loudly with the reason. You will find something broken that has been quietly broken for weeks.
- Add a cheap-first step to one repeated task. Pick the task you run most often and route it to the smallest model you have access to. Read ten outputs. Decide honestly how many needed the expensive model. In my experience the honest number is under half.
What to read next
While you are here, the back catalogue has more on this:
- The Limitation Inventory: Four Gaps That Cap What AI Does for You
- Every Agent Demo Sells the Ceiling. Adoption Is Decided by the Floor.
- The Loop Stack: How to Stop an AI Agent From Gaming Its Own Goal
Footer
Anees Merchant writes one essay every Tuesday about enterprise AI, agentic systems, and the human side of the work. He is the author of Merchants of AI, a TEDx speaker, and a doctoral researcher in Human-AI Communication at the Swiss School of Business and Management.
The newsletter version, with extra commentary, goes out separately on Mondays. Subscribe at /newsletter.
See you Tuesday.