There is a sentence that appears, in one form or another, in every discussion about a disappointing AI output. It goes: it depends how you prompt it.
The sentence is true. It is also the single clearest signal that you are not looking at a product. It relocates responsibility for the result onto the person who paid for it, and it does so in a way that cannot be argued with, because there is always a better prompt you did not write.
We want to take that sentence seriously rather than dismiss it, because understanding why it is so structurally convenient explains most of what is wrong with how AI is currently sold.
The accountability gap, defined
In any normal supply relationship there is a moment where responsibility transfers. You commission a survey and the surveyor is responsible for its accuracy. You buy a component and the manufacturer is responsible for its tolerance. The transfer is what you are paying for, as much as the artefact.
General AI assistants have no such moment. The vendor supplies capability. You supply the request. The output is jointly produced, and when it is wrong there is no principled way to allocate the failure. In practice it allocates to you, because you are the one who has to deal with it.
That is the accountability gap. It is not a bug in the models. It is a property of the commercial shape.
Four consequences that follow
You cannot buy it properly
Procurement functions exist to allocate risk. Ask any procurement team to evaluate a supplier who will not state what they deliver, will not price the unit, and disclaims responsibility for the result, and they will tell you the arrangement is unbuyable at any serious scale. Which is why most enterprise AI spend is still classified as tooling rather than as supply.
You cannot improve it systematically
If the quality of the output depends on the phrasing of the request, then quality is a property of your staff rather than of the system. Improvement means training people to phrase things better. That knowledge lives in individuals, leaves when they do, and cannot be audited.
You cannot delegate it
A junior member of staff cannot be handed a task that requires prompt craft to get right, because the craft is the job. The tool that was supposed to reduce the expertise required to do the work has instead moved the expertise somewhere less visible.
You cannot defend it
When something goes wrong and someone asks how the decision was made, “the assistant suggested it and Priya adjusted the prompt until it looked right” is not an answer that survives a regulator, a client or an internal audit. It is barely an answer that survives a manager.
Why vendors like it this way
It would be unfair to suggest this is a conspiracy. It is simply the path of least resistance, and every incentive points along it.
Specifying an outcome is expensive. It requires domain expertise the vendor may not have, an evaluation set somebody has to build, and a willingness to be measured. Selling capability requires none of that. Ship the interface, let a million users find the applications, and let each of them own their own results.
There is also a genuine technical difficulty. Committing to an outcome means committing to an error rate, and committing to an error rate means someone has to pay for the errors. Nobody wants to be first to write that number down.
And there is a marketing advantage in vagueness. A product that does everything is more exciting than a product that does one thing with a published accuracy figure. Right up until a buyer asks what it will do for them specifically.
What a promise actually requires
If you wanted to close the gap, what would you have to build? It is a short list and none of it is exotic.
- A specification the buyer reads before paying. What goes in, what comes out, in what shape.
- A measurement against held-out data. Not a benchmark, a set drawn from the buyer’s own kind of work.
- A refusal path. The system declines rather than guessing when it is unsure, and a refusal costs nothing.
- A remedy. Published terms for what happens when it is wrong anyway.
- A version pin. The buyer’s results do not silently change because the vendor swapped models on a Tuesday.
That last one is underrated and increasingly important. If your output depends on a model you do not control, and that model is updated without notice, then a thing that worked in March produces something different in June and nobody told you. A promise that cannot survive a version change is not a promise.
The counter-argument from capability
The honest counter is that constraining a general system throws away most of its value. The reason a general assistant is useful is precisely that nobody had to anticipate your use case.
We accept this for exploratory work and reject it for production work, and the distinction is about who carries the risk.
When you explore, you are the expert. You can tell whether the output is good, you are not depending on it, and the cost of a bad answer is a few wasted minutes. Open capability is exactly right, and constraining it would be perverse.
When you put a task into production, the person consuming the output is usually not the expert, cannot verify it, and is acting on it. The risk profile has inverted completely. Continuing to sell open capability into that situation is selling a research instrument to someone who needs a supplier.
How to tell which one you are being sold
There is a simple test, and it takes one question: what happens if it is wrong?
If the answer involves better prompting, more context, a different model, or the phrase “it depends”, you are being sold capability. That may be fine. Buy it knowing you own the results.
If the answer is a specific remedy with a number attached, you are being sold a job. Now you can evaluate it the way you evaluate any other supplier.
Neither answer is dishonest. The problem is that the first is frequently delivered in the packaging of the second, and buyers reasonably assume that something with a price and a logo comes with the accountability that every other priced thing they buy comes with.
What the gap costs, in practice
The accountability gap is not an abstraction. It shows up as specific, recurring costs inside organisations that have adopted general assistants without noticing what they bought.
Verification labour. Somebody checks the output. If the task was worth automating, the checking is usually done by someone senior enough to spot a subtle error, which means you have replaced a cheap task with an expensive one. We have seen teams where the review of AI-drafted work takes longer than writing it did, and nobody had measured this because nobody owned the number.
Inconsistency across a team. Five people doing the same task with the same tool produce five different standards, because each has developed their own prompting habits. The variance is invisible until an auditor or a customer sees two outputs side by side.
Silent drift. The vendor updates the model. Nothing announces this. Output that was reliably formatted a certain way starts occasionally deviating. Because there was never a specification, there is nothing to detect the deviation against, and it surfaces weeks later as a downstream integration failure.
Institutional knowledge with no home. The person who worked out how to get good results leaves. Their prompts might be in a shared document, or they might be in their own notes, or in their head. What they knew about which phrasings failed is gone entirely, because failures were never recorded.
None of these appear on the invoice. All of them are real costs of a purchase that looked cheaper than building something specified.
A question worth asking your own team
If you already have general assistants in use, there is a diagnostic that takes an afternoon and tends to be uncomfortable.
Pick a task the tool is used for regularly. Ask three people who do it to produce output for the same five inputs, independently, without comparing notes. Then have the person who owns the quality standard mark all fifteen results against what they would have accepted.
Two things usually come out. The variance between people is larger than anyone expected, and the marker has trouble articulating the standard they are marking against, because it was never written down. That second finding is the more important one: it means the specification does not exist anywhere in the organisation, and the tool has been quietly filling the gap with whatever each individual assumed.
That exercise is also, conveniently, the first half of building a proper evaluation set.
What we do about it
On client engagements we will not put a probabilistic system into a production path without the refusal rule and the escalation queue, and we say so before the proposal. It occasionally costs us work, because a queue implies a person, and a person implies the saving is smaller than the pitch suggested.
We would rather lose that argument at proposal stage than at the point where a wrong output reached a customer. A system that guesses confidently is not cheaper. It moves the cost from a queue you can see to an error rate you cannot.
The same reasoning drives BotUp. Every job carries a published deliverable, a price and refund terms, because the alternative is to sell something whose failure mode is the buyer’s problem. And it drives how we sell our own services: a written scope and a fixed price, with the comparison against conventional delivery stated in the proposal. If our estimate is wrong, that is ours to absorb.
The sentence, revisited
It depends how you prompt it is a true statement about a research tool and an evasion when it comes from a supplier. The next few years of this industry will be spent working out which of those each vendor is, and most of them will not choose voluntarily. Buyers will choose for them, by asking the only question that matters and declining to accept a technical answer to a commercial one.
If you are trying to work out which side of that line a system you depend on sits, we are happy to look at it with you. We will tell you if the honest answer is that it belongs in the chat window after all.