The chat window was never the product

Selling AI as a conversation hides the thing a buyer actually wants, which is a finished piece of work with a price and a deadline attached.

Open almost any AI product built in the last three years and you land in the same place: a text box, a blinking cursor, and an implicit instruction to be clever. The interface is an invitation to experiment. It is also, commercially, an abdication. It hands the entire burden of specification to the person who came to you because they did not want to do the work.

We have built enough of these systems for clients, and one of our own, to have formed a view. The chat window is a magnificent development environment and a poor product. It is where you discover what a model can do. It is not where most people should be buying what a model does.

The buyer did not come for a conversation

Consider what actually happens when a marketing manager needs forty product descriptions written to a house style. They do not want a partner in thought. They want forty product descriptions. The conversation is overhead, and every turn of it is a chance for the output to drift from what they meant.

What they get instead is a negotiation. They write a prompt. The result is close but the tone is wrong. They write a follow-up. Now the tone is right but it has invented a product feature. Another follow-up. Twenty minutes later they have something usable and a private theory about which magic words work. That theory is not transferable, not documented, and not the same as the theory their colleague developed for the same task.

The industry has a name for this. It calls it prompt engineering, and treats the skill as a feature of the ecosystem. From the buyer’s side it looks like being handed the raw materials and being told that assembly is part of the fun.

Three things a chat interface cannot give you

The gap is not aesthetic. There are specific commercial properties that a conversational interface structurally cannot provide, and every one of them is something a buyer expects from anything else they purchase.

A price known in advance

You do not know how many turns a task will take. Neither does the vendor. Pricing is therefore per token or per seat, which is a price for access rather than a price for outcome. The buyer carries the risk that the task turns out to be hard. In every other purchasing relationship they have, the supplier carries that risk.

A definition of done

When is the conversation finished? When you stop, which is usually when you are tired rather than when the work meets a standard. There is no acceptance criterion because nobody wrote one. The thing that makes fixed-price software delivery possible, a written scope both parties agreed to, has no equivalent here.

Somewhere to complain

If the output is wrong, whose fault is it? The honest answer under a chat model is that it is yours, because you asked badly. That answer is technically defensible and commercially useless. It means there is no refund, no service level, and no recourse. You bought access to a capability, and you got access to a capability.

Compare it to how you buy anything else

Think about commissioning a photographer. You do not buy an hour of their attention and iterate toward a photograph. You agree a brief: these subjects, this style, this many usable frames, delivered by this date, for this fee. If the shoot fails, that is their problem to fix. The brief is the product. The photography is the mechanism.

The same is true of a translation, an audit, a survey, a legal opinion. In each case the buyer describes an outcome and the supplier commits to it. The supplier’s expertise is expressed partly in doing the work and partly in being willing to name a price before knowing exactly how hard it will be.

AI vendors have inverted this. They have taken the most capable general-purpose tool built in a generation and sold it in the one shape that transfers all risk to the customer.

What changes when you sell the job instead

The alternative is not complicated to describe. Publish the job, not the capability. Say what will be delivered, what it costs, how long it takes, and what happens if it is wrong. Then take payment for that, not for access.

This is the argument behind BotUp, and it is worth being specific about what it changes:

  • The specification moves to the seller. Someone who understands the task writes down what a good result looks like before anyone pays. That work happens once and serves every buyer of that job.
  • The price is a number. Not a rate, not a credit balance. A figure you can put in a purchase order.
  • There is an evaluation harness behind it. If you promise a deliverable you need a way to know it was delivered. That is a test suite for a probabilistic system, and building good ones is a genuine speciality.
  • Refund terms exist. Published before purchase. This is the part vendors find uncomfortable, and it is the part that makes the rest credible.

None of that requires a better model. It requires a different commercial posture, and a fair amount of engineering behind the posture.

The objection, and why it is weaker than it looks

The standard objection is that general capability is the whole point. Constrain the model to defined jobs and you throw away the flexibility that makes it valuable. You end up with a worse version of traditional software.

There is something to this, and it is why the chat window will not disappear. For genuine exploration, for the tasks nobody has done before, an open interface is correct. Researchers, developers and anyone doing novel work should have one.

But look at what people actually buy. The volume is not in novel work. It is in the same forty tasks repeated across thousands of businesses: summarise these documents, extract these fields, draft this category of correspondence, categorise this backlog, translate this catalogue. Each of those is well defined. Each of them is being sold today as raw capability, and each of them is a job somebody would happily buy at a fixed price if it existed.

The flexibility argument protects the vendor, not the buyer. It is the reason the vendor does not have to write a specification.

What this looks like in an enterprise

We see the same pattern inside client organisations. A company buys seats for a general assistant. Six months later, adoption is uneven, nobody can say what it produced, and the renewal conversation has no numbers in it. The tool did useful things. Nobody can prove which.

Compare that with a system built around defined jobs. Invoice extraction runs nine thousand times a month, with a measured accuracy against a held-out set, and a queue where anything below a confidence threshold reaches a person. You know the volume, the error rate, the cost per unit and the human hours displaced. That is a line item a finance director can evaluate.

The difference is not the model. It is often the same model. The difference is that somebody did the work of turning a capability into a job.

The uncomfortable part for us

We build software for a living, and this argument cuts against a comfortable business. It is easier to sell a discovery phase, a platform build and an ongoing retainer than to name a deliverable and a price. The chat-shaped product has an analogue in professional services: billing for time rather than for outcome, which protects the supplier from ever having to be precise about what they are producing.

We changed our own model for exactly this reason, and wrote about what it cost us on the engagement models page. Every proposal now names a fixed price against a written scope, with the conventional cost stated beside it. If we cannot beat that meaningfully we say so and decline. That is uncomfortable in the same way publishing refund terms on an AI job is uncomfortable, and for the same reason: it removes the vendor’s ability to be vague.

Where the interface still matters

None of this is an argument that interfaces do not matter, or that everything should become a form. The interface still has to do three things well, and most job-shaped AI products do them badly.

It has to make the job’s boundaries obvious before purchase, so nobody buys the wrong thing. It has to show the work rather than only the result, because a buyer who cannot see how a conclusion was reached will not act on it. And it has to make the failure path visible: what happens when the job cannot be completed, who is told, and what the buyer gets instead.

Those are ordinary product design problems. They are considerably more tractable than teaching every buyer to be a prompt engineer.

What we would ask a vendor

If you are evaluating an AI product, the questions that separate a product from a demonstration are commercial rather than technical:

  • What exactly will this deliver, written down, before I pay?
  • What does it cost per unit of that deliverable, not per token?
  • How do you measure whether it succeeded, and can I see the measurement?
  • What happens when it fails? Who finds out, and what do I get?
  • If the underlying model changes next quarter, what protects the output I depend on?

A vendor selling a chat window cannot answer most of those. Not because they are evasive, but because the shape of the product does not contain the answers.

A worked example, in numbers

Abstractions are easy to agree with, so here is a real shape from a client engagement, with the figures rounded and the sector changed.

A distributor received about four thousand supplier documents a month: purchase order confirmations, delivery notes and invoices, in maybe thirty different layouts, mostly PDF, some of them scans of faxes. Three people spent most of their week reading these and typing eleven fields each into an ERP. The error rate on manual entry was around two percent, which sounds small until you notice it was two percent of four thousand, and each error cost roughly forty minutes to find and unwind.

The first proposal they received was seats for a general assistant. Staff would paste documents in and copy the extracted fields out. It would have worked, in the sense that the model was capable. It would also have kept every one of the properties described above: no price per document, no accuracy measurement, no definition of done, and an error rate that depended on how each of the three people happened to phrase things that day.

What we built instead was a job. One defined task: given a supplier document, return eleven fields with a confidence score on each. The specification was written first, with a held-out set of six hundred documents that a person had already keyed correctly. That set is the acceptance criterion. The system is measured against it on every deployment.

The commercial shape that came out of it: a cost per document of a little under two pence, a measured field-level accuracy of 99.4 percent against the held-out set, and anything below a confidence threshold routed to a person rather than guessed. Three people became one person handling exceptions. The finance director could evaluate that, because it was expressed in the units their business already used.

Nothing in that project required a frontier model. It required someone to write down what correct meant, and to build the harness that proves it. That is the work the chat window lets everybody skip.

Why vendors resist this

It is worth being fair about why the industry has settled where it has, because the reasons are not stupid.

Defining a job is expensive. You need a domain expert to say what good looks like, an engineer to build the evaluation, and a commercial person willing to name a price against an uncertain cost base. Selling access requires none of that. You ship the interface and let a million users discover the use cases for you, at their expense.

It is also genuinely hard to price probabilistic work. If your system is right 99 percent of the time, the one percent has a cost, and somebody has to carry it. Vendors selling access never have to answer that question. Vendors selling jobs have to answer it in the refund policy, in public, before anyone buys.

And there is a real fear of being pinned down. A published deliverable is a published limitation. It says here is what we do and, by omission, here is what we do not. Access-based products get to imply they do everything.

Those are understandable commercial instincts. They are also exactly the instincts that a buyer should be suspicious of, because every one of them moves risk from the party best able to manage it to the party least able to.

The short version

The chat window is where AI was discovered, and it deserves credit for that. It is a research instrument that escaped into the consumer market and became the default shape of an entire industry’s products by accident rather than by design.

The businesses buying AI are not buying a research instrument. They are buying work: specified, priced, delivered and accountable. The vendors who make that transition will look less impressive in a demonstration and considerably more attractive on a purchase order.

We are betting the second thing matters more. If you want to see what that looks like in practice, BotUp is the argument implemented, and we are happy to talk about applying the same shape to whatever you are trying to automate.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded