The agent marketplace thesis, in full

BotUp argues AI work should be bought as defined jobs with specs, prices, evaluation and remedies, rather than as access to a model.

This is the full version of an argument we have been making in pieces all month. It is the thesis behind BotUp, set out properly, including the parts we are least sure about.

Our claim is simple. AI capability is being sold in the wrong shape. The right shape is a marketplace of defined jobs, each with published specifications and prices. We think a set of unglamorous commercial and engineering artefacts, rather than model quality, is the thing standing between here and there. Almost nobody is building those artefacts.

AI is usually sold as access, which puts the risk on the buyer

Almost every AI product is sold as access. You buy tokens, or seats, or a rate limit. What you receive is the ability to attempt things.

That makes sense for a research instrument. It is a strange way to sell work. In every other market, a buyer describes an outcome and a supplier commits to it. The supplier’s expertise shows up in two ways. They do the work, and they are willing to name a price before they know exactly how hard it will be.

The AI industry has inverted that. It has taken the most capable general tool built in a generation and sold it in the one shape that transfers all risk to the customer.

AI was sold this way because access was easier to sell

This did not happen because of a conspiracy. It happened because access was the path of least resistance, and every incentive pointed in that direction.

Defining a job is expensive. It needs domain expertise, an evaluation set somebody builds, and a commercial person willing to price against an uncertain cost base. Selling access requires none of that. Ship the interface, let a million users find the applications, and let each of them own their results.

Pricing probabilistic work is also hard. The system can be right often and still be wrong sometimes. Commit to an outcome and you commit to an error rate, and somebody has to pay for the errors. Access-based products never face that question.

Vagueness also sells. A product that does everything is more exciting than one with a published accuracy figure, right up until a buyer asks what it will do for them specifically.

A defined AI job needs five artefacts

A defined job needs Five artefacts. Miss one and you are selling access with a price tag attached.

  • An input contract, saying what will be accepted and what happens outside those bounds.
  • An output contract, specifying precisely what comes back rather than describing it.
  • An evaluation set, held out, produced by someone who knows the domain, which makes the acceptance criterion executable.
  • A confidence and escalation rule, so the system declines rather than guesses and a refusal is not billed.
  • A remedy, published before purchase, saying what the buyer gets when it is wrong anyway.

The last artefact is where vendors resist hardest. It is also what makes everything above believable. Anyone can claim high accuracy. Only someone who means it writes down what happens in the remainder.

A marketplace fits the economics of specification work

The obvious alternative is that individual vendors sell individual jobs. Some will. But the specification work has an economic property that favours a marketplace.

Writing a job specification is expensive, and it is done once. After that, it serves every buyer of that job forever. That is a classic case for a market. The specification cost should sit with whoever has the domain expertise, and the demand should be able to find it.

A marketplace also solves discovery. A buyer with a task does not know whether it is specifiable, who can do it, or what it should cost. A marketplace of published specifications answers all three by existing.

It also creates the comparison that access-based selling prevents. Two suppliers offering the same job with different accuracy figures, prices and remedies is a market. Two vendors offering access is a brand preference.

Four conditions have to hold for this to work

Four conditions have to be true, and we are more confident about some than others.

  • Enough tasks must be specifiable. We think a great many are: document handling, extraction, categorisation, routine correspondence, enrichment, translation of non-critical content, review at volume.
  • Buyers must prefer accountability to flexibility. This is the load-bearing assumption.
  • Evaluation must be trustworthy. A marketplace of published accuracy figures only works if the figures mean something.
  • The economics must work without a large take. If the marketplace charges fifteen percent, the jobs have to be priced to absorb it and the model competes badly against a vendor selling directly.

Those specifiable tasks make up a large market, and it is currently sold as access. We believe organisations buying at scale want a price, a specification and a remedy more than they want an open interface. Individual enthusiasts want the opposite, and they are loud.

Trustworthy evaluation requires held-out sets, methodology that can be inspected, and some mechanism against sellers grading their own homework.

Trustworthy evaluation is the hardest condition

Trustworthy evaluation is the hardest of those conditions, by a distance.

A seller publishing their own accuracy figure has every incentive to choose a favourable set. The obvious answer is marketplace-run evaluation on a standard set. That pushes toward benchmarks, and benchmarks get gamed and stop reflecting real use.

Our current position is buyer-supplied evaluation. A buyer can submit their own sample and receive a measured result on their actual data before purchasing. This is expensive and honest. It also makes the number specific to that buyer, rather than to an average.

We are not confident this is the final answer. It is the best one we have, and it is the part of the thesis most likely to change.

We have built the job machinery, but the market is still thin

What we have built

We have built the job specification format with the five artefacts, the evaluation harness that backs each job, the refusal-and-no-charge rule, and published remedies.

What remains weak or missing

Discovery is weak, which is the same lesson we learned on Open Lance and apparently had to learn twice. Buyer-supplied evaluation is expensive and slow. The marketplace is thin, which is the ordinary cold-start problem and no thesis makes it easier.

If this is right, buyers choose measured jobs instead of model access

The end state is easier to judge as a concrete picture than as a principle.

A buyer with a recurring task searches for it the way they would search for a supplier. They find several published jobs with the same input contract, different accuracy figures, different prices and different remedies. They submit a sample of their own data, receive measured results from each, and choose.

Integration is a purchase order and an API key rather than a project. The job has a stable contract, so the buyer’s system is not coupled to any particular model or provider. When a better supplier appears, switching is a configuration change rather than a rebuild.

Sellers are not primarily model companies. They are domain experts who understand invoices, or clinical coding, or shipping documentation, and who have done the specification work. The model underneath is a commodity input they choose on price and performance, and change without their buyers noticing.

That last property is what makes it a market rather than a directory. Buyers purchase an outcome and become indifferent to the mechanism. That is the normal condition of every mature supply relationship, and the opposite of how AI is bought today.

Some work should stay outside this model

The boundary matters, because a thesis that claims everything explains nothing.

Genuine exploration stays in the open interface. When you are the expert, when you cannot say in advance what you want, when the cost of a poor answer is a few wasted minutes, an open tool is exactly right and constraining it would be perverse.

Decisions with legal or clinical consequence do not become jobs. The issue is accountability. It cannot be transferred to a supplier. Those need a named person and always will.

Anything genuinely novel resists specification by definition. If nobody has done it before, there is no ground truth, no evaluation set and nothing to promise.

What is left is still enormous: the repetitive, well-understood, high-volume work that constitutes most of what organisations actually need done.

This would reduce some consultancy work, including some of ours

This thesis is bad for a certain kind of consultancy revenue, including some of ours.

A large share of AI services work today is building bespoke pipelines for tasks that dozens of other organisations also need. If those become purchasable jobs, that work stops being a project and becomes a subscription, and the integrator’s role shrinks to connecting things.

We would rather be the marketplace than the integrator being disintermediated by it. That is a self-interested position, and we would rather say so plainly than dress it up as foresight.

Three outcomes would prove this thesis wrong

A thesis with no falsification condition is a slogan, so we should state what would prove us wrong.

  • If buyers overwhelmingly prefer open capability even at scale, we are wrong about the core assumption and BotUp is a niche for cautious enterprises rather than a market structure.
  • If the specifiable set turns out to be much smaller than we think, because real tasks resist specification more than they appear to, the addressable market is a fraction of what we assume.
  • If model capability advances such that a general system reliably self-specifies, producing a defensible contract for any task without human domain work, then the expensive artefact becomes cheap and the marketplace advantage disappears.

The third is the one we watch. It is not implausible, and it would not make the jobs framing wrong. It would make the marketplace unnecessary.

We are betting on defined jobs because the alternative is worse

We are betting on this because the alternative bet is worse. Building a general assistant means competing with organisations spending more on a single training run than our entire history of revenue.

The shape also follows from something we already believe about our own business. We sell software on fixed scope with a written definition of done and a comparison to the conventional alternative, and we decline work we cannot beat. Selling AI work as defined jobs with published specifications and remedies is the same argument applied to a different product.

If we are right that buyers want accountability more than capability, both positions are correct. If we are wrong, both are wrong together, which at least makes the bet legible.

BotUp is the implementation, and the reasoning about what a defined job contains ran through several pieces earlier this month. If you have a task you think is specifiable, tell us about it and we will tell you honestly whether it is.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded