What we think happens next

We close thirty pieces with our five predictions, the parts we doubt, and the bets we are willing to have checked later.

This closes a month of daily writing. We have published Thirty pieces, and most of them make the same broad argument. The interesting problems in AI are commercial and organisational rather than technical, and the industry is selling the wrong shape of thing.

It seems right to end by saying our position plainly. That includes the parts that could turn out to be wrong. Predictions with no falsification condition are decoration.

Some parts now look settled to us

Three things now feel established to us, rather than forecasts.

  • Capability is commoditising faster than anything else in the stack. The gap between the best model and a good enough one has narrowed. Everyone buys from the same handful of suppliers. Differentiation based on model choice is a procurement decision a competitor replicates in an afternoon.
  • The engineering around the model is where systems succeed or fail. Every serious production incident we have been called to was caused by truncation, retries, schema drift, permissions or monitoring gaps. Schema drift is when the shape of the data changes underneath the system. None was caused by insufficient reasoning.
  • The accountability line does not move. Across twelve industries the stopping point is identical: where a named human must carry responsibility for a specific decision. That line is set by regulation, liability and professional practice. It is not a capability gap waiting to close.

These are the five predictions we are signing

These are the five predictions, with the condition that would prove each one wrong.

  1. AI work gets bought as defined jobs rather than as access. Specification, price, measured accuracy and published remedy replace tokens and seats for the large repetitive middle of what organisations need. Wrong if buyers keep preferring open capability even at scale.
  2. Provability becomes the enterprise differentiator. Procurement asks what can you show me rather than what can it do. Wrong if regulatory attention fades or enterprises keep buying on demonstrations.
  3. Time-based billing for well-defined software work becomes unusual. It does not disappear, but the default inverts. Wrong if buyers continue to value flexibility over price certainty, which remains possible for exploratory work.
  4. The middle of the freelance market keeps thinning. Volume work compresses, accountable work strengthens, and competent execution of well-specified tasks at moderate rates continues to erode. Wrong if a new category of work emerges at that level, which has happened before.
  5. Reputation becomes portable, driven by enterprise procurement rather than platform goodwill. Wrong if no buyer ever has enough leverage to demand it, which is the status quo and may simply persist.

We are least sure about timing and system buyers

The timing on all five predictions is uncertain. There is also one substantive thing we are least sure about.

We do not know what happens when the buyer of work is itself a system. If an organisation’s process can purchase a defined job programmatically, a marketplace of people becomes something closer to a supply chain. Most of our reasoning about accountability assumes a person at both ends.

We are building for it, and we are not confident about the shape or the timeline. Anyone claiming to be is guessing with more conviction than we can manage.

We have changed our mind on three things

Over the past two years, three assumptions have changed.

  • We thought trust was the constraint on marketplaces. Discovery was the harder constraint. We engineered escrow, disputes, verification and reputation, and users struggle most with finding the right person at all. That was a correct answer to the wrong problem.
  • We thought engineering quality was portable across sectors. It travels, yet by itself it did not carry the work. The firms doing best are the ones who can discuss an audit, a settlement cycle or a school bell without needing it explained. Sector depth mattered more than we assumed and took longer to build than repricing did.
  • We thought the model was the interesting part. Everything we have learned since says the specification is the interesting part: writing down what correct means, precisely enough to measure. That is domain work, it is expensive, and it is the thing almost nobody wants to do.

The month compresses into one argument

Thirty pieces is a lot to ask anyone to read, so this is the argument compressed into one page.

  • Products. The chat window is a research instrument that escaped into the consumer market and became an industry’s default product shape by accident. It cannot give a buyer a price known in advance, a definition of done, or anywhere to complain. Selling defined jobs fixes all three and requires five artefacts most vendors will not build.
  • Marketplaces. Commission is charged as a percentage of value while the underlying costs are flat. That is a position rent, defended by contract terms preventing users from routing around it. Removing it is possible if escrow is cheap enough and disputes are designed out rather than adjudicated cheaply.
  • Engineering. The failures that matter are truncation, retries, false confidence, schema drift, silent model changes and leaked evaluation sets. Retrieval quality sets the ceiling. Retrieval is whether the system brings back the right material before it answers. Observability, the monitoring around the system, has to be rebuilt around whether the system is still as right as it was. Nothing announces itself when it is wrong.
  • Commerce. The hourly rate stopped describing anything useful when delivery time compressed. Fixed scope with a published comparison is the honest replacement, and it costs the supplier the ability to be vague.
  • Sectors. The constraint is never technological. It is a bell, a settlement cycle, an audit, a rights window. Understanding it is the work.

Three questions came up again and again

These questions came up more than once this month, so these are our brief answers.

How much of this is AI-specific

Less than it looks. Most of what we argue is ordinary supplier discipline applied to an unfamiliar noun: specify the deliverable, measure it, say what happens when it fails, and name who is accountable.

The genuinely new part is the component at the centre. It produces confident, well-formed, plausible output when it is wrong, so every safeguard has to assume wrong looks exactly like right.

Whether this argues against our own consultancy

Partly, yes. If defined jobs become purchasable, a large share of current AI services work stops being a project. We would rather be the marketplace than the integrator disintermediated by it. That is a self-interested position, and we would rather state it than dress it up.

Why we publish the method and the failures

A commitment nobody can check is a description of intentions. The failures are what make the rest credible. A firm that publishes only its successes is telling you what it wants you to believe rather than what it knows.

This is what we would tell someone starting now

The advice we would give someone starting now is already in the month of writing.

  • Build the evaluation before the feature. If you cannot score it you have not specified it, and you will not know whether anything you do afterwards helped.
  • Fix retrieval before touching the prompt. The ceiling on answer quality is whether the right material was present, and no prompt recovers from an index that did not surface it.
  • Design the refusal path first. A system that declines when unsure gets trusted and used. One that always answers gets verified constantly and saves nobody anything.
  • Write the mandate before issuing the credential. What may this touch, to what limit, on whose authority, and how is it reversed. One page, before any code.
  • Put the human queue in the business case. It is almost always the dominant cost and it is almost always left out.

The old rules of good software still apply

The things that made software work well before still make it work well now.

  • A written scope agreed before building.
  • Tests before a rewrite.
  • Reversible deployment.
  • Documentation for a stranger.
  • Someone accountable by name.

Every one of those predates this technology by decades. Every one matters more because of it, since the component at the centre is unusually good at hiding its own mistakes.

The main thing AI changed about our work is that discovery got dramatically cheaper. That made a category of project viable that never was. Good still looks the same.

Our bets all rest on the same proposition

We are betting on Two products. Both are bets on the same proposition: buyers want accountability more than capability.

  • Open Lance argues that marketplace commission is a tax on work rather than a price for a service. It also argues that a platform which cannot trap its users has to be worth staying with.
  • BotUp argues that AI should be sold as defined jobs with published specifications, prices and remedies rather than as access to a model.

Our services business runs on the same logic: fixed scope, fixed price, the conventional comparison stated beside our number, and work declined when we cannot beat it.

If the proposition is right, all three are correct together. If it is wrong, all three are wrong together. We would rather the bet be legible than hedged.

These indicators will tell us whether we are wrong

Predictions are cheap without indicators. These are the specific things that would tell us early whether the five predictions above are holding.

  • For the jobs thesis: whether any major provider starts publishing per-task accuracy figures with remedies attached rather than per-token pricing. That would be the clearest signal the shape is shifting, and it would come from an incumbent rather than from us.
  • For provability: whether evidence requests start appearing in standard enterprise questionnaires rather than only in regulated sectors. When a mid-sized manufacturer asks to see an audit record for a specific inference, the moat has generalised.
  • For billing: whether large consultancies begin publishing fixed prices with comparisons. They have the most to lose and will move last, which makes them the reliable indicator.
  • For the freelance middle: whether the category we currently see growing fastest, correcting and replacing generated output, turns out to be permanent or transitional. We do not know, and it is the single number we watch most closely.
  • For portability: whether a large buyer rather than a platform forces the issue. An enterprise engaging contractors across several marketplaces has a real problem and leverage no individual has.

Writing daily exposed weak claims quickly

One observation from writing thirty of these is relevant to the argument.

Writing daily forced a discipline that is easy to describe and hard to practise: every claim needed a mechanism attached. It is straightforward to write that something is faster or better. It is considerably harder to write why, in a way a reader can check. The pieces that were difficult to finish were invariably the ones where the mechanism turned out to be weaker than the claim.

That is the same test we apply to systems. A number without a method behind it, an accuracy figure without a held-out set, or a commitment without a way to verify it. In each case the missing part is the part that makes it real, and its absence is usually the point.

We intend to come back and score this

The five predictions above are dated and specific enough to be assessed. We intend to revisit them, publicly, and say which ones we got wrong, because a firm that publishes predictions and never scores them is doing marketing.

Thank you for reading this month. The full set is on Insights, the method behind it on how we work, and if any of it prompted a question about something you are building, an engineer will answer.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded