Every engineering team has a rollback story. Something shipped, something broke, someone ran the command, and the service came back. That is one of the steady comforts of conventional software work. The previous version still exists, so you can go back to it.
That comfort fades with probabilistic systems, where the system can keep running while giving answers that are subtly wrong. The hard part usually is not reverting the model itself. Reverting a model is trivial. The problem is everything the previous version did while it was running.
A model may have written to records, sent information to customers, supported decisions, or shown output to someone who acted on it. None of that has to produce an error message. The system can look healthy while the damage is spreading.
A model rollback has two jobs
Conventional software rollback usually assumes the fault was in the code, and the code stopped when you stopped it. Revert, restart, done. Any bad data written during the incident is usually visible, because the failure was visible.
Model rollback has to handle a different shape of problem. The system may have been working, in the sense that nothing errored, while producing outputs that were subtly wrong. Those outputs were consumed. They were written to records, sent to customers, used in decisions, and they did not announce themselves.
That leaves two questions. Teams often plan for the first and leave the second until the pressure is on:
- How do we stop it.
- What do we do about everything it already did.
The first question is mostly an engineering control. The second one becomes an operational, commercial and sometimes regulatory decision. A good rollback plan has to cover both.
Stopping the model should be fast and predictable
Stopping is the easy half. It still needs to be done properly, because plenty of systems cannot do even this quickly.
The system needs a known version to return to. Pin model versions explicitly rather than using a provider’s floating alias, where the provider can move the name to a newer underlying model. If you cannot name the exact version currently serving traffic, you cannot revert to a known state.
Prompts need the same treatment as code. Keep prompts in the repository under version control, deployed with the code. A prompt edited in a database by someone at 11pm is not revertible in any meaningful sense, because the team may have no clean path back to the exact previous behaviour.
The whole feature should sit behind a flag that can be disabled without a deploy. When the answer to how do we turn this off is deploy a revert, the mean time to stopping is thirty minutes rather than thirty seconds.
The plan also has to say what happens when the feature is disabled. The system might fall back to a previous version, to a rules-based path, or to a human queue. A feature flag that turns something off without saying what happens instead is an outage switch.
Those controls do not answer the harder question. They only stop more output from being produced. The next job is to work out where the earlier output went.
You need to know where the output went
Before you can decide what to do about past output, you have to know where it went. This part has to be designed in advance because it cannot be reconstructed afterwards with confidence.
Every output a model produced should be traceable. The record should show which model version produced it, when it was produced, what input it came from, and what happened to it downstream. It should show whether the output was written to a record, sent externally, used in a decision, or shown to someone who acted on it.
That trace is lineage, the trail from input to model output to later use. Without that trail, a rollback becomes an archaeology project. We have watched a team spend two weeks reconstructing which records had been touched by a misbehaving classifier, working backwards from timestamps and hoping nothing else wrote in the same window.
With lineage, the same work is a query. You ask the system for everything produced by the affected model version in the affected period, then decide what to do with the results. That difference is worth building for before you need it.
Past output falls into four consequence categories
What you can actually do depends on where the output went. The four categories have very different remedies.
- Internal and unread. Written to a field nobody has looked at. Correct it silently and move on. The easiest case and, encouragingly, the most common.
- Internal and acted upon. Someone made a decision using it. Correction is possible but people need telling, because their conclusion may have been wrong and they may have told someone else.
- External and reversible. Sent to a customer, but the effect can be undone. A wrong status, a wrong estimate, an incorrect categorisation. Requires a correction and a communication, and the communication is usually the harder part.
- External and irreversible. A payment made, a message sent, a document filed, a decision communicated to a regulator. Cannot be undone at all. This is why the mandate work matters so much: keeping actions out of this category is a design decision made long before any incident.
This category work matters because rollback is not a single action. The right response changes with the consequence. Some records can simply be fixed. Some people need to be told. Some outcomes have already left the building.
Fixing every record can create a worse problem
The instinct is to fix everything. That instinct is frequently wrong, and this judgement call should involve someone other than the engineers.
Consider a classifier that mislabelled a proportion of records over three weeks. You can identify and correct them. But if downstream systems consumed the wrong labels and produced their own outputs, correcting the source without addressing the derived data creates a system that is internally inconsistent. That can be worse than being consistently wrong, because different parts of the business now tell different stories.
Or consider outputs a person reviewed and accepted. Silently changing them undermines the review record. The record now shows someone approving something they never saw. In regulated work that is a worse finding than the original error.
So the options are correct, correct and notify, annotate without changing, or leave and document. Which one is right is a business and compliance question, and the incident is a bad time to be having it for the first time.
The engineering team can identify the affected output and explain the mechanics. The decision about correction, notification and documentation belongs with the people accountable for the business and compliance consequences.
A staged deployment limits how much you may have to unwind
The best rollback plan is a deployment approach that limits how much there is to roll back. The order matters, because each stage gives the system a chance to show disagreement before exposure grows.
- Shadow first. The new version runs against live traffic and its output is recorded and compared, but not used. Costs inference, meaning paid model calls, and reveals disagreements before anyone is affected.
- Then a small percentage, with the evaluation running against production samples rather than only the fixed set.
- Then a wider rollout, holding at each stage long enough for the slow signals to appear. Some problems only surface when a downstream weekly process runs, which means a same-day full rollout guarantees you find them at full exposure.
The cost is that changes take longer to reach everyone. The benefit is that when something is wrong, the blast radius is five percent of three days rather than everything for three weeks.
That trade is central to model rollback. A slower rollout does not remove the need for rollback, lineage or judgement. It makes the eventual incident smaller, clearer and easier to contain.
You only trust a rollback plan after a drill
A rollback plan that has never been executed is a document, and documents are wrong in ways nobody discovers until the worst moment.
We run it as a drill on every system we build. The drill follows the actions the team would need during a real incident:
- Disable the feature and confirm the fallback path actually works.
- Query the lineage for everything produced in the last hour and confirm the answer is complete.
- Execute a correction on a small set in production and confirm downstream systems behave.
That drill has found several kinds of problems that looked harmless until someone tried to run the plan.
- Fallback paths that had rotted because nobody exercised them.
- Lineage records missing the model version because a refactor dropped the field.
- Correction scripts that worked in staging and hit a foreign key constraint in production, where the database would not allow a change because another record depended on it.
Every one of those would have been discovered during a real incident, at the worst possible time, by someone under pressure.
A real incident shows what has to be ready
A classification system sorted inbound requests into eleven categories, routing each to a different team. It ran at around 94 percent accuracy for five months.
A provider updated the underlying model. Accuracy on the evaluation set fell to 88, which the scheduled run caught the following morning. But the drop was not uniform. Two categories that were rare in the evaluation set had collapsed to near-random, and those two categories mattered disproportionately because one of them was the escalation path for vulnerable customers.
The incident then moved through a sequence of decisions and checks.
- Hour one. Feature flag disabled. Fallback was routing everything to a single triage queue, which was slower but correct. This worked because it had been drilled two months earlier; the first drill had found the fallback broken.
- Hour two. Lineage query for everything classified since the model changed. Fourteen hours of traffic, about 900 items, of which 60 were in the two affected categories.
- Hour four. Categorising the consequences. Most had been routed and not yet worked, so a simple re-route. Eleven had been actioned by a team, meaning someone had responded on a wrong basis. Two had generated customer communications.
- Day one, afternoon. The judgement call. Silent correction for the unworked items. Direct notification to the teams for the eleven. For the two customer communications, the client’s compliance lead decided to contact both, which was the right call and a decision an engineer should not have made alone.
- Week one. Version pinned to the previous model. Evaluation set expanded so the two rare categories had proper representation, which was the actual root cause: the set had been unrepresentative and had been reporting a number nobody should have trusted for those categories.
Total exposure fourteen hours, because the scheduled evaluation ran daily. Without it, the first signal would have been a team noticing odd routing, and the honest estimate is two to three weeks.
The incident was survivable because the basics were already in place
Four things made this survivable, all decided before the incident and none of them heroic.
- The scheduled evaluation ran daily rather than only on deploy, which is what bounded the exposure to hours.
- The feature flag existed and the fallback had been tested, which is what made stopping a two-minute decision.
- Lineage recorded the model version against every output, which turned a reconstruction project into a query.
- The escalation path to a compliance decision-maker was known in advance, so the awkward question about customer notification had an owner rather than a committee.
None of those are expensive. All of them are the sort of thing that gets deferred because they produce no visible feature, and each one converted a potential multi-week problem into a single working day.
The important detail is the timing. These controls helped because they were already present when the model changed. Building them after the incident would only have helped the next time.
Settle the customer and compliance decision before launch
Before deployment, someone should answer this in writing: if this system has been producing wrong output for three weeks and it has reached customers, what do we do?
The technical remedy matters, and the answer also has to cover the commercial and regulatory one. Do we tell them. Do we correct. Who decides. What is the threshold at which this becomes a disclosure.
Almost nobody asks this before launch, and it takes twenty minutes. Asking it during an incident takes days, involves lawyers, and produces a worse answer because everyone is defending a position rather than choosing one.
It also tends to change the design. A team that has thought properly about the three-week scenario builds a narrower mandate, a smaller blast radius and better lineage than a team that has not, without anyone having to argue for those things separately.
The one thing we would change with hindsight is the evaluation set itself. It had been built from a representative sample of traffic, which sounds correct and meant the rare categories had a handful of cases each. A set weighted toward the categories with the highest consequence, rather than the highest volume, would have caught the problem in the first hour rather than the fourteenth. We now build sets that over-sample the cases that matter most rather than the ones that occur most.
More on limiting what a system is permitted to do in the piece on narrow mandates, and on how we structure delivery in how we work.








