Somebody wants to replace a system. It is fifteen years old, nobody enjoys working on it, the original team has gone, and there is a proposal on the table to rewrite it properly.
That proposal contains an estimate. The estimate is wrong, and the estimator’s competence is not the reason. The estimate is wrong because nobody knows what the system does.
People usually know what the system is for. That is different from knowing exactly how it behaves. The behaviour that has to be preserved is spread across a decade of small changes. Each change answered a real problem at the time, and almost none of those decisions were written down.
Test the old system before replacement work starts
We do not replace a system we cannot characterise. Before any replacement work begins, we cover the existing behaviour with tests. That includes behaviour nobody documented and some behaviour nobody intended.
Those tests run against the old system first. That is the crucial part. The tests record what the old system does, rather than what people believe the new system should do. The gap between those two things is where every rewrite disaster lives.
Characterisation records reality, even when reality is wrong
The distinction matters. A verification test checks that behaviour is correct. A characterisation test records current behaviour, whether it is correct or wrong.
You write these tests by calling the system and recording what comes back. Then you assert that the same response keeps coming back. There is no judgement involved at the point of writing, which feels wrong to engineers who have been taught that a test encodes intent. Here the intent is to detect change.
Some of what you capture will be bugs. You capture them anyway, and you mark them, because the decision about whether to preserve a bug is a business decision. Frequently the answer is yes.
Downstream systems have adapted. Users have workflows built on it. A rounding error that has been consistent for eleven years is now the definition of correct as far as the annual accounts are concerned.
The first surprise is usually how much nobody knew
Every time we have done this, the exercise has surfaced behaviour that nobody in the organisation knew about.
The surprises are usually ordinary business rules and old decisions that became invisible. They look like this:
- Discount logic that applies differently in one region because of a change made for a customer who left in 2018.
- A validation rule that rejects a legitimate input pattern, which the sales team learned to work around so thoroughly that they had forgotten it was a workaround.
- Two code paths that produce different results for the same input depending on which entry point is used.
- A scheduled job whose output nothing consumes, running nightly for six years.
None of that is in any document. All of it would have been discovered after the replacement went live, one angry phone call at a time.
Put the tests where the business depends on the system
The instinct is to test at the unit level, and it is the wrong instinct here. Unit tests are coupled to structure, and the structure is exactly what you are about to throw away.
A thousand unit tests against classes that will not exist in the new system are a thousand tests you will delete. They may say a lot about the old design, but they will say little about whether the replacement keeps the business behaviour intact.
Characterisation tests belong at the boundary that will survive. Those boundaries are the things the rest of the organisation depends on:
- The API.
- The message contract.
- The file format.
- The database state after an operation.
Those are the things that have to remain true. The practical version is usually to record real inputs and outputs from production traffic, sanitise them, and replay them against both systems.
Replace the system in slices, then compare every difference
With characterisation in place, the replacement stops being a single event and becomes a sequence. Route a fraction of traffic to the new implementation, compare the outputs against the old one, and investigate every difference.
Running both systems and comparing the results is the technique that makes large replacements survivable. It is more work than switching over. It also means that on the day you finally retire the old system, you have already been running the new one against real traffic for weeks with the differences reconciled.
Each difference is investigated rather than assumed. Sometimes the new system is right and the old one had a bug. Sometimes the reverse is true. The investigation is the value.
The cost objection misses the cheaper risk
The objection is cost. Characterising a large legacy system can take weeks, and it produces no user-visible progress. Somebody senior will ask why the team has spent a month writing tests for code that is being deleted.
The answer is that the tests serve the new code. They are the specification for the new code, obtained by measurement rather than by archaeology, and there is no cheaper way to get one.
The alternative is to write the new system against an assumed specification and discover the differences in production. That alternative defers the cost and makes it more expensive.
Use the coarsest useful boundary when isolation is impossible
Occasionally the existing system is impossible to exercise in isolation. It writes to a mainframe, or it requires a device, or it has no interface other than a user interface that cannot be automated.
Then the answer is to characterise at whatever boundary is available, even if it is coarse. The available boundary may be one of these:
- Database state before and after a business process.
- A daily output file.
- A log.
Coarse characterisation is much worse than fine characterisation and enormously better than none. It usually still finds several surprises.
Model-assisted rewrites need the same rule even more
Legacy modernisation is increasingly done with model assistance, and the same rule applies with more force. A model can read a large codebase and produce a plausible description of what it does.
That description will be substantially correct and will contain confident errors. By inspection, the errors are indistinguishable from the correct parts.
Characterisation tests are the mechanism that turns a plausible description into a verified one. Use the model to generate candidate tests quickly, absolutely. Then run them against the old system and treat anything that fails as a question about the model’s understanding rather than a bug in the legacy code.
A worked example shows the numbers clearly
A distributor asked us to replace an order processing system written in 2009. Their previous supplier had estimated sixteen weeks against a specification assembled from interviews and the original design documents.
We spent five weeks characterising instead. The suite ended up with about nine hundred recorded cases, replayed from a year of production traffic, asserting on the resulting database state and the outbound messages.
It found thirty one behaviours nobody had documented. Six were bugs the business had adapted to and wanted preserved. Four were bugs nobody knew about that were quietly costing money, the largest being a tax rounding rule applied inconsistently between two order entry paths, worth a few thousand a year in the wrong direction. Two were features that had been requested, built, and never used by anyone.
The build then took nineteen weeks rather than sixteen, so the total was longer than the original estimate by about half. What did not happen was a single production incident at cutover, because the new system had been running in parallel against real traffic for six weeks with every difference reconciled.
The client’s own assessment was that the previous estimate had been for a different and easier project than the one that actually existed.
Parallel running has costs you should plan for
Running two systems against the same traffic sounds straightforward, and there are three practical problems worth planning for.
Side effects need to be disabled or redirected
The first is side effects. If both systems send emails, customers get two. If both write to the ledger, you have double counted.
So the shadow system runs with its outbound effects disabled or redirected. That means the thing you are testing differs slightly from the thing you will run, and the difference has to be small and understood.
Comparison noise can bury the useful differences
The second is comparison noise. Timestamps differ. Generated identifiers differ. Ordering within a collection may differ without meaning anything.
Without a comparison function that normalises those, you drown in false differences and stop reading the report, which is worse than having no report. Building that normaliser is a day of work, and it is the difference between the technique working and being abandoned in week two.
The old system needs an agreed end date
The third is cost, literally. You are running two systems and storing both sets of outputs. For a few weeks that is fine.
Agree the end date in advance, because parallel running has a way of continuing indefinitely once nobody is willing to be the person who turns the old system off.
Stop when the recorded cases cover real use
The characterisation phase can expand forever, because there is always another edge case. We use coverage of production traffic rather than coverage of code as the stopping rule.
When the recorded cases exercise the paths that account for the overwhelming majority of real usage, plus every path that touches money, that is enough.
The paths that remain uncovered are usually genuinely rare. For those paths, log loudly when they are hit in the new system, so that the first occurrence after cutover is a notification rather than a silent divergence.
The client should receive the evidence, not just the replacement
At the end of this work, we hand the client these artefacts:
- The characterisation suite, running against the old system, green.
- A list of behaviours we found that nobody knew about, with a decision recorded against each: preserve, fix, or drop.
- The comparison harness used during the slice-by-slice replacement.
- The differences found during parallel running, and their resolutions.
That list of undocumented behaviours is frequently the most valuable artefact of the whole project. It is often the first time anybody has written down what the system actually does.
A rewrite is a bet on what you know
A rewrite is a bet that you understand the current system well enough to replace it. Characterisation tests are how you find out whether that bet is safe, before you have placed it, for a cost measured in weeks rather than in a failed migration.
Every rewrite disaster we have been asked to rescue had the same root cause: the new system was correct according to a specification that did not match reality. The new system being badly built was never the cause. Nobody discovered the mismatch until the old system had been switched off.
More on how we approach this in how we work, and on the wider argument in custom software development.








