Legacy modernisation is now an AI problem because the hardest first job has changed. A system runs the business, nobody fully understands it, the people who wrote it have retired, and replacing it means first working out what it does.
That first job is the archaeology. Historically, it has consumed a third of the budget and most of the calendar. Before a team could safely replace anything, it had to read enough of the old system to know what mattered.
That is the part that has changed most. The change is larger than anything AI has done to writing new code. Discovery that took six weeks takes four days. The result is bigger than a lower bill. Projects that were previously not worth attempting are now viable.
Old systems hide the answers in the code
Understanding an old system means answering questions that only the code can answer. The problem is that the code is not organised around the questions a modernisation team needs to ask.
The questions are usually specific and scattered across the system:
- What are all the paths that write to this table.
- Which of these sixty stored procedures are actually called.
- What happens on the branch that only executes in February.
- Which of these config flags is still live.
- Where does this magic number come from.
Each question is answerable by a person reading carefully. Each one takes hours. There are hundreds of them, and you do not know which ones matter until you have answered enough to see the shape of the system.
The traditional approach is familiar. Interview whoever remains, read whatever documentation exists, trace the main paths, and accept that the rest will surface during implementation. That acceptance is where legacy projects overrun.
AI has changed the discovery work
A model can read the entire codebase. It can read across all of it at once and answer questions that used to require slow manual tracing.
On a recent engagement, the system had roughly 400,000 lines of a language nobody on the team wrote professionally. Within three days, we had a working view of the system:
- A map of every entry point.
- A call graph, meaning a map of which parts of the code call which other parts.
- A list of database objects with the code paths that touch each.
- An inventory of external integrations.
- A ranked list of the twenty most complex procedures by branching depth.
A team of three would have taken five to six weeks to produce that. It would also have been less complete.
This matters because the discovery phase used to decide whether the business case survived. When understanding the system gets cheaper, the whole case can change.
The model gives you a map, then you verify it
This is where the honesty has to come in. The map is not the territory, and a confidently wrong map is worse than no map.
Everything the model produces is a hypothesis. It is a well-informed hypothesis, and it is wrong often enough to matter. Verification still matters. The saving comes because verifying a claim is dramatically cheaper than discovering it from nothing.
Our verification is layered, because different claims need different checks:
- Anything structural, such as whether this function calls that one, is checked mechanically with a parser, a tool that reads the code structure.
- Anything behavioural, such as whether this branch handles the annual reconciliation, is checked by finding the code and reading it.
- Anything historical, such as whether this was added for a client who left, is checked against version control or a person, and is treated as unreliable until then.
The rule we work to is simple. A model may tell you where to look. It may not tell you what is true.
The model is strongest where manual reading is slowest
Four uses stand out. They are often different from what people expect.
- Dead code.
- Implicit business rules.
- Cross-cutting inconsistency.
- Translating unfamiliar languages.
Dead code
The model is good at identifying paths that cannot execute, given the current configuration and callers. This is tedious for a person, fast for a model, and mechanically checkable.
On the engagement above, roughly 30 percent of the codebase turned out to be unreachable. That is 30 percent nobody has to migrate.
Implicit business rules
The model is also good at finding conditions buried in code that nobody wrote down. These include discounts that apply below a threshold, exceptions for one customer, or a rounding rule that exists because of a decision in 2011.
These rules are what modernisation projects fail on. They are invisible until the replacement gets them wrong.
Cross-cutting inconsistency
Old systems often implement the same concept three different ways in three places, with subtly different edge-case behaviour. That is extremely hard to spot by reading sequentially.
It becomes much more straightforward when the whole codebase can be considered at once.
Translating unfamiliar languages
A model can explain what a piece of COBOL, PL/SQL or VB6 does to engineers who have never used it. We use it to make the old code legible enough that decisions can be made about it.
The model is weakest where the code cannot know the answer
The list of weak spots is shorter, but sharper.
- Why.
- What is load-bearing.
- Runtime reality.
Why
A model can tell you what the code does and will confidently invent a reason it does it. The reason is usually a plausible engineering rationale. The truth is usually a commercial decision from a decade ago.
Treat every explanation of intent as fiction until confirmed.
What is load-bearing
Which parts genuinely matter to the business is not in the code. It is in who complains when something breaks, and that is a conversation with people.
Runtime reality
Static analysis, meaning analysis of the code without running it, cannot tell you that a nightly job usually finishes at 3am but occasionally runs until 9 and blocks the morning batch. Only production tells you that.
The hidden rules are where projects fail
The implicit business rule category deserves examples, because it is where modernisation projects actually fail. The examples are always slightly absurd.
One system had a discount that applied to orders placed between 2pm and 4pm. It turned out to have been a promotion in 2013 that was never removed and had been quietly applying ever since. Nobody had noticed because it was small and the reports aggregated it away.
Another had a rounding rule that rounded down on one code path and to nearest on another. It produced a discrepancy of a few pence per transaction that a reconciliation process silently absorbed into a suspense account. The account had grown for eleven years.
One exception branch skipped a validation step for a single customer identifier, hard-coded. That customer had been acquired by a competitor four years earlier and the branch was still there.
Another date comparison behaved differently in February, because someone had implemented month-end using a fixed day count. It produced a wrong result on two or three days a year, and the resulting queries were handled manually by a team who assumed it was normal.
Each of these was discovered in the first fortnight rather than during a parallel run. Each one generated a business decision rather than an engineering one.
That is the value. The analysis found the rules early enough for someone to decide deliberately whether to carry the behaviour forward.
The first two weeks now start with a verified map
The shape of the engagement has changed enough to be worth setting out. It now moves through a clear sequence.
- Days one to three. Ingest everything: source, database schema, job schedules, configuration, whatever documentation exists. Produce the structural map and verify it mechanically. Nobody is interviewed yet, deliberately, because arriving with a map produces far better conversations than arriving with questions.
- Days four to six. Produce the inventory of implicit rules, ranked by how surprising they are. This is the document that goes to the business, and it is usually the moment the engagement becomes real for them, because it contains things they did not know about their own operation.
- Days seven to nine. Run interviews, now specific. The conversation moves from open requests such as tell us how this works to questions such as this branch appears to apply a different tax treatment for one region, do you know why. People answer that kind of question well and cannot answer the open version at all.
- Day ten onward. Build characterisation tests against the behaviour that has been confirmed, starting with the highest-consequence paths. From here it is ordinary careful engineering, and the archaeology is behind you rather than ahead.
Two weeks to a verified map and a test suite. That used to be the point at which discovery was roughly half done.
Discovery can now keep improving through the project
The traditional order was discover, then plan, then build. Discovery was front-loaded and expensive enough that everyone wanted to start building.
Because discovery is now cheap, it can be continuous. We map the system in the first week, and then re-run the analysis as understanding improves and questions sharpen.
The map gets better throughout the project. It no longer has to be a fixed artefact produced at the start and increasingly wrong.
It also changes what you can promise. We now quote fixed prices on modernisation work we would previously have insisted on doing time and materials. The main source of estimate risk was the unknown, and the unknown got smaller.
The safe modernisation method stays the same
The method around the discovery work has not changed at all. This is the part that would be a mistake to skip because discovery got faster.
The safe method still includes the same controls:
- Characterisation tests before any replacement, capturing what the system currently does including behaviour nobody intended.
- Replacement in slices, each independently reversible.
- Never a rewrite with a launch date.
- Running old and new in parallel and comparing output on real traffic before switching.
- A rollback that has been executed as a drill rather than assumed.
Faster understanding leaves a big-bang rewrite unsafe. The gain is that the safe approach becomes cheaper, which is a better outcome and a less exciting one.
Some shelved modernisation projects now make commercial sense
The commercial consequence matters more than the technical one.
There is a large population of systems where modernisation was never attempted because the business case failed on discovery cost alone. These are often twenty-year-old systems running a genuine but not enormous part of an operation, where six weeks of archaeology before any value could not be justified.
Those projects now work. We have taken on several that the client had quoted elsewhere years earlier and shelved. The system did not get easier. Understanding it got cheaper, and that was the term the case was failing on.
This is most visible in sectors with the oldest systems: utilities, manufacturing and telecommunications, where equipment and software both outlast the people who specified them.
A supplier should be able to explain the method
If someone proposes a modernisation project on this basis, three questions separate a real method from a claim.
- How do you verify what the analysis tells you?
- What will you do about the rules you find that nobody knew existed?
- How is the replacement sequenced, and what is reversible at each stage?
An answer that treats structural facts, behavioural claims and historical intent as the same thing has not thought about verification properly.
The right answer on hidden rules involves surfacing them to the business for a decision. Silent reimplementation misses the decision the business has to make.
If the answer on sequencing is a rewrite with a cutover date, the discovery speed has not made anything safer.
More on how we approach this work on custom software development, and on the standard we hold on every engagement in how we work.








