The handover is the product

Handover should leave your team able to run and change the system without us. That means code, runbooks, decision records and rehearsal.

There is a moment on every project where the software works, you are pleased, and everyone starts talking about wrapping up. What happens in the following fortnight decides whether the system can survive without the people who built it.

Our view is blunt: the handover is the product. The running software shows that the handover is real. If the system works only while the original supplier is nearby, you have bought a dependency as well as an application.

A working system can still be impossible to maintain

We have inherited a number of systems that worked perfectly and could not be maintained. The same pattern appears each time. The original team knew things that were never written down. A person performed the deployment by hand, instead of a script doing it the same way every time. Configuration on the servers had drifted away from anything held in version control, the place where changes should be tracked. The choices that shaped the architecture lived only in shared memory among people who had moved on.

The software was fine. Your organisation’s ability to change it was zero. That turns working software into a liability with a working interface.

This is the failure handover is meant to prevent. The end state should be simple to describe: somebody who never met us can understand, operate and change the system.

What we hand over is the ability to run the system

Handover has to include the things a future engineer will need after the original project has faded from memory. We hand over the working software, and we also hand over the record of how to keep it working.

  • The source, with its history intact. The final state in a zip file is too little. The commit history, which records the sequence of changes, shows why things are as they are. It is worth more than most of the documentation.
  • Infrastructure as code. Every environment is defined in version control and rebuildable from that definition. A written description alone does not prove that it can be rebuilt.
  • A runbook written for a stranger. Deploy, roll back, restore, rotate a credential, add a user, investigate the three most likely failures.
  • Architecture decision records. Every significant choice, with the options we rejected and why. In three years this is worth more than the code it explains.
  • The tests, and what they cover. Including an honest note on what is not covered.
  • The evaluation set, where a model is involved, because without it nobody can tell whether a change made things better or worse.

Those items are practical, rather than ceremonial. They are the difference between being shown a system and being able to own it.

We test the handover on someone who was not on the project

Documentation written by the people who built something is systematically wrong in one specific way. It leaves out everything the authors find obvious, and what they find obvious is exactly what a newcomer needs.

So we test it. Somebody who was not on the project takes the documentation and performs a deployment and a common change, with no access to the team. Whatever they get stuck on is a defect. We log it and fix it like any other defect.

The first time we did this, the tester took four hours to get an environment running, against an estimate of forty minutes. None of the blockers were interesting. They were small things the original team no longer saw.

  • An environment variable mentioned nowhere.
  • A step that assumed a tool was installed.
  • A permission that had been granted by hand months earlier.

All trivial, all invisible to us, all fatal to somebody working alone on a Friday.

Decision records stop future engineers undoing important choices

Of everything in the handover list, architecture decision records are the item clients initially value least and later value most.

The reason is that most maintenance disasters begin with somebody changing something that looks unnecessary. A newcomer sees a cache, which stores data temporarily, with an oddly short expiry. They see a retry with a strange backoff, meaning the system waits in an unusual pattern before trying again. They see a denormalised column, which duplicates data. Each one can look like a mistake. Each was a considered response to a problem that will recur if it is removed.

A decision record is three paragraphs:

  • What we decided.
  • What else we considered.
  • What would have to change for this to be reconsidered.

That last section is the valuable one, because it tells a future engineer whether the constraint still applies. Without that record, the next person has to guess whether a strange-looking choice is accidental or load-bearing.

Handover needs its own time and budget

Treating handover as the last week guarantees that it will be compressed, because the last week is where slippage lands. We schedule it as a phase with its own duration and its own deliverables, typically the final three weeks. We do not compress it when earlier work runs late.

That occasionally means moving a launch date, which is an unpopular conversation. It is a better conversation than the one where you are live on a system nobody on your side can operate.

When handover has its own budget, the work is visible. The runbook, rebuild, decision records, drills and watched operations are deliverables, rather than afterthoughts squeezed between launch tasks.

Your team should run it while we watch

The strongest handover mechanism we have found is having your team perform the operations while we watch. For a defined period, your team runs the system, rather than watching us run it.

They do the deployment. They respond to the alert. They perform the restore. We are available, and we do not touch the keyboard. Everything that goes wrong in that period is enormously valuable, because it goes wrong while somebody who knows the answer is sitting there.

Ninety days is our usual period. By the end of it, your team has operated the system, which is a different state from having been shown how. They have dealt with the real sequence of tasks, with the real permissions, alerts and deployment process.

The hardest knowledge to transfer is judgement

Deployment steps and configuration are the easy part of a handover, because they are procedural and a runbook can hold them. Judgement transfers less easily. A future engineer needs to know which alerts matter, which parts of the codebase are fragile, and which apparent inefficiencies are load-bearing.

We try to capture that explicitly, instead of hoping it comes across in conversation. The runbook carries a section on the failures we consider most likely. For each one, it gives the symptom, the probable cause and the first thing to check.

That failure section covers the same practical chain a person follows during an incident:

  • The symptom.
  • The probable cause.
  • The first thing to check.

The decision records carry the rejected options. Those rejected options matter because they show what was considered and why it was left aside.

We also keep a short, honest document listing the parts of the system we are least happy with, why they are that way, and what we would do about them given time. That document is uncomfortable to write. Clients consistently say it is among the most useful things they receive.

It is the difference between inheriting a system with a map of its weak points and inheriting one where every weakness has to be discovered during an incident.

The first real incident proves whether handover worked

The acceptance meeting is a poor test of handover quality. The real test is the first genuine production incident after the supplier has gone. That usually happens somewhere between two and six months later, and usually at an inconvenient hour.

The outcome depends on whether the person on call can answer three questions without phoning anyone:

  • What changed recently.
  • What does normal look like for this metric.
  • How do I put it back the way it was.

If the deployment history is visible, the dashboards have baselines rather than only current values, and the rollback has been executed as a drill, the incident is twenty minutes. If those things are missing, it is a night.

So we rehearse it. Before handover completes, we break something in a controlled way and let your team resolve it, with us present and silent. What that exercise surfaces is rarely technical.

It often shows that the human and operational pieces were unclear. Nobody knew who had authority to roll back. The alert routed to a person who had left. The runbook assumed access somebody did not have. Those are handover defects, and it is far better to find them while we are still there.

Sometimes the team receiving the handover does not exist yet

A harder case, and a common one, is handing over to an organisation that has not yet hired the people who will maintain the system. There is nobody to sit beside for ninety days, and the real recipient is a job advert.

This changes what a good handover looks like. The technology choices matter more, because you need to hire against them. A stack with a small hiring pool becomes an active liability rather than a preference. The documentation has to assume no institutional context whatsoever.

In that case, we usually recommend a retained arrangement covering the first months of whoever is eventually hired. It is priced as a small monthly amount. The reason is that the new engineer will need somebody to ask while they learn the system.

We are explicit that this is a temporary arrangement with an end date. A support contract that quietly becomes permanent is the lock-in we said we would not create, arriving through a side door.

After go-live, fewer calls are the measure of success

The metric we track after go-live is support contacts per week, and the target is that it trends to zero.

This is an unusual thing for a supplier to optimise for, because contact is revenue. We track it because a call is evidence of a documentation defect. We also track it because the alternative business model, where you cannot operate without us, is one we decided not to have.

That metric makes handover quality measurable rather than assertable. If the calls keep coming, the handover did not finish the job.

The contract should make leaving possible

Everything above is undermined if the commercial arrangement makes leaving expensive, so we write the exit into the contract. The practical commitments are specific.

  • Full intellectual property in your repository.
  • Mainstream technology throughout, with nothing in the stack that requires us specifically.
  • No minimum term on retained work.
  • Thirty days notice, and we help with the transition.

The self-interested version of this argument is that a supplier who is difficult to replace has an incentive structure that produces worse work over time. The client-facing version is simpler: if you stay, we would like it to be because the work is good.

This costs real money

Real money. Three weeks of handover on a fourteen week project is a fifth of the engagement spent on something a competitor would do in two days. It is priced into the proposal rather than absorbed.

We win less work on price as a result and we are asked back more often, which is a trade we are content with. It also means we can take on maintenance of systems we did not build, because we know what a maintainable system looks like from having had to produce them.

You can ask any supplier for the same proof

You do not have to take a supplier’s handover claims on trust. Ask for evidence before the project ends, while the work can still be corrected.

  • Ask to see the runbook before the final payment, and have somebody outside the project try to use it.
  • Ask for a full environment rebuild from source, performed in front of you.
  • Ask for the decision records. If there are none, ask why the architecture is as it is and see how long the answer takes.
  • Ask what happens if you terminate in ninety days, and get the answer in writing.
  • Run one deployment yourselves, with the supplier watching, before you accept.

A supplier who has built for handover will welcome all five. One who has not will explain why each is unnecessary. The explanations will be reasonable, and you should insist anyway.

Build for the people who will inherit the system

Software outlives the relationship that produced it. The useful question is whether it can be understood, operated and changed by people who have never met the supplier, because every system eventually reaches that state.

Building for that from the first week costs perhaps fifteen percent more and it is the difference between an asset and a dependency.

More on the commitments behind this in how we work, and on the sector where it matters most in energy and utilities.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded