The migration nobody noticed

The safest migration has no cutover moment. Old and new run side by side until the new path is proven and the old one can be removed.

Ask anybody who has been in this industry for a decade about a migration weekend and you will get a story. The plan was for six hours. It took nineteen. Somebody discovered at four in the morning that a table had a foreign key, a rule linking one table to another, that nobody had accounted for. The rollback was attempted and turned out to fail, because it had never been tested against a partially migrated state.

That story is so common that people have started to expect migrations to look that way. They do not have to. The technique for avoiding that kind of weekend is well understood, and it is unglamorous.

The best migration is the one the client knows about, the team can observe, and users never feel. We get there by avoiding a single moment where the old thing stops and the new thing starts.

A single cutover moment creates most of the risk

Almost everything that goes wrong starts with the decision to have a moment. At time T the old system is authoritative. At time T plus one the new one is. Every condition that must be true has to become true inside a window that somebody has committed to.

That structure has no tolerance. If a step takes three times longer than expected, there is no slack. If something is discovered mid-way, the team has two bad choices. They can continue into unknown territory, or they can roll back a process that is only partly complete. Both choices are bad at four in the morning.

So we design migrations that avoid the cutover moment altogether. The work still happens, and it still has to be correct. The difference is that the system never depends on two halves of a change becoming true at the same instant.

Break the change into expand, migrate and contract

The old pattern that supports this is expand, migrate, contract. It is the foundation of everything else. Any schema change, meaning any change to the shape of the stored data, is broken into three deployments that are separately safe.

The three deployments happen in this order:

  1. Expand: add the new structure alongside the old. Nothing reads it yet. This deployment is trivially safe because nothing depends on the new thing.
  2. Migrate: write to both, backfill the historical data in batches, and begin reading from the new structure once it is verified complete.
  3. Contract: once nothing reads the old structure, remove it. Usually weeks later, deliberately, when everybody has forgotten it was ever exciting.

During the migrate phase, each part is separate. Writing to both, backfilling historical data, and reading from the new structure can each be controlled and reversed on its own.

That separation removes the sharp edge. At no point does the system require both halves of a change to become true simultaneously.

Backfill old data in small batches

The single most common cause of a migration incident is a single enormous statement that locks a table for longer than anybody expected. The system is technically doing what it was told to do, but the business experiences it as an outage or a serious slowdown.

We backfill in batches sized so that each one completes in well under a second. We pause between them, and the job runs continuously in the background for as long as it takes.

A table with four hundred million rows might take three days to backfill. That is entirely acceptable, because nothing is waiting for it and the system is serving traffic normally throughout.

The batch job needs three properties:

  • Resumable, because it will be interrupted.
  • Idempotent, because it will be re-run over rows it has already done.
  • Observable, because somebody needs to know it is progressing without reading logs.

Idempotent sounds technical, but the idea is plain: running the same part again over rows already handled must leave the data in the same correct state.

Dual writing works only if divergence is caught

During the middle phase, the application writes to both old and new structures. The trap is what happens when one write succeeds and the other fails.

If both writes are in a single database transaction, they succeed or fail together, and this part is straightforward. If the new structure is in a different store, the two writes lose that shared success or failure. This is where migrations quietly corrupt data.

A failure between the two writes leaves the copies diverged. Each system can still look fine on its own, so nobody sees the problem by checking either side in isolation.

The answer is a reconciliation job comparing old and new continuously and reporting differences. It runs for the entire duration of the dual write period.

A one-off check before switching over misses the problem. Divergence is produced by transient failures that occur at random, and those failures would not necessarily be present during a scheduled check.

Use real traffic to prove the new reads

Before switching reads over, there is an intermediate state that is extremely useful. Read from the new structure. Also read from the old. Compare them. Serve the old one. Log any difference.

This runs in production, on real traffic, with no user-visible effect. It answers the question that staging cannot answer, no matter how much testing happens there. The question is whether the new structure returns the same thing as the old one for the actual distribution of data your users generate.

This costs a little latency and some compute for a period. Every time we have done it, it has found something. The something was always in a data shape that did not exist in the test fixtures.

Test rollback against a half-finished migration

A rollback plan that has never been executed is a hypothesis. We execute ours, in production-like conditions, against a partially migrated state, before the migration begins.

The partial state is the important detail. Most rollback plans are written for the case where the change is complete and needs undoing. The dangerous case is where the change is half done. That is exactly the case nobody rehearses.

Testing rollback this way proves the plan against the state most likely to be painful. It also avoids learning, during the incident, that the rollback only worked on paper.

Some old records will fail to move cleanly

Every migration of an old system encounters records that cannot be moved cleanly. The records usually expose old assumptions that stopped being true years ago.

The common cases include:

  • Dates that are not dates.
  • Foreign keys pointing at rows that were deleted years ago.
  • Text in a field that was supposed to be numeric.
  • Duplicates that should have been impossible.

The wrong response is to fix them silently with a best guess. The right one is to quarantine them, count them, and put the count in front of the business with examples.

Frequently the answer is that a category of records from before some date does not matter and can be dropped. That is a decision the client is entitled to make explicitly, rather than have made for them by a coercion rule in a script.

Moving to another database adds another class of problems

Everything above holds when the change is within one database. It gets harder when the destination is a different engine, because the dual write is no longer transactional and the two stores disagree about what a value even is.

The differences that bite are rarely the ones in the migration guide. They show up in how the two databases compare, sort, round and represent absence.

These are the differences that matter:

  • Text comparison and sorting behave differently, so a uniqueness constraint that held in one store may be violated in the other.
  • Timestamp precision differs, so records that were distinguishable become identical.
  • Numeric types round differently, which matters enormously if any of the numbers are money.
  • Empty string and null are the same thing in some engines and emphatically separate in others, which will silently change the behaviour of every query that tests for absence.

We find these by comparing rather than by reading documentation. The reconciliation job that runs during dual writing is what surfaces them.

That comparison should be semantic rather than byte for byte. In other words, it should compare what the data means, not only the exact stored representation. Otherwise the report fills with differences that do not matter, and nobody reads it.

Old structures usually have hidden dependants

A migration plan usually accounts for the application. What it misses is everything else that has quietly taken a dependency on the old structure over the years.

Those dependencies can be all over the organisation:

  • A reporting tool connected directly to the database, built by somebody in finance, which nobody in engineering knows exists.
  • A nightly export to a partner.
  • Scheduled jobs, some of them on a server that predates the current infrastructure.
  • Saved queries in a business intelligence tool.
  • A materialised view.
  • On one occasion we encountered, a spreadsheet with a live connection that a director used every Monday.

The contract phase is where these surface, painfully, because dropping the old structure is the moment they break. So we look before we drop.

We check query logs for anything touching the old tables, sorted by source, over a period long enough to catch a monthly job. Then we contact the owners.

The list is always longer than expected. This exercise is the reason we leave weeks between switching reads and dropping anything.

The business keeps changing data while the migration runs

One more constraint deserves stating, because it is the reason big-bang cutovers keep being proposed. The business does not stop while you migrate. New records arrive during the backfill, and some of them will be edits to records the backfill has already copied.

Dual writing handles the new ones. The edits are the subtle case.

A row backfilled on Monday, modified on Tuesday through a path that writes to both, is fine. A row modified by a batch job that was never updated to dual write will be silently stale in the new structure with nothing to indicate it.

This is why the reconciliation job compares continuously rather than once at the end. It is also why it compares a moving window of recently modified records more frequently than it sweeps the whole table.

Finding a divergence within an hour of it occurring means you can still work out what caused it. Finding it three weeks later means reading a lot of logs.

People still need to know the migration is running

A migration with no user-visible moment still deserves communication. Otherwise an unrelated incident during the migration period may get attributed to it. Worse, a real migration problem may get dismissed as coincidence.

We publish a simple timeline internally. It says when dual writing starts, when backfill is expected to complete, when reads switch, and when the old structure will be removed.

The timeline contains the same basic shape every time:

  1. Dual writing starts on this date.
  2. Backfill expected to complete around then.
  3. Reads switch on that date.
  4. Old structure removed a fortnight later.

It costs nothing. It also means that when somebody sees something odd, they know who to ask.

A good migration is boring to everyone using the system

The last large migration we ran moved about two hundred million records between two quite different structures. The expand deployment was on a Tuesday. Backfill ran for nine days.

Verified reads ran for a further two weeks and found two data shapes we had not anticipated. Both were fixed without drama. Reads switched on a Thursday morning during business hours. The old structure was dropped nineteen days later.

There was no weekend. Nobody was on call. The client’s team were aware it was happening and experienced nothing, which was the objective.

This approach takes more work and removes a category of risk

This way of migrating is more total work than a big-bang cutover. It is spread over weeks rather than concentrated in a weekend, which makes it harder to plan around and harder to declare finished.

It also requires the application to tolerate an intermediate state. That means writing code that handles both structures for a period, and engineers dislike that because it is temporary complexity.

The exchange is clear. Temporary complexity removes a permanent category of risk. That is a trade worth making every time.

More on how we approach change in how we work, and on the related discipline in DevOps.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded