Feature flags are the cheapest insurance you will ship

Feature flags separate deployment from release, so teams can turn risky behaviour off in seconds and investigate after the system is healthy.

Every release process reaches the same moment. Somebody asks whether we are confident.

The honest answer is always no, not completely. A system of any size has behaviour that only appears when real traffic meets real data. Tests, reviews, staging environments and deployment pipelines all reduce risk, but they do not remove that fact.

So the useful question is different. The useful question is how fast you can undo the change when production teaches you something.

Separate deploying code from showing it to users

The single most useful idea is to separate two events that often get treated as one.

  • Deployment is putting code on a server.
  • Release is making behaviour visible to users.

When those two events are bundled together, every release carries the full risk of a deployment. Every rollback also requires a redeployment, so the undo path has the same machinery, delay and stress as the original change.

A feature flag, meaning a switch in configuration that controls whether a piece of behaviour is active, changes that shape. Code can ship dark and dormant, days before anyone sees it. It has already been through the deployment pipeline. It has run alongside production traffic. The moment of turning it on becomes a configuration change rather than a code change.

The rollback then costs ten seconds and no build.

Fast rollback changes how people act under pressure

A rollback by redeployment takes as long as a deployment. For most organisations, that is between five and forty minutes. During an incident, that difference matters. It can be the difference between an event nobody outside the team notices and one that reaches a customer.

The cost of undoing a change also changes how engineers behave. When undoing something is expensive, engineers hesitate. They investigate. They try to understand the problem before acting, because reverting has a cost and might be unnecessary.

When undoing is free, the better response is available. Turn it off immediately. Then investigate afterwards, with the system healthy and no clock running.

That second posture produces much better outcomes. It is created entirely by the cost of the action.

Put flags around changes where failure would matter

A codebase where every branch is conditional becomes unreadable and untestable. The flags themselves then become a source of bugs, so flags need to be used where they are worth the extra branch.

Our rule of thumb is to flag these kinds of changes:

  • Anything that changes behaviour a user can observe.
  • Anything that touches money or inventory.
  • Anything that alters a data path.
  • Anything whose failure mode we cannot fully predict.

Refactoring with no behavioural change does not need one. A new algorithm for calculating a price certainly does.

Every temporary flag needs an expiry date

The failure mode of this technique is flag accumulation. Every flag is a branch. Every branch doubles the state space. A codebase with two hundred live flags has a combinatorial testing problem nobody is solving.

That is why every flag is created with an expiry. A suggestion is too weak. It needs a date.

When the date passes, only two things can happen:

  • The flag gets removed.
  • Somebody explicitly extends it with a reason.

We run a report of flags older than their expiry. We treat it as a backlog with the same seriousness as a bug list.

Removing a flag is a real task with real work. The work is small, but it has to be scheduled:

  • Delete the branch.
  • Delete the tests for the abandoned path.
  • Delete the configuration.

It takes twenty minutes and it never happens unless somebody schedules it.

Keep release flags and operational flags separate

There are two kinds of flag, and they have different jobs.

  • Release flags are temporary. They exist to decouple deployment from release, and they should be removed within weeks.
  • Operational flags are permanent. They exist to let you degrade the system deliberately, and they should never be removed.

The second category is the one that saves a peak trading day. These switches let the team keep the system working by turning off behaviour that can wait.

Operational flags include controls such as:

  • A switch that disables recommendations.
  • One that turns off a non-essential third party call.
  • One that puts the system into a reduced mode.

These are controls, and they belong in a documented list that an on-call engineer can find at three in the morning. Treating release flags and operational flags as the same thing is how a useful kill switch gets deleted during a tidy-up.

Test both paths, or the switch is unsafe

A flag creates two code paths. The test suite must exercise both. Otherwise, you have shipped an untested path and armed a switch that turns it on.

The practical approach is to run the suite twice for flags covering significant behaviour:

  • Once with the flag on.
  • Once with the flag off.

That is usually better than trying to construct a matrix across every combination. It catches the majority of problems at a manageable cost. It is also one of the reasons to keep the number of live flags small.

Use progressive exposure to compare groups

Turning a flag on for one percent of traffic, then ten, then fifty, is standard practice. It is frequently done for the wrong reason. People often think mainly about reducing the number of affected users. The stronger reason is that it gives you a comparison.

With one percent enabled, you have ninety nine percent as a control group at the same moment, under the same conditions, with the same data. If error rates or conversion differ between them, that difference is caused by the change. You know that with far more confidence than comparing today against yesterday.

So the exposure ramp should be paired with a comparison, rather than just a wait. Turning it on for one percent and watching an overall dashboard tells you almost nothing, because the overall dashboard is dominated by the ninety nine percent.

Database changes need a different discipline

Flags work beautifully for behaviour and poorly for schema, the structure of the database. You cannot flag a dropped column back into existence.

Anything that changes data structure needs the expand and contract discipline instead. In this case, the flag controls which structure is read. It does not decide whether the structure exists.

The general rule that follows is worth stating plainly: never combine a schema change and a behaviour change in the same release.

Ship the structure first, dormant. Ship the behaviour behind a flag afterwards. Each is separately reversible, and that is the property you are buying.

The flag system must fail safely

There is an irony to plan around. The mechanism that exists to protect you from outages is itself a dependency in the request path. If it fails, every request that consults a flag is affected simultaneously.

Three rules keep that from happening:

  • Every flag has a default compiled into the application, so that if the configuration store is unreachable the system behaves in a known way rather than throwing.
  • Values are cached locally with a generous stale window, so a brief outage of the store changes nothing at all.
  • Evaluation is local rather than a network call per request, because a synchronous lookup on every request adds latency in the best case and an outage in the worst.

The default value deserves thought. It should not be whatever the type initialises to.

For a release flag, the default is off, because the safe state is the old behaviour. For an operational kill switch, the default is on, because the safe state is the system working normally. Getting that backwards means a configuration outage silently disables your recommendations, or worse, silently enables a half-finished feature.

Target named accounts as well as traffic percentages

Targeting by percentage of traffic is the common case. Being able to enable something for a named list of accounts changes what the tool is for.

Account targeting lets you use flags in several useful ways:

  • Turn a feature on for your own staff first, in production, with real data, which finds a class of problem no staging environment will.
  • Enable something for the one customer who asked for it while you finish the general version.
  • Disable a feature for a single account that is hitting a bug, buying time to fix it properly rather than reverting for everybody.
  • During an incident, exclude the account generating the problematic traffic without affecting anyone else.

The cost is that the evaluation now depends on request context. That context has to be plumbed through to wherever the flag is checked. That plumbing is mildly annoying to add later and trivial to include at the start, which is a reason to build the mechanism before you think you need it.

A shipping calculation shows the difference

A client changed the way shipping was calculated. The change was well tested, reviewed, deployed behind a flag on a Tuesday, and enabled for five percent on Wednesday morning.

Within twenty minutes, the comparison showed the enabled cohort converting about two percent lower. There was no error. There was no exception. Nothing showed on any technical dashboard.

The new calculation was correct in every case the tests covered. It produced a slightly higher figure for multi-item orders to one region, because of a rounding rule that had been applied per order and was now applied per item.

The engineer who noticed turned it off before the stand-up. Total exposure was around forty minutes at five percent of traffic. The change was fixed and re-enabled two days later.

Under a deploy-and-revert model, that would have gone to a hundred percent of traffic. The symptom would have been a conversion dip nobody attributed to a deployment. It would plausibly have run for a week before somebody connected the two.

The flag did not prevent the bug. It reduced its cost by a factor of several hundred.

The build cost is small, and the discipline is the real work

A basic implementation has three parts:

  • A configuration store.
  • A client library with sensible caching.
  • An interface for changing values with an audit log of who changed what.

That is roughly a day of work, and there are mature hosted options if you would rather buy it.

The ongoing cost is the discipline. Expiries enforced. Both paths tested. The operational set documented. Perhaps an hour a month.

Against that, the last three production incidents we were involved in were resolved by turning something off, in under a minute, by whoever noticed rather than by whoever had deployment access.

A redeployment rollback is a repair story

A rollback that requires redeployment is a repair story with a longer name. If your response to an incident depends on somebody deciding whether the change was the cause, you have introduced a human judgement into the fastest path you own.

Making the undo free removes both problems. It is one of the very few things in software that costs a day and pays for itself the first time it is used.

More on the engineering commitments behind this in how we work, and on deployment practice in DevOps.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded