Ask most teams for their performance target and you will usually hear a shape of answer, not a number. It should feel fast. We aim for under a second. Nobody has complained recently.
Those answers sound reasonable until you need to test them. None of them can be tested. That means none of them can be defended when somebody proposes a feature that would breach them.
That is the actual problem. Performance works best as a product decision that engineering implements. It goes wrong when nobody makes that decision clearly enough to write it into the work.
Averages hide the customers who feel the system at its worst
The first correction is to stop looking at means. An average response time is dominated by the requests that were easy. The easy requests are rarely the ones costing you anything.
A system with an average of 300 milliseconds and a 99th percentile of nine seconds may look fine on a dashboard. For a user, it means one request in a hundred is unusable. Those requests are also rarely spread evenly across all customers.
They cluster on your largest customers, your longest histories, and your most complex orders. In other words, they cluster on the accounts you can least afford to lose.
So the target is always a percentile. A percentile tells you how the system behaves for a share of users or requests, rather than smoothing everything into one average. We usually set two targets: a 95th for the ordinary experience and a 99th for the worst acceptable case.
The gap between those two numbers matters. It is a design decision about how much variance you are willing to accept.
Set the target for the action the user is taking
Page-level metrics are too coarse to act on. Users do not experience pages as technical units. They experience actions.
Those actions include:
- opening a record
- adding a line
- saving
- searching
- exporting
Each action gets its own budget, because the expectations are different. A search should feel immediate. An export can take several seconds provided it says so. A save must be fast because it is on the critical path. Slow saves also teach people to batch their work and lose it.
Writing budgets per interaction makes the conversation with a product owner much easier. They can reason about opening a customer record takes under 600 milliseconds for 95 percent of customers. They cannot reason in the same way about a page weight.
Put the number inside the definition of done
If the number is missing from the scope beside the behaviour, it is an aspiration. Aspirations lose to deadlines every time.
We put performance budgets in the main body of the definition of done. We include the data volume they apply to, because a target without a volume is meaningless.
Six hundred milliseconds for a customer with eleven orders is a different engineering problem from six hundred milliseconds for a customer with forty thousand.
Once the number is written down, a feature that would breach it becomes a visible trade instead of a silent regression. Somebody has to decide, explicitly, that the new widget is worth two hundred milliseconds on the record page.
Sometimes it is. What matters is that a person made the decision.
Use test data with the same shape as real data
The most common reason performance work fails to help is that testing used data that does not resemble reality. A seeded test database with a thousand customers, each with a handful of records, will hide the problems that appear at scale.
Volume matters, and shape matters too. Shape means the distribution of sizes, the outliers, the customer with fifteen years of history, and the order with four hundred lines.
We build test data from the production distribution, including the tail. Then we run the budgets against the 95th percentile customer rather than the median one.
Measure first, then fix the dominant cost
Performance work done by intuition is wasted work. Engineers are consistently wrong about where time goes. That includes experienced engineers, and it includes us.
The order is always the same:
- measure
- find the dominant cost
- fix that
- measure again
Frequently the dominant cost is structurally boring rather than algorithmically interesting. The usual causes are ordinary things hiding in the path of the request.
They include:
- a query in a loop
- an index that does not exist
- serialising a payload three times
- a synchronous call to a service that did not need to be synchronous
We have never once found that the answer was to rewrite something in a faster language. We have several times found that a client had been told it was.
Most performance problems come from four causes
The causes we find most often are simple once they are visible. The hard part is identifying which one is actually dominating the time.
- Doing work per row that could be done per request. The classic query in a loop, in all its disguises, including the one hidden behind a lazy-loading ORM relationship.
- Fetching data nobody displays. Loading a full object graph to render three fields. Trivial to fix, enormously common.
- Waiting sequentially for things that could happen concurrently. Four independent calls taking four hundred milliseconds each, in sequence, when they could take four hundred total.
- Doing work at request time that could have been done earlier. Computing something on every request that changes daily.
Between them these account for the overwhelming majority of what we find. Once identified, none require deep expertise to fix. The expertise is in the identifying.
Guard the budget or the regression returns
Performance regresses. It regresses because every feature adds a little. No individual addition is large enough to notice. The person adding it is also testing on a small dataset on a fast machine.
The mechanism that works is a check in the pipeline. For the critical interactions, on representative data, assert the budget and fail the build when it is exceeded.
A report somebody reads is too weak. A failure gets attention.
Teams push back on this because a failing build for a performance regression feels disproportionate. It is proportionate, because the alternative is worse. The regression ships and is discovered six weeks later, mixed in with thirty other changes. At that point, finding it costs a day instead of a minute.
Some improvements come from reducing the wait, not the work
Some of the most valuable improvements are about avoiding unnecessary waiting. The underlying work may still happen, but the person using the system no longer sits through all of it.
Those product choices include:
- optimistic updates that show the result immediately and reconcile afterwards
- rendering the part of the page that is ready rather than waiting for everything
- doing slow work in the background with an honest indicator rather than a spinner that means nothing
- prefetching the thing the user is about to ask for
These are product decisions with engineering consequences. That is the theme running through all of this. They frequently deliver a larger improvement in how the system feels than any amount of query tuning.
Measure browser speed from the user side
Server response time is the part engineers measure because it is the part they own. For somebody using the system, server time is frequently a minority of the wait.
The rest of the wait comes from several places:
- network latency, which is not zero for anybody outside your office and is substantial on mobile
- the size of the JavaScript bundle, which has to be downloaded, parsed and executed before anything interactive happens
- fonts and images
- layout work the browser does
- the sequence of requests
The sequence matters because a page that needs four round trips before it can render is bounded by four times the latency, regardless of how fast the server is.
That means an interaction budget has to be measured from the user’s side rather than from the load balancer. We measure both. The difference between them tells you which half of the problem you have.
Teams that only measure server side reliably conclude that the system is fast while their users disagree.
Slow systems create data quality problems
A pattern is worth naming. When something is slow, people work around it. The workaround usually damages the data as well as the experience.
If saving is slow, people batch their work and enter it later from memory. The timestamps are wrong, and some of the work is missing.
If search is slow, people stop searching and rely on recall. They miss the record that already existed and create a duplicate.
If a report takes four minutes, somebody exports it once a month and works from a stale copy. Decisions get made on figures from three weeks ago.
Those problems rarely show up as performance incidents. They show up as data quality problems, duplicate records and decisions based on stale numbers. They are almost never traced back to latency.
This is the strongest argument we know for treating performance as a product requirement rather than an engineering preference. We have watched it play out in clinical settings, where the consequences are considerably worse than an annoyed user.
Cost belongs beside speed
The conversation usually stops at speed. There is a second dimension that matters increasingly: cost.
The same inefficiencies that make a system slow also make it expensive. At scale, the bill is a better forcing function than the stopwatch.
A query that runs once per row rather than once per request costs proportionally more compute and database capacity. An endpoint that fetches an object graph nobody displays pays for that in transfer and memory.
Work done at request time that could have been precomputed is paid for on every request forever. Under a metered model provider, a prompt that carries more context than it needs is money spent per call.
We put cost per transaction alongside latency in the reporting for systems where volume is significant. It changes which optimisations get prioritised. It also gives the work a business justification that survives contact with a roadmap discussion, which it feels sluggish does not.
We write these commitments into the proposal
When we commit to performance in a proposal, we make the commitment specific.
- Named interactions with 95th and 99th percentile budgets.
- The data volume and shape those budgets apply to.
- The measurement method, so nobody argues later about how it was tested.
- A pipeline check that fails on regression.
- What happens if a budget cannot be met, which is a conversation rather than a quiet acceptance.
Writing down the number gives performance an owner and a deadline
Nobody schedules performance work in advance because it has no deadline and no owner. Writing a number into the scope gives it both. It converts an eventual crisis into an ordinary requirement that gets built like everything else.
The systems we have had to rescue for performance reasons were never slow because the team lacked skill. They were slow because nobody had ever written down how fast was fast enough. Every individual decision was locally reasonable, and the aggregate was a system people had learned to work around.
More on how we specify work in engagement models, and on the delivery method in how we work.








