Two structural choices shape how this company works more than anything else we have written down. We do not run a contractor bench, and we do not force a distribution on performance reviews.
Each choice is common enough to understand on its own. Together, they rule out the operating model most firms of our size use. They also create consequences we live with every week.
A bench smooths demand by moving risk onto workers
A bench exists to solve a real problem. Client demand is lumpy. A project ends in March and the next one starts in May, and in between you have people with nothing to do.
A firm has two choices in that gap. It can carry those people, which costs money. Or it can avoid employing them until they are needed, which is what a bench of contractors achieves.
That answer is sensible, and the difficulty it solves is real. We understand why the model is common. The business problem is uneven demand, and the bench makes that problem easier for the firm to absorb.
The cost lands somewhere else. The volatility of client demand moves onto the people doing the work. The firm’s revenue becomes smoother, and the individual’s income becomes lumpy. The person absorbing that risk is the one least able to diversify against it.
We hire permanent staff and carry the gaps ourselves
Everyone here is permanent and salaried. When there is a gap between projects, we carry it ourselves.
That decision affects how we grow. We hire more slowly than demand grows. We sometimes turn down work because we do not have the people, instead of assembling a team from a network in a fortnight.
It also changes how a quiet month feels inside the company. The cost of that month lands on us. That focuses the mind on pipeline in a way that a variable cost base does not.
The gaps have produced work we now depend on
The interesting consequence is what happens in the weeks between projects. That time exists, and it has to be used for something.
Nearly everything we have that is worth having came out of it. These were built by people who had three weeks and a thing that had been annoying them:
- The internal tooling.
- The characterisation test harness we use on legacy work.
- The evaluation framework we now use on every model project.
- Two products, one of which is live.
None of those were funded as initiatives. They were built because the time existed and the problem was close enough for someone to fix.
A firm with a bench does not get that outcome, because the person with three spare weeks is not employed during them. This is the strongest argument for the model, and it took us two years to notice it.
Clients get continuity when the same people stay
The other benefit is the one clients care about most. The people who built your system are still here when you call about it eighteen months later.
Under a bench model, the team that delivered your project has dispersed by the time you need them. The knowledge that was in their heads has gone with them. There is always knowledge in heads, regardless of how good the documentation is.
We take documentation seriously precisely because we know that. Even so, documentation is still an incomplete substitute for the person who remembers why a decision was made.
We do not stack-rank people at review time
The second choice is about performance reviews. We do not stack-rank people. There is no forced distribution, no requirement that a proportion of people be rated below expectations, and no quota of top ratings.
The theory behind forced distribution is familiar. Managers are too generous, a curve corrects for that, and identifying the bottom decile drives improvement. On a large population, there is an argument for the first two. The third part is where the model goes wrong.
Forced ranking makes a colleague’s success feel like your loss
On a team of our size, forced ranking is actively destructive for a reason that is easy to state. It makes your colleague’s success your loss.
If a fixed number of top ratings exist, helping somebody else do excellent work reduces your own expected outcome. Nobody says this out loud, and everybody responds to it.
The behaviours are subtle, but they are visible. The model creates:
- Less willingness to spend an afternoon unblocking somebody.
- Less appetite for the unglamorous work that keeps a codebase healthy and shows up on nobody’s individual review.
- A quiet preference for projects with visible outcomes over projects that matter.
We depend heavily on people doing exactly the things that ranking discourages. These are ordinary parts of the work, and they matter:
- Reviewing carefully.
- Writing the decision record.
- Attacking a colleague’s work in red-team week and finding something.
Every one of those acts improves somebody else’s output at the cost of your own visible throughput.
Reviews are written against the level, not against colleagues
Reviews happen twice a year. They are written, and they are against expectations for the level rather than against other people.
The question is whether somebody is doing the job at the standard we expect. If they are not, the review has to say specifically what is different and what support is available.
Compensation is reviewed against market twice a year and adjusted. It is not used as a lever attached to a rating. We do this because pay drifting behind market is the most common reason good people leave, and because tying pay to a rating recreates the ranking dynamic through a side door.
Performance problems need direct conversations earlier than review time
The obvious objection is how to deal with somebody who is not performing when there is no curve.
We deal with that directly, and much earlier than a review cycle. If somebody’s work is not at the standard, that is a conversation in the week it becomes apparent. We do not wait to turn it into a rating in six months.
The conversation is specific. It comes with support. If the work does not improve, it ends in a departure.
We have had those departures, and they are unpleasant. What a forced distribution offers is a way of making the decision feel procedural rather than personal. It converts a difficult conversation into an arithmetic exercise, and the manager gets to say the system required it.
That is a comfort for the manager. It is worse for everybody else, including the person leaving, who deserves a straight answer rather than a curve.
We avoid easy engineering metrics because they measure the wrong thing
The absence of a curve creates an obvious temptation. If people are not ranked against each other, surely they can be measured against numbers. The numbers are already in the tooling.
We do not use those numbers for assessment, because every available number measures something adjacent to the work rather than the work. The common numbers reward volume and visible motion:
- Commits.
- Pull requests.
- Lines changed.
- Tickets closed.
- Story points completed.
None of them capture the afternoon somebody spent deleting code. They miss the design conversation that prevented three weeks of building the wrong thing. They also miss the review comment that caught an authorisation gap.
The specific damage is that measured people optimise for the measurement, and they do it honestly rather than cynically. An engineer who knows tickets closed is counted will pick up the small ones. That is a rational response to a stated priority, and the fault is entirely with whoever chose the metric.
So assessment is a written judgement by people who have seen the work. That is subjective and slower and occasionally wrong. We prefer an honest subjective assessment to a precise measurement of the wrong thing, and we say so in the review document.
We gave up utilisation because it stopped matching how we work
The number a professional services firm normally runs on is utilisation, meaning the proportion of available hours that are billable. It is easy to compute. It correlates with revenue under an hourly model, and it is the metric the bench exists to protect.
Once we moved to fixed price against a written scope, utilisation stopped meaning anything. Nobody is paying us for hours, so an hour that is not billable is not lost revenue.
What matters is whether the scope shipped and whether the client would sign again. Those are the two things we actually track.
Giving up utilisation is harder than it sounds, because it is the number that makes a services business legible to itself. Without it, a quiet week looks like a quiet week rather than a costed deviation. Somebody has to be comfortable with that.
The compensation is visible in the work that comes out of quiet time. The three weeks somebody spent building the evaluation harness would have scored zero on utilisation. It has since paid for itself many times over. Under the old number, that work was a failure.
These choices make each hire matter more
Both decisions push the same way. They make each individual hire matter more.
A firm with a bench can afford a mediocre contractor for one project. A firm of permanent staff carrying its own gaps cannot treat the decision the same way, because the decision is open-ended rather than for three months.
That is the honest reason our process is slower and more effortful than it might be. It is also why an engineer reads every application rather than a filter.
This is not a virtue we adopted. It is what the structure requires. If we ran a bench, we would almost certainly hire faster and more loosely, because the cost of being wrong would be smaller.
These choices cost us real efficiency
We should be honest about the price.
No bench means we scale more slowly than demand. We have turned down work we wanted, twice in the last year, because taking it would have meant hiring in a hurry. It also means a quiet quarter is a real cost that lands entirely on us.
No forced ranking means we carry the risk of being too generous in assessment. That is the failure mode the curve exists to prevent. We mitigate it with written reviews that have to be specific, and we are aware that mitigation is not the same as elimination.
Both choices trade efficiency for something else. We would not claim they are correct for every firm. A large organisation with thousands of engineers has calibration problems we do not have and cannot solve the way we do.
We wrote this down because the structure explains what people notice
People considering working here should know what they are joining, including the parts that constrain us. Clients also notice several consequences of these choices:
- Continuity of team.
- Willingness to say no to work.
- A bias toward the unglamorous engineering that nobody gets individually credited for.
Those are not culture statements. They are downstream consequences of two structural decisions about employment and assessment, and they would not survive if we changed either.
Our open roles are on the careers page. What each person here is accountable for is on leadership.








