The last week before handover on every project is spent trying to break what we built. We have already been testing throughout the project. This week has a different purpose: a team member who did not write the code tries to make it fail, misbehave, leak or corrupt, with no constraints on what they try.
We have done this on every engagement for three years. It has never once found nothing.
This answers a different question from ordinary tests
Tests assert that the system does what it is supposed to do. Red-team week asks what else the system can be made to do. That question needs a different posture.
The person doing ordinary testing wrote or reviewed the requirements, so they are looking inside the space of intended behaviour. The person attacking the system is looking outside it. The most productive question they ask is "what happens if I do something nobody expected".
Tests are written once and run forever, and that is their strength. An attack is creative and unrepeatable, and that is its strength. The two methods find different kinds of problems.
Fresh eyes find the problems
The attacker must be somebody who did not build the thing. This is the single factor that most determines what gets found.
The author of a piece of code cannot see it freshly. They know why each decision was made. That knowledge suppresses the exact question that finds problems: why is this like this.
We rotate the role. The person attacking this project built the last one. That also means everybody knows their work will be attacked by a colleague who knows what they are doing.
The same kinds of issues keep appearing
After enough of these weeks, the pattern is remarkably stable. These are the findings we see, in order of frequency.
- Authorisation checked in the interface but not in the endpoint.
- Object identifiers that can be enumerated.
- Uploads treated as trusted.
- Error messages that reveal structure.
- Rate limits missing on the paths that matter.
- State machines that can be driven backwards.
Authorisation checked in the interface but not in the endpoint
This is the single most common finding, and it appears in perhaps two thirds of engagements. The button is hidden from users who should not have it. The endpoint behind the button, meaning the server path that performs the action, will happily serve anyone who calls it directly.
It happens because the interface is built first. Once the button disappears, the check feels done.
Object identifiers that can be enumerated
This usually means a sequential number in a URL, paired with an endpoint that does not verify the requester owns the object. It often appears on a secondary object rather than the main one. The invoice is protected, and the attachment on the invoice is not.
Uploads treated as trusted
We see files whose declared type is believed rather than verified. We see paths constructed from a user-supplied name. We see content served back with a type that lets it execute in a browser.
Error messages that reveal structure
This can be a stack trace in a response, an error that distinguishes between wrong password and unknown user, or an exception carrying a query.
Rate limits missing on the paths that matter
Login is usually protected. Password reset, search, export and anything that sends an email frequently are not.
State machines that can be driven backwards
This means submitting a form twice, replaying a request, or completing step three without step two. On anything involving money or inventory, this is where the expensive bugs are.
The week follows a fixed structure
Left unstructured, red-team week becomes an unfocused poke around. So we work through domains.
- Authentication and session handling.
- Authorisation on every endpoint.
- Input handling and uploads.
- State transitions and concurrency.
- Error handling and information disclosure.
- Where a model is involved, injection and privilege.
Each finding gets written up with the request that produced it, the response, and an assessment of what it would allow. Anything that permits data to be read or modified across a tenancy boundary stops the handover. A tenancy boundary is the line between one customer, client or account space and another.
Most findings come from our logic
Dependency scanning is automated and continuous. It catches known vulnerabilities in things we did not write. Red-team week almost never finds those, because the scanner already did.
What red-team week finds is logic. It finds authorisation that was checked in one place and not another. It finds an assumption that a value could not be negative. It finds a workflow that nobody considered running in reverse.
Those are invisible to a scanner because they require understanding what the application is for. They are also the majority of what actually goes wrong.
A search index exposed document snippets
On one engagement, we built a document management system for a professional services client. Everything was correct on the obvious paths. Documents were scoped to a matter, matters were scoped to a client, and access was checked on every document endpoint.
The search index was different. Search returned snippets. The snippets were generated from document content. The index had been built without the access metadata because it was added in a later sprint.
Nobody could open a document they were not entitled to open. Anybody could search for a phrase and read a hundred characters of it in the results.
For a firm whose entire obligation is that matters do not leak into each other, that was the most serious finding of the project. It was found on the Wednesday of red-team week by somebody typing a client’s name into a search box on a hunch.
The tools stay simple
People expect this to involve specialist software. Mostly it involves a small set of tools.
- A proxy that lets you see and modify every request the application makes.
- A terminal.
- A text file of notes.
The proxy is the important one. It lets the attacker edit a request before it is sent. The whole premise is that the browser is not the only client, and most logic flaws become obvious once you can change the request directly.
Automated scanners have a place, and that place is narrow. They are good at finding known vulnerabilities in dependencies. They are also good at spotting missing headers. Both should already be handled by continuous checks rather than by a person in the final week.
Scanners are close to useless at logic. They do not know that an invoice belongs to a client. They do not know that a refund should not exceed the payment.
What helps is a second account. Almost every serious finding we make comes from having two users in different tenancies open side by side, then trying to make one see the other’s data. That takes no tooling at all. It is the exercise most often skipped, because setting up realistic multi-tenant test data is tedious and nobody budgets for it.
Models add one more place to attack
Where an engagement includes a model, the week gains a section. The mindset shift is the same one that governs the rest of the exercise. We do not spend long trying to make the model say something inappropriate, which is interesting but rarely consequential. We spend the time trying to make the model cause an action.
That means planting instructions in every piece of content the system will ingest.
- Documents.
- File names.
- Metadata.
- The middle of long texts where attention is weaker.
- Any field a user can type into that later reaches a context window.
Then we watch what the deterministic layer does with the resulting requests. The finding we are looking for is that the system carried out the action. The model frequently complies, but the important part is what the surrounding system does next.
The other productive avenue is the output path. We check the code around the model output.
- Whether model output is rendered as markdown with remote images enabled.
- Whether links are constructed from model output without validation.
- Whether anything downstream parses the response in a way that could be manipulated.
That class of problem is invisible from the model’s behaviour. It lives entirely in the code around it.
Findings go into the backlog and the handover
Findings go into the same backlog as everything else, with a severity. Anything that crosses a tenancy or authorisation boundary is fixed before handover, full stop. Anything that requires an implausible chain of conditions is documented and prioritised with the client rather than fixed silently.
The write-up goes to the client, including the findings we did not fix and why. That is occasionally an awkward document to hand over. It is considerably less awkward than the same information arriving from a penetration tester the client hired afterwards.
The habit changes how engineers build
The effect we did not anticipate is on the work that happens before red-team week. Once engineers know a colleague will spend five days trying to break what they wrote, and that the findings are shared, the code that arrives at red-team week is different.
Authorisation gets checked at the endpoint because nobody wants to be the person whose button-hiding was the finding. This is a slightly cynical mechanism. It works better than any amount of guidance about secure coding, because it is specific, social and imminent.
Small teams can run a compressed version
For a small team, releasing somebody for a week is real. The compressed version that still works is two days, one person, working strictly through the domain checklist rather than freestyling. It focuses on authorisation and state transitions, because that is where the serious findings cluster.
The author attacking their own code does not work. Reviewing the checklist as a thought experiment, without actually issuing the requests, does not work either. The value is in the doing. More specifically, it is in the surprise when something that should have failed returns a two hundred.
We say it always finds something because it does
There is a version of this post that says our process is so rigorous that red-team week usually comes up clean. That would be a better advertisement, and it would be false.
Every project has findings. Competent engineers, reviewing each other’s work, with tests and static analysis, still ship authorisation gaps. Building a system and attacking it are different cognitive activities, and the first does not produce the second.
A supplier telling you their process prevents these entirely is telling you they do not look. The useful question is whether somebody went looking before handover, and what they did with what they found.
If you take one thing from this, make it the second account. Set up two users in different tenancies, log into both, and spend an hour trying to make one of them see something belonging to the other. You do not need a security specialist, a budget or a tool anybody has to buy. In our experience that single hour finds more than a week of reading code.
More on the delivery method in how we work, and on our security practice in cybersecurity services.








