Almost every software engagement of any size begins with something called discovery. It lasts between two and twelve weeks, it produces documents, and it is billed.
That description covers two very different things. Some discovery is the most valuable work on the project. It finds the facts that decide whether the build can be priced and delivered sensibly. Some discovery is a way for a supplier to be paid while delaying the moment they have to commit to anything.
Those two versions are difficult to tell apart from the outside. That is why the second version keeps turning up.
Discovery should remove the unknowns that matter
Most projects start with real uncertainty. Someone has to remove it before a serious price or plan means much. The useful questions are usually specific, and they cannot be answered by guessing.
These are the kinds of things discovery needs to find out:
- What the legacy system actually does, as opposed to what its documentation claims.
- What is in the data.
- What the people who will use this need it to do, as distinct from what the person commissioning it believes they need.
- What the integration partner’s API does when you send it something unusual.
A legacy system is the existing system the business already depends on. An API is the way one system talks to another. Both can behave differently from their documents, especially around old data, odd inputs and error cases.
Every one of those questions has an answer that can only be obtained by looking. Looking takes time. Work that removes those unknowns is valuable, and a supplier who skips it and quotes anyway is either padding heavily or about to have a difficult project.
The weak version gives you a report instead of an answer
The version that avoids analysis has a recognisable shape. It produces a document rather than an answer. It ends with a recommendation to proceed to the next phase. It contains a great deal of context, several diagrams and a summary of workshops.
The missing thing matters most. It contains no number the supplier is willing to be held to.
That is the tell, and it is close to the only one that matters. A discovery phase that does not end in a commitment has reduced the supplier’s risk, while leaving yours where it was. You are several weeks and a substantial invoice further along, with the same fundamental uncertainty you started with.
The test is whether the supplier will be bound by a price
Before agreeing to a discovery phase, ask one question: at the end of this, will you give me a fixed price against a written scope, and if not, what specifically will still be unknown.
The answers separate very cleanly. A supplier doing real analysis will say yes, or will name precisely what will remain uncertain and why.
A supplier doing the other thing will often explain several points:
- Software is inherently unpredictable.
- Estimates depend on many factors.
- They prefer to work collaboratively on a time and materials basis.
Each of those statements is individually defensible. Together they mean the risk is staying with you.
Discovery should be small, separately priced, and finished
Ours are typically one to three weeks and they are priced separately and modestly, because the purpose is to buy down uncertainty rather than to be a revenue line. A discovery phase that costs a significant fraction of the build is the project in a costume.
It also has a defined end with a defined deliverable. The deliverable is a definition of done and a price. Context and diagrams may accompany it, but the price and scope are the part that matter.
A definition of done sets out what finished means for the work being quoted. Without that, a fixed price is only a number beside a loose description.
During discovery, we test the risky parts directly
The useful work is usually hands-on. We try to replace assumptions with evidence, and we focus on the things that could make the build harder than it first appears.
This is what we actually do during it:
- Connect to the things. We do not read about them. We actually authenticate against the legacy system and the third party API and observe their behaviour, including error cases and rate limits.
- Look at the real data. We use a full export or a genuinely random sample, never a curated one. We count the nulls, find the schema changes, and look at the oldest records.
- Watch somebody work. We sit with the people who do the job today. This consistently contradicts something in the brief.
- Establish the numbers. Volume, concurrency, growth. If nobody knows, measure.
- Build the smallest risky thing. If one part is genuinely novel, we build a rough version of just that part to find out whether the approach works.
That last one is the highest value activity and the one most often missing. A week spent proving that the hard part is possible is worth more than four weeks of workshops about everything else.
Workshops can miss the risks that matter
Workshops feel productive. Everyone is engaged, the walls fill with notes, and the output looks substantial. They are also the easiest way to spend three weeks without learning anything that was not already known by somebody in the room.
Workshops surface opinions and requirements. Most project risk sits in the data, the integrations and the legacy behaviour, none of which will attend a workshop. We keep workshops short and use the time to look at systems instead.
A previous discovery report rarely gives us enough to quote
We are sometimes handed a discovery document produced by another supplier and asked to quote against it. We read these documents carefully, and we usually still do a short investigation of our own. Clients occasionally find that frustrating.
The reason is that these documents almost never contain what we need. They describe scope and rarely establish behaviour. There will be a paragraph on the integration and no evidence anybody called it. There will be a data model and no counts.
That does not mean we doubt the previous supplier’s competence. We are unable to commit to a fixed price on the basis of somebody else’s untested assumptions, and saying so is more honest than quoting and hoping.
Access problems stretch discovery before analysis starts
The most common reason a discovery phase takes four weeks instead of two has nothing to do with analysis. It is that the supplier spent eleven days waiting for credentials.
The usual access needs are ordinary, but they still take time:
- A read-only account on the legacy database.
- A sandbox key for the payment provider.
- A copy of production data with an approval from someone in compliance.
- An hour with the person who understands the pricing rules.
None of that is difficult and all of it takes longer than anybody expects, because each request goes to a person for whom it is not a priority. A sandbox is a safe test version of a live service, and it often has to be issued by the provider.
So we list the access we need in the proposal itself, with the dates by which we need it, and we treat it as a binding assumption rather than a polite request. If the sandbox is not available by day three, that is a defined event with a defined consequence. We do not absorb it quietly as a surprise.
Clients sometimes read this as bureaucracy at the start of a relationship. We see it as the supplier telling you, before you have spent anything, what you will need to do for this to go well. A supplier who does not raise it has either not thought about it or intends to bill for the waiting.
Model work has to prove quality on your data
Where a model is in scope, discovery has a different centre of gravity, and it is worth being explicit because the sector is full of proposals that skip it. The uncertainty is almost never whether a system can be built. It is whether it can reach a quality bar on this particular data, and no amount of architecture discussion answers that.
So the deliverable is a measurement. We assemble a held-out set from examples somebody has already completed correctly, run a baseline against it, and report a number with the failure cases attached. A held-out set is a group of examples kept back for testing. A baseline is the first measured run, used as the point of comparison.
That takes a few days and it converts the entire commercial conversation from speculation into arithmetic. If the baseline is already close to the bar, the remaining work is small and we can price it. If it is far away, the client finds out in week one for a small sum rather than in month four for a large one.
It also produces the artefact the project will need anyway. The evaluation set built during discovery is the same set that guards against regression for the life of the system, so none of the effort is discarded when the build starts. Regression means the system getting worse against checks it used to pass.
The commercial model changes how suppliers behave
Under time and materials, discovery that runs long is revenue. Under fixed price, discovery that runs long is cost, and discovery that fails to find something is a much larger cost later.
That difference produces different behaviour, and it is the strongest argument for the commercial model rather than for any particular process. We are motivated to find the awkward thing in week one because we will be paying for it in week nine, and the client gets the benefit of an incentive rather than of a promise.
One project changed when we tested the data
A client came to us with a discovery document from a previous supplier, running to sixty pages, recommending a twenty two week build. The document described the legacy system’s data model in detail, taken from its documentation.
We spent nine days. Four of them were on the data, where we found that the field the entire migration depended on had been repurposed in 2019 and contained two different kinds of value with no flag distinguishing them. Establishing which was which required a rule that only two people in the organisation knew, and one of them was retiring.
That single finding changed the shape of the project substantially. It also meant our quote was higher than the previous supplier’s estimate, which is an uncomfortable position to be in commercially. We were able to show exactly why, with the query and the counts, and the client took it.
Twenty two weeks would not have been achievable. The overrun would have surfaced in month four, and it would have been characterised as a scope disagreement rather than as an assumption nobody tested.
These are the things to ask for
If you are being asked to pay for discovery, the questions should force the work toward evidence and commitment.
- What will you have connected to, by the end, rather than read about?
- Will you have looked at the full production data or a sample somebody chose?
- Which part of this do you consider most likely to be harder than it looks, and how will you find out during discovery?
- Does this end in a fixed price, and if not, what remains unknown?
- What would you have to discover for you to recommend not proceeding?
The last question is the most revealing. A supplier who cannot describe a finding that would lead them to advise against the project is not going to produce one, and a discovery phase that can only end in a recommendation to proceed is not an investigation.
More on how we price in engagement models, and on the method in how we work.








