How to do discovery research
Discovery is how a hypothesis stops being a hunch. Keep the evidence linked to the work, not buried in an archive.
- steps
- 9
- named failure modes
- 5
- definition-of-done criteria
- 6
How to do discovery research, in one paragraph
Discovery research turns an assumption you are pricing into evidence that could have proved you wrong. In The Tenhaw Way it runs in an epic's research phase, before design and before the ready-for-dev gate, and its only job is to test whether the epic's planned share of the outcome's currency target survives contact with reality. Good discovery is timeboxed against the quarter it has to ship in, ends in a revised number rather than a report, and leaves its evidence attached to the epic where the gate and outcome validation will both look for it.
What these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.
Run discovery when an epic's planned currency share rests on something you cannot yet evidence: that the behaviour you are pricing exists, that enough people have the problem, or that the value is the size you claimed. If no finding would change the number, the scope or the decision to build, skip discovery and start building.
- Time to run it once
- Two weeks, capped, and one week under about £100k of value
- What you need
- A markdown repository for the evidence corpus
- The epic the research is attached to
Step by step
Each step is deep-linkable, so you can send a colleague the one that is in dispute.
- 1
Name the number the research could move
Discovery is not a standing activity.
Open the epic and write three figures at the top of the research file: the outcome's currency target, the epic's planned share of it, and the threshold at which your decision changes. Something like: below £120k this epic leaves Q3. Without that threshold you will produce findings nobody can act on, because no result is ever obviously bad enough. If the epic does not exist yet, this is outcome shaping rather than discovery, and the question is whether the outcome can carry a credible number at all.
- 2
Write the hypothesis so it can fail
State the belief in one sentence with the number attached: trade customers reorder near-identical baskets often enough that one-tap templates lift repeat order rate by two points, worth £300k of gross profit this quarter. Underneath it write the kill condition as a measurable result, not a mood: fewer than four of eight interviewees rebuild orders from history, or under a quarter of accounts show repeat baskets in the data. Have product and engineering sign off both before fieldwork starts. Arguing about what would change your mind is cheap beforehand and impossible afterwards.
- 3
Timebox discovery against the quarter
The roadmap is one quarter, twelve to thirteen weeks, and the epic has to ship inside it.
Cap discovery at two weeks, one week for anything under roughly £100k of planned value, and put the end date on the epic before you start. If two weeks cannot answer the question, that is itself a finding: either the epic is too speculative for this quarter and moves out, or you cut the question down to the single assumption that most threatens the number and answer only that. Open-ended research is how epics quietly leave the roadmap without anyone deciding they should.
- 4
Pick methods that can return a no
Match the method to the risk.
For desirability, six to eight interviews with people who have the problem and are not your fans, recruited on day one from support tickets, churned accounts and lapsed users rather than the customer advisory board. For the size of the prize, query behaviour you already hold: order history, funnel drop-off, ticket volumes, cancellation reasons, and pull the denominator first so you know how many people the epic can reach. For feasibility, put an engineer in the room for an afternoon. Run at least one method that returns a number; interviews alone will not defend a currency figure at the gate.
- 5
Build the evidence corpus in markdown
Convert every source into markdown in a repository the whole team and your models can read: interview transcripts in full, the query you ran alongside its output, ticket exports, competitor behaviour described in prose. One source per file, each opening with a date, who or what it came from, and how it was collected. Do not summarise on the way in. A tidy summary written during collection is where the inconvenient quote disappears, and it strips out the raw material the next step works on. Name files so a person and a model can both tell what is inside without opening them.
- 6
Run a contradiction pass, then check it
Ask a model to read the whole corpus and return, with file and line references, what the evidence supports, what it contradicts, and where it is silent. Silence is the useful category: it tells you which part of your number nothing in the corpus touches. Then open every reference and confirm the quote exists and says what the summary claims. Expect to reject some of it, and record which claims you dropped and why, because that record is what stops the same claim reappearing at the gate. Synthesis across forty documents is what AI is good at. Being the only reader of the evidence is what it is not.
- 7
Re-price the epic in the open
Discovery ends with a number, not a narrative.
Take the epic's planned share and confirm it, change it, or set it to zero and close the epic. Show the arithmetic in four lines: how many people the behaviour applies to, the adoption or conversion rate the evidence supports, value per event, and the realisation factor your own delivery history justifies. If you land seventy per cent of what you plan, apply seventy per cent now rather than discovering it in month three. Record the baseline the epic starts from as well, because validation cannot prove a lift without one. A smaller number is a successful discovery.
- 8
Write the decision and clear the gate
Discovery is finished when a written decision sits on the epic: proceed at this number, reshape to this smaller scope, or stop.
Underneath it, list the questions you knowingly proceeded without and name who owns each one. Proceeding on an unanswered question is a decision someone owns, not an omission nobody noticed. Then take it through the gate your team's mode requires. In an AI-augmented team the discovery output becomes the first product-approved story attached to the epic, which ready-for-dev demands alongside product and engineering approval. If the epic cannot clear the gate, say so in the decision rather than letting it drift.
- 9
Book the rematch at outcome validation
Write the discovery estimate, in currency, into the epic itself, so outcome validation reads it without archaeology.
Every live outcome is validated monthly, and the epic sits in value monitoring until its value is confirmed or deliberately written off. When that happens, compare what landed against what discovery predicted and record the ratio as a single number. After three or four quarters you have a calibration curve built from your own work, and discovery stops being an argument about optimism and becomes an argument about your own history. This is the loop almost nobody closes, and it is what makes the next discovery cheaper.
Worked example: saved order templates at a trade distributor
Kesteven Supply is an invented UK trade distributor running a trade account portal, and every number below is made up to show the shape of the work. The outcome: lift repeat order rate on the portal, target £1.4m of additional gross profit across four quarters. One Q3 epic, saved order templates, carries a planned share of £300k. Hypothesis: trade customers reorder near-identical baskets often enough that one-tap templates lift repeat order rate by two points. Kill condition: fewer than four of eight interviewees rebuild orders from history, or under a quarter of accounts show repeat baskets in twelve months of order data. Threshold for action: below £120k the epic leaves Q3. Nine working days of a two-week box. The order history query returns 4,200 active trade accounts, of which 1,600, or 38 per cent, place orders where at least 80 per cent of lines match a previous order. Eight interviews recruited from lapsed and mid-tier accounts rather than the top twenty: six rebuild baskets by scrolling order history, two paste product codes from their own spreadsheet. Both spreadsheet users said templates would save them time, then said they order weekly whatever happens. A fake-door save-as-template button shown to 900 accounts for two weeks was clicked by 31 per cent. The kill condition was not met, so the epic survived. The number did not. Re-priced: 1,600 eligible accounts, 20 per cent sustained use rather than the 31 per cent click, giving 320 accounts. Of those, the evidence supports extra orders only from the group citing hassle on small top-ups, about 40 per cent, so 128 accounts, one extra order a month, £340 average gross profit per order, over three months: £130,560. Apply the 70 per cent realisation factor from the last four quarters and it is £91k. Decision: proceed at £91k, sequenced behind two larger epics. Baseline recorded: 21 per cent repeat order rate over the prior 90 days. Proceeding without an answer on whether templates cannibalise higher-margin phone orders through the sales desk, owned by the commercial lead. The visible consequence is £209k of the outcome's target with no epic behind it, argued in week three of the quarter rather than week thirteen.
Where this goes wrong
- 01
Research designed to confirm.
If the sample is customers who already love the product and the questions ask whether they would like a new feature, everyone says yes and nothing is learned. Recruit for the problem rather than for the affection, and write the kill condition before the first conversation rather than after it.
- 02
No baseline captured.
Discovery predicts a two-point lift and nobody records what the rate was the week before the epic shipped. Four months later, outcome validation has a number with nothing to compare it against, and the epic gets closed on an argument instead of evidence. Capture the baseline while you are already in the data.
- 03
Discovery as a parking bay.
An epic nobody wants to kill can sit in the research phase indefinitely and look busy. If an epic has been in research for more than two weeks with no decision, that is not research, it is an unmade prioritisation call. Move it out of the quarter and say so out loud.
- 04
Treating AI synthesis as the evidence.
A model summarising forty transcripts produces something coherent and compresses away the outlier that changes the answer. Ask for file and line references, open them, and treat any claim without a traceable source as not yet true.
- 05
Leaving the planned value untouched.
Teams run discovery, learn something material, then ship the epic carrying the number it was given in planning. If the evidence did not move the currency figure or explicitly confirm it, discovery has not finished.
Done means
- The hypothesis is written as a statement that could have been falsified, with a measurable kill condition, and the evidence says which way it went.
- Every claim traces to a named source in the corpus, a transcript, a query, a ticket, that someone else can open without asking you.
- The epic's planned currency share is confirmed or changed, with the arithmetic, the baseline it starts from and the realisation factor shown.
- A written decision sits on the epic: proceed, reshape or stop, plus the questions you knowingly proceeded without and who owns them.
- The evidence and the decision are attached to the epic and its outcome where the work is tracked, not in a shared drive.
- The epic can clear its mode's ready-for-dev gate, or discovery has stated plainly that it cannot yet and why.
In an AI-augmented team, discovery output becomes the first product-approved story attached to the epic, and that story is one of the three things ready-for-dev checks. In an AI-native team there is no child story, so discovery has to produce the key user journeys and the test requirements that go into the outcome ticket itself. That changes the fieldwork, not just the write-up: you need the sequence people actually follow, the edge cases and the failure states, rather than sentiment about whether they would like the feature. It also raises the bar on the corpus, because the same markdown files are what the model builds from.
The two delivery modes, side by side →Questions
How long should discovery take?
Two weeks at most, and one week for anything under roughly £100k of planned value. The constraint is not research quality, it is the quarter: the epic has to ship inside the same twelve to thirteen weeks, so every day in research is a day off the build. If two weeks cannot answer the question, reduce the question to the single assumption that most threatens the number, or move the epic out of the quarter and say why.
Does every epic need discovery?
No. Discovery is for epics whose planned currency share rests on an assumption you cannot evidence. If the behaviour is already visible in your data, the value is a rate change you can calculate, and nothing you could learn would change the scope or the number, skip it and build. Running discovery on an epic whose answer is already known is the most common way a research phase turns into a parking bay.
Who should run discovery?
The product manager who owns the epic, with an engineer present for at least one session and a second person reading the raw evidence. Whoever wants the epic to succeed should not be the only one reading the transcripts. It is the same reason the kill condition is written and signed off before fieldwork rather than after the results arrive.
Can AI do the discovery for us?
It can do the synthesis, which is the slow part: reading forty documents and returning what they support, contradict and leave silent, with references. It cannot be the only reader. Models compress away the outlier, and the outlier is often the finding. Ask for file and line references, open them, and treat anything without a traceable source as not yet true.
More on setting up the work
These guides are written to be read in order.
How to set an outcome
An outcome is a business result with a price tag. If you cannot price it, it is not an outcome, it is a wish.
How to put together a quarterly roadmap
A roadmap is one quarter, twelve to thirteen weeks. Outcomes can span quarters; epics cannot.
How to write a value-focused epic
An epic is your unit of value contribution to an outcome. If it does not carry a number, it is a feature wishlist.
How to break an epic into stories
A story is the smallest piece of user-visible value the team can ship. If it does not change the user's experience, it is a chapter, not a story.
Want help installing this?
These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.
Most organisations start with a fixed-price Agent-Readiness Audit · £30k–£90k · 6–8 weeks