Setting up the work

How to run product discovery research

Discovery is how a hypothesis stops being a hunch. Keep the evidence linked to the work, not buried in an archive.
steps
9
named failure modes
5
definition-of-done criteria
6

How to run product discovery research, in one paragraph

Discovery research turns an assumption you are pricing into evidence that could have proved you wrong. In The Tenhaw Way it runs in an epic's research phase, before design and before the ready-for-dev gate, and its only job is to test whether the epic's planned share of the outcome's currency target survives contact with reality. Good discovery is timeboxed against the quarter it has to ship in, ends in a revised number rather than a report, and leaves its evidence attached to the epic where the gate and outcome validation will both look for it.

That is the procedure. The call is where it meets your delivery structure.

Talk it through

What these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.

When to use this

Run discovery when an epic's planned currency share rests on something you cannot yet evidence: that the behaviour you are pricing exists, that enough people have the problem, or that the value is the size you claimed. If no finding would change the number, the scope or the decision to build, skip discovery and start building.

Time to run it once
Two weeks, capped, and one week under about £100k of value
What you need
  • A markdown repository for the evidence corpus
  • The epic the research is attached to

If you do not have those in place, the call is a good place to work out what comes first.

Talk it through
On this page
9 steps

Step by step

Each step is deep-linkable, so you can send a colleague the one that is in dispute.

  1. 1

    Name the number the research could move

    Discovery is not a standing activity.

    Open the epic and write three figures at the top of the research file: the outcome's currency target, the epic's planned share of it, and the threshold at which your decision changes. Something like: below £120k this epic leaves Q3. Without that threshold you will produce findings nobody can act on, because no result is ever obviously bad enough. If the epic does not exist yet, this is outcome shaping rather than discovery, and the question is whether the outcome can carry a credible number at all.

  2. 2

    Write the hypothesis so it can fail

    State the belief in one sentence with the number attached: trade customers reorder near-identical baskets often enough that one-tap templates lift repeat order rate by two points, worth £300k of gross profit this quarter. Underneath it write the kill condition as a measurable result, not a mood: fewer than four of eight interviewees rebuild orders from history, or under a quarter of accounts show repeat baskets in the data. Have product and engineering sign off both before fieldwork starts. Arguing about what would change your mind is cheap beforehand and impossible afterwards.

  3. 3

    Timebox discovery against the quarter

    The roadmap is one quarter, twelve to thirteen weeks, and the epic has to ship inside it.

    Cap discovery at two weeks, one week for anything under roughly £100k of planned value, and put the end date on the epic before you start. If two weeks cannot answer the question, that is itself a finding: either the epic is too speculative for this quarter and moves out, or you cut the question down to the single assumption that most threatens the number and answer only that. Open-ended research is how epics quietly leave the roadmap without anyone deciding they should.

  4. 4

    Pick methods that can return a no

    Match the method to the risk.

    For desirability, six to eight interviews with people who have the problem and are not your fans, recruited on day one from support tickets, churned accounts and lapsed users rather than the customer advisory board. For the size of the prize, query behaviour you already hold: order history, funnel drop-off, ticket volumes, cancellation reasons, and pull the denominator first so you know how many people the epic can reach. For feasibility, put an engineer in the room for an afternoon. Run at least one method that returns a number; interviews alone will not defend a currency figure at the gate.

  5. 5

    Build the evidence corpus in markdown

    Convert every source into markdown in a repository the whole team and your models can read: interview transcripts in full, the query you ran alongside its output, ticket exports, competitor behaviour described in prose. One source per file, each opening with a date, who or what it came from, and how it was collected. Do not summarise on the way in. A tidy summary written during collection is where the inconvenient quote disappears, and it strips out the raw material the next step works on. Name files so a person and a model can both tell what is inside without opening them.

  6. 6

    Run a contradiction pass, then check it

    Ask a model to read the whole corpus and return, with file and line references, what the evidence supports, what it contradicts, and where it is silent. Silence is the useful category: it tells you which part of your number nothing in the corpus touches. Then open every reference and confirm the quote exists and says what the summary claims. Expect to reject some of it, and record which claims you dropped and why, because that record is what stops the same claim reappearing at the gate. Synthesis across forty documents is what AI is good at. Being the only reader of the evidence is what it is not.

  7. 7

    Re-price the epic in the open

    Discovery ends with a number, not a narrative.

    Take the epic's planned share and confirm it, change it, or set it to zero and close the epic. Show the arithmetic in four lines: how many people the behaviour applies to, the adoption or conversion rate the evidence supports, value per event, and the realisation factor your own delivery history justifies. If you land seventy per cent of what you plan, apply seventy per cent now rather than discovering it in month three. Record the baseline the epic starts from as well, because validation cannot prove a lift without one. A smaller number is a successful discovery.

  8. 8

    Write the decision and clear the gate

    Discovery is finished when a written decision sits on the epic: proceed at this number, reshape to this smaller scope, or stop.

    Underneath it, list the questions you knowingly proceeded without and name who owns each one. Proceeding on an unanswered question is a decision someone owns, not an omission nobody noticed. Then take it through the gate your team's mode requires. In an AI-augmented team the discovery output becomes the first product-approved story attached to the epic, which ready-for-dev demands alongside product and engineering approval. If the epic cannot clear the gate, say so in the decision rather than letting it drift.

  9. 9

    Book the rematch at outcome validation

    Write the discovery estimate, in currency, into the epic itself, so outcome validation reads it without archaeology.

    Every live outcome is validated monthly, and the epic sits in value monitoring until its value is confirmed or deliberately written off. When that happens, compare what landed against what discovery predicted and record the ratio as a single number. After three or four quarters you have a calibration curve built from your own work, and discovery stops being an argument about optimism and becomes an argument about your own history. This is the loop almost nobody closes, and it is what makes the next discovery cheaper.

Bring a real piece of work to the call and we will walk it through these.

Talk it through
Worked example

Worked example: saved order templates at a trade distributor

Kesteven Supply is an invented UK trade distributor running a trade account portal, and every number below is made up to show the shape of the work. The outcome: lift repeat order rate on the portal, target £1.4m of additional gross profit across four quarters. One Q3 epic, saved order templates, carries a planned share of £300k. Hypothesis: trade customers reorder near-identical baskets often enough that one-tap templates lift repeat order rate by two points. Kill condition: fewer than four of eight interviewees rebuild orders from history, or under a quarter of accounts show repeat baskets in twelve months of order data. Threshold for action: below £120k the epic leaves Q3. Nine working days of a two-week box. The order history query returns 4,200 active trade accounts, of which 1,600, or 38 per cent, place orders where at least 80 per cent of lines match a previous order. Eight interviews recruited from lapsed and mid-tier accounts rather than the top twenty: six rebuild baskets by scrolling order history, two paste product codes from their own spreadsheet. Both spreadsheet users said templates would save them time, then said they order weekly whatever happens. A fake-door save-as-template button shown to 900 accounts for two weeks was clicked by 31 per cent. The kill condition was not met, so the epic survived. The number did not. Re-priced: 1,600 eligible accounts, 20 per cent sustained use rather than the 31 per cent click, giving 320 accounts. Of those, the evidence supports extra orders only from the group citing hassle on small top-ups, about 40 per cent, so 128 accounts, one extra order a month, £340 average gross profit per order, over three months: £130,560. Apply the 70 per cent realisation factor from the last four quarters and it is £91k. Decision: proceed at £91k, sequenced behind two larger epics. Baseline recorded: 21 per cent repeat order rate over the prior 90 days. Proceeding without an answer on whether templates cannibalise higher-margin phone orders through the sales desk, owned by the commercial lead. The visible consequence is £209k of the outcome's target with no epic behind it, argued in week three of the quarter rather than week thirteen.

Yours will look different. Thirty minutes is enough to see how.

Talk it through
Failure modes

Where this goes wrong

  1. 01

    Research designed to confirm.

    If the sample is customers who already love the product and the questions ask whether they would like a new feature, everyone says yes and nothing is learned. Recruit for the problem rather than for the affection, and write the kill condition before the first conversation rather than after it.

  2. 02

    No baseline captured.

    Discovery predicts a two-point lift and nobody records what the rate was the week before the epic shipped. Four months later, outcome validation has a number with nothing to compare it against, and the epic gets closed on an argument instead of evidence. Capture the baseline while you are already in the data.

  3. 03

    Discovery as a parking bay.

    An epic nobody wants to kill can sit in the research phase indefinitely and look busy. If an epic has been in research for more than two weeks with no decision, that is not research, it is an unmade prioritisation call. Move it out of the quarter and say so out loud.

  4. 04

    Treating AI synthesis as the evidence.

    A model summarising forty transcripts produces something coherent and compresses away the outlier that changes the answer. Ask for file and line references, open them, and treat any claim without a traceable source as not yet true.

  5. 05

    Leaving the planned value untouched.

    Teams run discovery, learn something material, then ship the epic carrying the number it was given in planning. If the evidence did not move the currency figure or explicitly confirm it, discovery has not finished.

Definition of done

Done means

  • The hypothesis is written as a statement that could have been falsified, with a measurable kill condition, and the evidence says which way it went.
  • Every claim traces to a named source in the corpus, a transcript, a query, a ticket, that someone else can open without asking you.
  • The epic's planned currency share is confirmed or changed, with the arithmetic, the baseline it starts from and the realisation factor shown.
  • A written decision sits on the epic: proceed, reshape or stop, plus the questions you knowingly proceeded without and who owns them.
  • The evidence and the decision are attached to the epic and its outcome where the work is tracked, not in a shared drive.
  • The epic can clear its mode's ready-for-dev gate, or discovery has stated plainly that it cannot yet and why.

If you recognise one of those already happening, that is a good call to have.

Talk it through
If your team is AI-native

In an AI-augmented team, discovery output becomes the first product-approved story attached to the epic, and that story is one of the three things ready-for-dev checks. In an AI-native team there is no child story, so discovery has to produce the key user journeys and the test requirements that go into the outcome ticket itself. That changes the fieldwork, not just the write-up: you need the sequence people actually follow, the edge cases and the failure states, rather than sentiment about whether they would like the feature. It also raises the bar on the corpus, because the same markdown files are what the model builds from.

The two delivery modes, side by side →

Which mode your team is actually in is the first thing we establish on a call.

Talk it through
The subject behind the procedure

Where this sits in a programme

The procedure is the same whatever you are building. These cover what it runs into when the thing being built is agentic.

If you want this run inside a programme rather than read, that is the conversation.

Talk it through
book a call

Want help installing this?

These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.

most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes

// pick a slot · cal.com/tenhaw/professional-servicesLIVE CALENDAR

Calendar not loading? Open it on cal.com or email hello@tenhaw.com.

Questions

How long should discovery take?

Two weeks at most, and one week for anything under roughly £100k of planned value. The constraint is not research quality but the quarter. The epic has to ship inside the same twelve to thirteen weeks, so every day in research is a day off the build. Tenhaw's own evidence work runs in that box, and a proof of concept for a London specialty insurance business took PDFs through to business intelligence on Azure in a fortnight, over ground the client had circled for roughly a year. If two weeks cannot answer the question, reduce the question to the single assumption that most threatens the number, or move the epic out of the quarter and say why.

Does every epic need discovery?

No. Discovery is for epics whose planned currency share rests on an assumption you cannot evidence. If the behaviour is already visible in your data, the value is a rate change you can calculate, and nothing you could learn would change the scope or the number, skip it and build. Running discovery on an epic whose answer is already known is the most common way a research phase turns into a parking bay.

Who should run discovery?

The product manager who owns the epic, with an engineer present for at least one session and a second person reading the raw evidence. Whoever wants the epic to succeed should not be the only one reading the transcripts. Tenhaw staffs its own engagements the same way, with teams of two or three senior people and partner oversight on every one, so the second reader is a named person rather than a spare pair of hands. It is the same reason the kill condition is written and signed off before fieldwork rather than after the results arrive.

Can AI do the discovery for us?

It can do the synthesis, which is the slow part, reading forty documents and returning what they support, contradict and leave silent, with references. Tenhaw runs the same split on paid work, a model doing the corpus pass and a named person checking every reference, with partner oversight on every engagement. It cannot be the only reader. Models compress away the outlier, and the outlier is often the finding. Ask for file and line references, open them, and treat anything without a traceable source as not yet true.

What is a kill condition in product discovery?

A kill condition is the measurable result that would make you stop or reshape the work, written down before the research starts. It has to be a number, not a mood. Fewer than four of eight interviewees rebuild orders from history, say, or under a quarter of accounts show the behaviour in the data. Product and engineering sign it off before fieldwork begins, because agreeing what would change your mind is cheap beforehand and nearly impossible once the results are in. Without one, no finding is ever obviously bad enough to act on, and discovery drifts into confirming what the team already wanted to build.

What happens if discovery shows the epic is worth less than planned?

You change the number in the open and treat that as a success, because a smaller honest figure beats a large fiction discovered in month three. Tenhaw reports a month that delivers no measurable value as a failed month. Show the arithmetic in four lines: how many people the behaviour reaches, the adoption rate the evidence supports, value per event, and the realisation factor your delivery history justifies. If the revised figure still clears the threshold you set at the start, proceed at the new number and re-sequence it against the other epics. If it falls below that threshold, the epic leaves the quarter, and the gap it leaves in the outcome's target gets argued in week three rather than week thirteen.

Where should discovery evidence be kept?

Attached to the epic it belongs to, in a markdown repository the whole team and your models can read, not in a shared drive nobody revisits. One source per file: full interview transcripts, the query you ran alongside its output, ticket exports, each opening with a date, where it came from and how it was collected. Do not summarise on the way in, because a tidy summary written during collection is where the inconvenient quote disappears. The approval gate and monthly outcome validation both look for evidence on the epic, so keeping it anywhere else guarantees archaeology later.

Can you run discovery with interviews alone?

Not if the output has to defend a currency figure. Interviews are the right tool for desirability, six to eight with people who actually have the problem, recruited from support tickets, churned accounts and lapsed users rather than your fans. But at least one method has to return a number, which means querying the behaviour you already hold, order history, funnel drop-off, ticket volumes, and pulling the denominator first so you know how many people the work can reach. People also say things their behaviour contradicts; in one worked example, two interviewees praised a proposed template feature and then explained they would order weekly whatever happened. The data settles it.

What is a realisation factor and how do you set one?

A realisation factor is the share of planned value your team historically delivers, applied to every new estimate before anyone commits to it. If your last four quarters landed seventy per cent of what was planned, apply seventy per cent now rather than discovering the shortfall in month three. You build it by closing the loop: write the discovery estimate into the epic, compare what landed at outcome validation, and record the ratio. Tenhaw publishes the method in full, free to adopt, so a team can run that loop without hiring anyone. After three or four quarters you have a calibration curve built from your own work, and estimating stops being an argument about optimism and becomes an argument about your own history.

Should discovery end in a report?

No. Discovery ends in a revised number and a written decision on the epic: proceed at this figure, reshape to a smaller scope, or stop. Underneath the decision, list the questions you knowingly proceeded without and name who owns each one, because proceeding on an unanswered question is a decision someone owns, not an omission nobody noticed. The 500-team target operating model James Rooney co-led at HSBC Global Payment Solutions was piloted rather than presented. A report is where findings go to be admired; a decision is something the approval gate can act on. If the research cannot yet produce one, say that plainly and move the epic out of the quarter rather than letting it sit in research looking busy.

Continuous discovery vs a two-week timebox: which is better?

They do different jobs, and the timeboxed pass is the one an approval gate can act on. A continuous habit keeps a research cadence running whatever the roadmap is doing, which is useful for building a picture of your users over time. The timeboxed pass runs inside one epic's research phase, is capped at two weeks and at one week for anything under roughly £100k of planned value, and ends by confirming, changing or zeroing that epic's planned share of the outcome's currency target. Its trigger is narrower too. If no finding would change the number, the scope or the decision to build, you skip it and start building.

Is a fake door test enough evidence to price a feature?

It is good evidence of interest and weak evidence of value, so discount it hard before it reaches the number. In one worked example a save-as-template button shown to 900 accounts for two weeks was clicked by 31 per cent, and the re-price used 20 per cent sustained use instead, because a click costs nothing and changing how you order every week costs something. Pair it with behaviour you already hold, such as order history or funnel drop-off, then apply the realisation factor your own delivery history justifies. A click-through rate carried straight into a currency figure is the quickest route to arriving short in month three.

What are the steps in a two-week discovery sprint?

Eight steps inside the box, plus a rematch at outcome validation later. Write three figures at the top of the research file: the outcome's currency target, the epic's planned share of it, and the threshold at which your decision changes. State the hypothesis with a measurable kill condition and have product and engineering sign both off before fieldwork starts. Put the end date on the epic, recruit interviewees on day one, and run at least one method that returns a number rather than only conversations. Convert every source into markdown, run the contradiction pass and check its references yourself, re-price the epic, then write the decision on it. The worked example lands in nine working days of the two-week box.

We only have an idea, not an epic yet. Is that discovery?

No, that is outcome shaping, and it is a different question. Discovery tests whether an epic's planned share of a priced outcome survives contact with reality, so it cannot start until you can write three figures at the top of the research file: the outcome's currency target, the epic's planned share of it, and the threshold at which your decision changes. If nothing has been priced, the question in front of you is whether the outcome can carry a credible number at all. Set the outcome first, break it into epics, then run discovery on the ones whose share rests on something you cannot yet evidence. Fieldwork first produces interesting findings and no decision.

Does discovery change if AI is writing the code?

Yes, and it changes the fieldwork rather than just the write-up. In an AI-augmented team, where a developer writes the code and the model assists, this guide runs as written. In an AI-native team there is no child story to hold the detail, so discovery has to produce the key user journeys and the test requirements that go into the outcome ticket itself. That means capturing the sequence people actually follow, the edge cases and the failure states, not sentiment about whether they would like the feature. It raises the bar on the evidence corpus, because the same markdown files are what the model builds from. Tenhaw's engineering handbook targets that reader, 72 rules on GitHub written to be enforced by an agent.

Why capture the baseline during discovery rather than at launch?

Because you are already in the data, and nobody goes back for it later. Discovery predicts a two-point lift, the epic ships, and if nobody recorded what the rate was the week before, validation four months on has a number with nothing to compare it against and the epic gets closed on an argument instead of evidence. Record it in the same pass that produces the denominator and the adoption rate, and write it into the epic beside the revised currency figure. The worked example on this page notes a 21 per cent repeat order rate over the prior 90 days, because a lift needs a starting point.

How do you work out how many customers a feature would reach?

Pull the denominator before you count anything else. Query behaviour you already hold, order history, funnel drop-off, ticket volumes or cancellation reasons, and start from the total population the epic could ever touch, so every percentage after it has a ceiling. In one worked example that is 4,200 active trade accounts, of which 1,600, or 38 per cent, place orders where at least 80 per cent of lines match a previous order. The 1,600 is the reachable group, not the value. Adoption, the share of that group whose behaviour actually changes, and value per event come after it. Skip the denominator and you finish with a rate and no idea what it applies to.

What do you do when the research says nothing either way?

Name the silence and then decide what to do about it in the open. A corpus that neither supports nor contradicts part of your number is telling you which assumption you never actually tested, which is more useful than a weak signal you might have leaned on. Two honest options fit inside the timebox. Spend what is left of it on that single gap, or re-price on what you do have and carry the question forward with an owner beside it. In the worked example the team proceeded without knowing whether templates cannibalise higher-margin phone orders through the sales desk, and named the commercial lead as the person who owns it.