Discovery and chapters, answered in full.
- questions in this group, each answered in full
- 36
- pages the answers are written on, every one linked
- 2
- questions across the whole FAQ
- 1424
36 questions on discovery and chapters, answered by Tenhaw, a UK AI consultancy and AI delivery partner based in London. Nothing here is a summary: each answer is the exact text from the page that owns it, and every group links back to that page for the context around it.
On this page
How to run product discovery research
Answered on How to run product discovery research, and rendered here in the same words.
Read the page these answers live on →
How long should discovery take?
Two weeks at most, and one week for anything under roughly £100k of planned value. The constraint is not research quality but the quarter. The epic has to ship inside the same twelve to thirteen weeks, so every day in research is a day off the build. Tenhaw's own evidence work runs in that box, and a proof of concept for a London specialty insurance business took PDFs through to business intelligence on Azure in a fortnight, over ground the client had circled for roughly a year. If two weeks cannot answer the question, reduce the question to the single assumption that most threatens the number, or move the epic out of the quarter and say why.
Does every epic need discovery?
No. Discovery is for epics whose planned currency share rests on an assumption you cannot evidence. If the behaviour is already visible in your data, the value is a rate change you can calculate, and nothing you could learn would change the scope or the number, skip it and build. Running discovery on an epic whose answer is already known is the most common way a research phase turns into a parking bay.
Who should run discovery?
The product manager who owns the epic, with an engineer present for at least one session and a second person reading the raw evidence. Whoever wants the epic to succeed should not be the only one reading the transcripts. Tenhaw staffs its own engagements the same way, with teams of two or three senior people and partner oversight on every one, so the second reader is a named person rather than a spare pair of hands. It is the same reason the kill condition is written and signed off before fieldwork rather than after the results arrive.
Can AI do the discovery for us?
It can do the synthesis, which is the slow part, reading forty documents and returning what they support, contradict and leave silent, with references. Tenhaw runs the same split on paid work, a model doing the corpus pass and a named person checking every reference, with partner oversight on every engagement. It cannot be the only reader. Models compress away the outlier, and the outlier is often the finding. Ask for file and line references, open them, and treat anything without a traceable source as not yet true.
What is a kill condition in product discovery?
A kill condition is the measurable result that would make you stop or reshape the work, written down before the research starts. It has to be a number, not a mood. Fewer than four of eight interviewees rebuild orders from history, say, or under a quarter of accounts show the behaviour in the data. Product and engineering sign it off before fieldwork begins, because agreeing what would change your mind is cheap beforehand and nearly impossible once the results are in. Without one, no finding is ever obviously bad enough to act on, and discovery drifts into confirming what the team already wanted to build.
What happens if discovery shows the epic is worth less than planned?
You change the number in the open and treat that as a success, because a smaller honest figure beats a large fiction discovered in month three. Tenhaw reports a month that delivers no measurable value as a failed month. Show the arithmetic in four lines: how many people the behaviour reaches, the adoption rate the evidence supports, value per event, and the realisation factor your delivery history justifies. If the revised figure still clears the threshold you set at the start, proceed at the new number and re-sequence it against the other epics. If it falls below that threshold, the epic leaves the quarter, and the gap it leaves in the outcome's target gets argued in week three rather than week thirteen.
Where should discovery evidence be kept?
Attached to the epic it belongs to, in a markdown repository the whole team and your models can read, not in a shared drive nobody revisits. One source per file: full interview transcripts, the query you ran alongside its output, ticket exports, each opening with a date, where it came from and how it was collected. Do not summarise on the way in, because a tidy summary written during collection is where the inconvenient quote disappears. The approval gate and monthly outcome validation both look for evidence on the epic, so keeping it anywhere else guarantees archaeology later.
Can you run discovery with interviews alone?
Not if the output has to defend a currency figure. Interviews are the right tool for desirability, six to eight with people who actually have the problem, recruited from support tickets, churned accounts and lapsed users rather than your fans. But at least one method has to return a number, which means querying the behaviour you already hold, order history, funnel drop-off, ticket volumes, and pulling the denominator first so you know how many people the work can reach. People also say things their behaviour contradicts; in one worked example, two interviewees praised a proposed template feature and then explained they would order weekly whatever happened. The data settles it.
What is a realisation factor and how do you set one?
A realisation factor is the share of planned value your team historically delivers, applied to every new estimate before anyone commits to it. If your last four quarters landed seventy per cent of what was planned, apply seventy per cent now rather than discovering the shortfall in month three. You build it by closing the loop: write the discovery estimate into the epic, compare what landed at outcome validation, and record the ratio. Tenhaw publishes the method in full, free to adopt, so a team can run that loop without hiring anyone. After three or four quarters you have a calibration curve built from your own work, and estimating stops being an argument about optimism and becomes an argument about your own history.
Should discovery end in a report?
No. Discovery ends in a revised number and a written decision on the epic: proceed at this figure, reshape to a smaller scope, or stop. Underneath the decision, list the questions you knowingly proceeded without and name who owns each one, because proceeding on an unanswered question is a decision someone owns, not an omission nobody noticed. The 500-team target operating model James Rooney co-led at HSBC Global Payment Solutions was piloted rather than presented. A report is where findings go to be admired; a decision is something the approval gate can act on. If the research cannot yet produce one, say that plainly and move the epic out of the quarter rather than letting it sit in research looking busy.
Continuous discovery vs a two-week timebox: which is better?
They do different jobs, and the timeboxed pass is the one an approval gate can act on. A continuous habit keeps a research cadence running whatever the roadmap is doing, which is useful for building a picture of your users over time. The timeboxed pass runs inside one epic's research phase, is capped at two weeks and at one week for anything under roughly £100k of planned value, and ends by confirming, changing or zeroing that epic's planned share of the outcome's currency target. Its trigger is narrower too. If no finding would change the number, the scope or the decision to build, you skip it and start building.
Is a fake door test enough evidence to price a feature?
It is good evidence of interest and weak evidence of value, so discount it hard before it reaches the number. In one worked example a save-as-template button shown to 900 accounts for two weeks was clicked by 31 per cent, and the re-price used 20 per cent sustained use instead, because a click costs nothing and changing how you order every week costs something. Pair it with behaviour you already hold, such as order history or funnel drop-off, then apply the realisation factor your own delivery history justifies. A click-through rate carried straight into a currency figure is the quickest route to arriving short in month three.
What are the steps in a two-week discovery sprint?
Eight steps inside the box, plus a rematch at outcome validation later. Write three figures at the top of the research file: the outcome's currency target, the epic's planned share of it, and the threshold at which your decision changes. State the hypothesis with a measurable kill condition and have product and engineering sign both off before fieldwork starts. Put the end date on the epic, recruit interviewees on day one, and run at least one method that returns a number rather than only conversations. Convert every source into markdown, run the contradiction pass and check its references yourself, re-price the epic, then write the decision on it. The worked example lands in nine working days of the two-week box.
We only have an idea, not an epic yet. Is that discovery?
No, that is outcome shaping, and it is a different question. Discovery tests whether an epic's planned share of a priced outcome survives contact with reality, so it cannot start until you can write three figures at the top of the research file: the outcome's currency target, the epic's planned share of it, and the threshold at which your decision changes. If nothing has been priced, the question in front of you is whether the outcome can carry a credible number at all. Set the outcome first, break it into epics, then run discovery on the ones whose share rests on something you cannot yet evidence. Fieldwork first produces interesting findings and no decision.
Does discovery change if AI is writing the code?
Yes, and it changes the fieldwork rather than just the write-up. In an AI-augmented team, where a developer writes the code and the model assists, this guide runs as written. In an AI-native team there is no child story to hold the detail, so discovery has to produce the key user journeys and the test requirements that go into the outcome ticket itself. That means capturing the sequence people actually follow, the edge cases and the failure states, not sentiment about whether they would like the feature. It raises the bar on the evidence corpus, because the same markdown files are what the model builds from. Tenhaw's engineering handbook targets that reader, 72 rules on GitHub written to be enforced by an agent.
Why capture the baseline during discovery rather than at launch?
Because you are already in the data, and nobody goes back for it later. Discovery predicts a two-point lift, the epic ships, and if nobody recorded what the rate was the week before, validation four months on has a number with nothing to compare it against and the epic gets closed on an argument instead of evidence. Record it in the same pass that produces the denominator and the adoption rate, and write it into the epic beside the revised currency figure. The worked example on this page notes a 21 per cent repeat order rate over the prior 90 days, because a lift needs a starting point.
How do you work out how many customers a feature would reach?
Pull the denominator before you count anything else. Query behaviour you already hold, order history, funnel drop-off, ticket volumes or cancellation reasons, and start from the total population the epic could ever touch, so every percentage after it has a ceiling. In one worked example that is 4,200 active trade accounts, of which 1,600, or 38 per cent, place orders where at least 80 per cent of lines match a previous order. The 1,600 is the reachable group, not the value. Adoption, the share of that group whose behaviour actually changes, and value per event come after it. Skip the denominator and you finish with a rate and no idea what it applies to.
What do you do when the research says nothing either way?
Name the silence and then decide what to do about it in the open. A corpus that neither supports nor contradicts part of your number is telling you which assumption you never actually tested, which is more useful than a weak signal you might have leaned on. Two honest options fit inside the timebox. Spend what is left of it on that single gap, or re-price on what you do have and carry the question forward with an owner beside it. In the worked example the team proceeded without knowing whether templates cannibalise higher-margin phone orders through the sales desk, and named the commercial lead as the person who owns it.
If the sources do not answer it, a call will.
Talk it throughHow to write a chapter (a sub-task of a story)
Answered on How to write a chapter (a sub-task of a story), and rendered here in the same words.
Read the page these answers live on →
Can a chapter have chapters of its own?
No. Chapters are the last level. A chapter you cannot finish in two days is telling you the seam is in the wrong place, or that the parent should have been more than one story. Move the seam first, and if that does not work, stop and take the story back to product rather than inventing a level below.
Do chapters get story points?
No. Points stay on the parent story, and that story keeps the size it was given before the build started. Size chapters in days, for sequencing only, and never sum those days back onto the story. The moment chapter sizes feed your velocity, the throughput data that produces your p50 and p85 stops being comparable across quarters.
Does product need to approve chapters?
No. The developer creates and closes them, and they should still be visible on the board so the chain from chapter to story to epic to outcome holds. If product is being asked to make a call on a chapter, whatever is under discussion is user-visible and should have been raised as a story.
What if a chapter turns out to be user-visible after all?
Convert it. Raise it as a story under the same epic, get product approval, and let product decide whether it belongs in this quarter or the next. Do not ship a user-visible change under a chapter because the branch and the flag happen to be there already, which is exactly how work disappears from the board product reads.
What is a chapter in The Tenhaw Way?
A chapter is The Tenhaw Way's name for a sub-task, the fourth and optional level of the breakdown, below the outcome, the epic and the story. A developer creates chapters mid-build, once a story turns out to be bigger than refinement thought. A chapter carries no currency value, no product approval and no user-visible slice, because the story above it holds all three. A good one names one technical deliverable, merges on its own behind the parent story's flag, and takes half a day to two days, about four per story at most. Chapters exist so a big story stays visible and reviewable, not so it stays hidden. Tenhaw publishes the whole method, these rules included, free to adopt without hiring us.
How do you split a story that turns out to be bigger than estimated?
Split it mid-build along the system's existing seams: a schema change, a service boundary, a contract between two components, a feature flag. Each sub-task should name one technical deliverable, merge on its own behind the parent story's flag, and take between half a day and two days, with no more than about four in total. Do not nest a level below; if a piece will not fit in two days, move the seam instead. And keep the split invisible to the user. Anything a user could see when shipped alone is a new story that needs product approval, not a sub-task.
What is the difference between a sub-task and a story?
A story changes something a user can see and needs product approval; a sub-task is invisible to the user and exists only so the story's acceptance criteria can pass. Two questions settle it. If the work shipped alone, could a user do something new? If yes it is a story. A new screen, a new field or a new email all qualify, however small. Do the parent story's acceptance criteria fail without it? If they pass regardless, it is not a sub-task either, it is tech debt or a bug and belongs against those budget epics. A genuine sub-task, like a migration or an idempotency key, is invisible to the user and something the criteria cannot pass without.
Should sub-tasks be created during planning or during development?
During development. Create sub-tasks when the code has told you something planning could not: a service returns stale data, a migration has to run before an endpoint changes, a dependency needs versioning first. If you can list the sub-tasks before anyone opens the editor, you have not found sub-tasks, you have found a story that is too big. Take it back, split it into stories that each change something a user can see, and get product approval on each one. Sub-tasks written in advance read as diligence, but they are a sizing failure in disguise, and they become a private backlog nobody outside the team ever reads.
Should developers work on sub-tasks in parallel or in sequence?
In sequence, with exactly one in progress at a time. Order them by what unblocks what, putting schema and contract changes first so later work builds on a settled shape. Three sub-tasks running in parallel across two developers finishes the story later than one developer taking them in order, because the merges fight each other and the review queue backs up. Each sub-task should be one pull request that merges on its own behind the parent story's flag, so a colleague can review it without waiting for the next. If a later one becomes unnecessary once earlier work lands, close it with the reason rather than deleting it.
Should a refactor found mid-build be a sub-task of the story?
Only if the story's acceptance criteria fail without it. A genuine sub-task is work the story cannot pass without, and that is the whole test. A refactor the criteria do not depend on belongs against the quarter's Tech Debt epic instead, raised properly rather than buried under a story because raising it takes longer. Those budget epics exist so the quarter's burn rate stays visible, and every item hidden under a story makes that number a lie. Defects you trip over mid-build follow the same rule. If the story's criteria pass regardless, they are raised against the Bug Budget epic, not slipped in as a sub-task.
Is it OK to split a ticket into backend and frontend sub-tasks?
No, and it is the split that most often defeats the point of splitting. Backend and frontend are two halves of one change. The reviewer cannot judge either half without the other, and nobody reading the board can tell what is actually left. A sub-task has to stand up on its own, one pull request that merges behind the parent story's flag, so cut at a joint the system already has instead. Names like Part 1 of 3 fail the same test, because they describe the order somebody happened to work in rather than what the work is.
What should a sub-task ticket actually contain?
Five fields, and stop, in under two hundred words in total. A title that is a verb plus the thing, such as Make the renewal accept endpoint idempotent. One link to the parent story, nothing else. What changes: the files, services or tables affected. How you will know: the test or check that proves it, an engineering check rather than a product criterion. And out of scope: the neighbouring work you are deliberately not doing here. Anything beyond that is either restating the story above it or hiding a second sub-task inside this one.
Is splitting a story mid-sprint just extra admin?
It costs about an hour, once, on the story that surprised you. Four sub-tasks at most, five short fields each, under two hundred words in total, written at the keyboard by the developer who found the problem. Tenhaw's engineers work that way on client delivery, against a published standard of 72 rules written to be enforced by an agent. Set that against the alternative, where a story sized at two days quietly takes four, one large pull request lands at the end, and until it does nobody outside the team can see what is left or review any of it in pieces. The admin is not the split. It is what you pay when a big story stays hidden instead of staying visible and reviewable.
Is a story finished once all its sub-tasks are closed?
No. A sub-task closes on engineering when its own check passes and a human has reviewed the change. The story moves to in review only once every sub-task is closed and its acceptance criteria are demonstrated end to end, not inferred from three green sub-tasks in a row. That gap is where defects live, because each piece was verified against its own technical check and never against what the user was actually promised. Sub-tasks do not release either, and value monitoring stays at epic level, where the currency target sits.
How often is it normal for a story to split mid-build?
Two or three times a quarter is normal, and it deserves one line at the next fortnightly refinement rather than a post-mortem, covering what the story looked like going in and what it turned out to contain. The pattern is what matters. The same shape every fortnight, always around the same service or the same kind of change, is a sizing signal to act on. Tenhaw watches that rate on client delivery, the same attention to flow that cut delivery lead times by 60% within six months at Globelynx. At that refinement, resist re-pointing a story already in flight to tidy the burn-down, because it corrupts the throughput history your p50 and p85 forecasts are built on.
Why can't a sub-task carry its own value target?
Because the value is already held above it, and putting a number on the sub-task counts it twice. Tenhaw reports the money at outcome and epic level for the same reason. The epic carries the planned share of the outcome in currency and the story carries the acceptance criteria, so a sub-task with its own figure appears a second time when somebody sums the quarter, and the roadmap starts claiming more value than the outcome was ever priced at. Giving one its own product acceptance does the same damage in a different direction. You end up with two competing versions of done, and a developer choosing between them mid-build.
How do we introduce sub-tasks to a team that doesn't use them?
Do not roll them out as a process. A sub-task is what a developer writes at the keyboard, once the code has told them something refinement could not, so the habits matter more than the tooling: create one only mid-build; cap them at about four per story, half a day to two days each; keep exactly one in progress; give each five fields and no currency value; and report the split in a line at the next refinement without re-pointing a story already in flight. Tenhaw calls them chapters and passes habits like these on by pairing with a client's own engineers. The name matters less than the rules. Run it for a quarter and your real sizing problems surface on their own.
Should you split a ticket before handing it to an AI coding agent?
No. An AI-native team, where the model builds and the developer directs and verifies, runs two levels, so there is nothing to split. The outcome ticket is the unit of work and it goes over whole. If you feel the urge to decompose it, read that as a signal about the ticket rather than about the size of the work. It is usually thin, missing a key user journey or a test requirement, and the fix is upstream. Add what is missing, hand the ticket over again, and do not invent a sub-level the mode does not have. Tenhaw staffs that mode as an Agentic Build Team of three practitioners under partner oversight, with James Rooney on every engagement.
If the sources do not answer it, a call will.
Talk it through1424 questions, grouped by subject
Every question answered anywhere on tenhaw.com sits in one of 51 groups. This is one of them.
- Target operating model design18
- The Tenhaw Way18
- The AI-native target operating model18
- Building with AI22
- The AI-native delivery lifecycle19
- The how-to library15
- Outcomes and roadmaps36
- Measuring value18
- Epics and stories54
- Forecasting and dates36
- Running delivery day to day54
- Bugs and root cause36
- Risks and release notes36
All 1424questions, and every group →
Or ask the question directly and skip the categories.
Talk it throughStill have a question?
A 30-minute discovery call with James Rooney. Bring the question this page did not answer. You'll leave with a rough scope whether you engage us or not.
most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes
Calendar not loading? Open it on cal.com or email hello@tenhaw.com.