How to measure value (outcome validation)
A shipped epic is not a delivered epic. Validation is what closes the loop between \"we built it\" and \"it worked\".
- steps
- 8
- named failure modes
- 5
- definition-of-done criteria
- 6
How to measure value (outcome validation), in one paragraph
Outcome validation is the monthly check on whether the money you planned has arrived. Every live outcome, every month. Value is measured at the epic, because the epic is what carries a priced share of the outcome's target. Good looks like this: the metric and the attribution method were fixed in the epic before release, the baseline was exported before the change shipped, and every epic in value monitoring has a realised-to-date number and a dated decision from the last calendar month. A shipped epic is not a delivered epic.
What these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.
Run it on a fixed date every month, on every outcome with at least one epic in value monitoring, from the first release under that outcome until the last epic closes. Set the measurement up far earlier, when the epic is written, because a baseline cannot be captured retrospectively.
- Time to run it once
- About three hours a month, of which the session is forty-five minutes
- What you need
- The reporting system the measure is pulled from
- The outcome's recorded baseline
- The delivery tracker
Step by step
Each step is deep-linkable, so you can send a colleague the one that is in dispute.
- 1
Define the measure before you build
Measurement belongs in the epic, not a follow-up ticket.
Before it leaves Ready for Dev, write five things in: the metric, the exact source of the number (system, table or saved report, plus the query), the arithmetic that converts that metric into currency, the window the value needs to accumulate over, and one person who can pull it. Written out it reads like this: weekly completed checkouts from the orders table, saved query val_checkout_v1, times £62 average order value at 41% gross margin, read monthly for six months, pulled by the team's analyst. If nobody can name the query, the epic is not ready for dev.
- 2
Capture the baseline before you release
Pull the metric for the four quarters before the change ships and paste the raw weekly numbers into the epic, not a dashboard link whose definition will drift. You want three things from it: the level, the week-to-week spread, and the seasonal shape, because a 2% lift in November proves nothing if November is always up 2%. Note anything else landing in the same window: a price change, a campaign, another epic touching the same journey. If the source system only retains ninety days, start the export today and set a weekly snapshot. An epic without a baseline produces an argument, not a number.
- 3
Choose an attribution method and write it down
Record the strongest method you can afford in the epic before release.
A holdout or A/B split gives you a counterfactual and is the default where traffic allows, but check the volume can detect the effect you priced: a 1% conversion lift needs tens of thousands of sessions per arm. Next best is a staged rollout by region or cohort, released against not-yet-released. Below that, interrupted time series against the baseline trend, credible only when you can name what else moved in the window. Last is a declared assumption, signed off by someone commercial and labelled an assumption for good. Never change the method after seeing the result.
- 4
Move the epic into value monitoring
Live monitoring and value monitoring are different jobs.
For seven days after release you are triaging stability: customer impact, FAQ, support macro, and the continue, watch or rollback call. That tells you nothing about value. On day eight, do four things. Move the epic, not the story, into value monitoring, because value sits on the thing that was priced. Name an owner, usually the product manager who priced it. Record the first read date. Move the outcome into value monitoring too, so it stops being reported as delivered. An epic in value monitoring is not finished work, and a roadmap review that counts it as delivered is reporting fiction.
- 5
Run the validation monthly, every live outcome
Same date each month, forty-five minutes, one session across every live outcome.
In the room: the product manager for each outcome, an engineering lead, and someone commercial who can challenge the conversion. Numbers are pulled and circulated the day before, never queried live: the session that waits for a dashboard is the one that gets cancelled. Epic by epic, state the measured number, realised value to date and forecast at close, then take one of four decisions: on track, needs longer with a next read date, partially realised so re-forecast, or not realised so close it. Record the number, the decision and the date, even in a month where nothing moved.
- 6
Convert to currency the same way every time
Keep one conversion sheet per outcome and show the arithmetic on it.
Write down the unit economics in use, average order value, gross margin, cost per support contact, fully loaded hourly rate, with the source and date of each figure. Use margin, not revenue, whenever the outcome is stated in profit. For cost savings, count only money that leaves the profit and loss or hours redeployed to something named; hours saved in the abstract are not money. Refresh the figures once a quarter and re-run last month's numbers when one changes. Check across epics that no pound is claimed twice, which happens whenever two epics touch the same journey.
- 7
Roll epic actuals up and re-forecast
After the epic pass, sum realised value to date and forecast at close across every epic under the outcome, and set both against the outcome's target. Three numbers, and they are not meant to agree. Apply your calibration factor to anything still unmeasured: if the organisation has realised 70% of planned value across the last four quarters, forecast unmeasured epics at 70% of plan. Measured epics use their measurement, never the factor. When the roll-up falls short, the decision is explicit and taken in the room: add an epic to the next quarter's roadmap, or restate the target with a written reason and the name of whoever agreed it.
- 8
Close honestly, including the epics that missed
An epic leaves value monitoring one of two ways: value confirmed against the method recorded before release, or deliberately closed with a note that the expected value did not land. Name the assumption that broke: demand, adoption, unit economics or attribution. Closing at 40% of plan with the reason written beats an epic left open for a year. When every epic under an outcome is closed, close the outcome with its reason: all linked epics done, value realised, or accepted as not realised. Then do the arithmetic that pays for all of this, realised divided by planned across the last four quarters, and plan next quarter with that factor.
A worked example: a £200k checkout epic
An online retailer prices an epic at £200k: remove the forced account-creation step in checkout. The measure is weekly completed checkouts from the orders table, saved query val_checkout_v1, converted at £62 average order value and 41% gross margin. The baseline for the four quarters before release is 41,000 sessions a week at 2.9% completion, with November running two points above the rest of the year. The method is a 50/50 holdout, chosen because 41,000 sessions a week can detect the priced effect inside three weeks. Month one: holdout 2.9%, treatment 3.3%, which is 164 incremental orders a week across full traffic and £4,170 of margin a week, so £18k realised. Forecast at close over twelve months, £216k. Decision: on track, next read 14 March. Month five: the lift settles at 0.3 points once the launch novelty fades. Realised to date £71k, forecast at close £165k. Decision: partially realised, re-forecast. The outcome target is £500k across three epics, this one at £200k and two others at £150k each. The roll-up now reads £500k planned, £71k realised, £375k forecast, because the two unmeasured epics are forecast at the organisation's calibration factor of 70%. The £125k gap goes on next quarter's roadmap as a named epic rather than into the following quarter's optimism.
Where this goes wrong
- 01
Reporting activity instead of value.
"Fourteen thousand people used the new flow" is a usage number, and it gets offered up because usage data exists on day one while margin data lags by weeks. Write the currency conversion into the epic before release and the usage number has somewhere to go.
- 02
Changing the measure after seeing the result.
The primary metric is flat, so a friendlier secondary metric appears in the pack. It happens whenever the measure was chosen after release rather than fixed in the epic before it.
- 03
No baseline, so validation becomes anecdote.
Nobody exported the pre-change numbers, the source system retains ninety days, and the meeting turns into who remembers what conversion used to be. It is the least recoverable failure on this list.
- 04
Counting the same pound twice.
Two epics touching the same checkout claim the same uplift, the outcome reports 140% of target, and finance can see revenue is flat. It comes from two people pricing epics against one metric without a shared conversion sheet.
- 05
Cancelling the month because it is too early to tell.
Skip validation twice and it stops existing, and six months later nobody can say what the quarter produced. Run it anyway and record "no movement, next read 3 September".
Done means
- Every live outcome has a validation record dated within the last calendar month, containing a number rather than a comment.
- Every epic in value monitoring carries a named measure, a source query someone can run, an exported baseline, an attribution method recorded before release, and a realised-to-date figure.
- Each outcome shows three numbers side by side: planned value, realised to date, and forecast at close, with unmeasured epics forecast at the calibration factor.
- No pound of value appears under two epics, checked against the single conversion sheet for that outcome.
- Every epic that has left value monitoring left with either confirmed value or a written reason the value did not land, naming the assumption that broke.
- The organisation has a calibration factor from the last four quarters of realised against planned, and next quarter's plan is built with it.
Nothing about the measurement changes; where it is written does. An AI-native team has no stories, so the measure, the source query, the attribution method and the baseline go into the outcome ticket alongside the key user journeys and the test requirements, and the ticket is not approved without them. That is an advantage, because the unit of work and the unit of value are the same object and nobody has to reconstruct which epic a release belonged to. The risk runs the other way. AI-native teams ship more per quarter, so the number of epics in value monitoring grows faster than the ritual scales and forty-five minutes stops being enough. Split the session by outcome before you start skipping epics. Agents can pull and format the monthly numbers, and should, but the four decisions stay with the people in the room.
The two delivery modes, side by side →Questions
How long should an epic sit in value monitoring?
As long as the money takes, and you decide that when you write the epic rather than when someone asks. Divide the planned value by a realistic monthly run rate: a £200k epic earning £40k a month needs five months at minimum, plus whatever the adoption ramp adds. If an epic would take more than about two quarters to prove, schedule interim reads at thirty, sixty and ninety days and record a forecast at each one rather than going quiet until the end. The epic will sit in value monitoring long after the quarter it shipped in, which is expected: epics belong to one roadmap, but the money does not stop at the quarter boundary.
What if we cannot attribute the value cleanly?
Say so in the epic, take the strongest method you can afford, and label the number with the method that produced it. Holdout first, then staged rollout by cohort or region, then interrupted time series against the baseline trend, then a declared assumption signed off by someone with commercial ownership. A weak method that is disclosed is workable, because everyone reading the number knows what it is worth. An unlabelled claim built on an assumption is worse than no number, because it gets planned against. If attribution is impossible in principle, say that in the epic before it is approved rather than discovering it in the validation session six months later.
Does every epic need this, including tech debt and bugs?
No. The tech debt epic and the bug budget epic that open every roadmap are fixtures rather than priced contributions to an outcome, so validating them in currency invents numbers nobody believes. Track those two on burn rate instead: how much debt was linked and cleared, how many bugs were raised and closed, and whether the trend across quarters is improving or rotting. Everything else in the roadmap carries a currency share of an outcome and goes through the full validation, including epics that are enablers for later work. Price those against the value they unlock, not at zero.
How do we stop this becoming a blame exercise?
The number, not the person, and two habits keep it that way. The decision options include not realised so close it, which makes closing an epic short a normal result of the ritual rather than an escalation. And the calibration factor is an organisational figure rather than a scorecard: realising 70% of plan is common, and knowing it lets you plan headroom instead of pretending. The failure mode to watch for is the product manager who quietly stops bringing epics to the session. If attendance starts slipping, the honesty has already gone, and no amount of template will bring it back.
More on measuring and managing
These guides are written to be read in order.
02How to manage day-to-day product delivery
A product manager's daily job is sequencing decisions, not status updates. If you are spending all day in chat, something is wrong.
← All seventeen guidesWant help installing this?
These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.
Most organisations start with a fixed-price Agent-Readiness Audit · £30k–£90k · 6–8 weeks