How to measure the value a project delivered
- steps
- 8
- named failure modes
- 5
- definition-of-done criteria
- 6
How to measure the value a project delivered, in one paragraph
Outcome validation is the monthly check on whether the money you planned has arrived. Every live outcome, every month. Value is measured at the epic, because the epic is what carries a priced share of the outcome's target. Good looks like this: the metric and the attribution method were fixed in the epic before release, the baseline was exported before the change shipped, and every epic in value monitoring has a realised-to-date number and a dated decision from the last calendar month. A shipped epic is not a delivered epic.
That is the procedure. The call is where it meets your delivery structure.
Talk it throughWhat these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.
Run it on a fixed date every month, on every outcome with at least one epic in value monitoring, from the first release under that outcome until the last epic closes. Set the measurement up far earlier, when the epic is written, because a baseline cannot be captured retrospectively.
- Time to run it once
- About three hours a month, of which the session is forty-five minutes
- What you need
- The reporting system the measure is pulled from
- The outcome's recorded baseline
- The delivery tracker
If you do not have those in place, the call is a good place to work out what comes first.
Talk it throughOn this page
Step by step
Each step is deep-linkable, so you can send a colleague the one that is in dispute.
- 1
Define the measure before you build
Measurement belongs in the epic, not a follow-up ticket.
Before it leaves Ready for Dev, write five things in: the metric, the exact source of the number (system, table or saved report, plus the query), the arithmetic that converts that metric into currency, the window the value needs to accumulate over, and one person who can pull it. Written out it reads like this: weekly completed checkouts from the orders table, saved query val_checkout_v1, times £62 average order value at 41% gross margin, read monthly for six months, pulled by the team's analyst. If nobody can name the query, the epic is not ready for dev.
- 2
Capture the baseline before you release
Pull the metric for the four quarters before the change ships and paste the raw weekly numbers into the epic, not a dashboard link whose definition will drift. You want three things from it: the level, the week-to-week spread, and the seasonal shape, because a 2% lift in November proves nothing if November is always up 2%. Note anything else landing in the same window: a price change, a campaign, another epic touching the same journey. If the source system only retains ninety days, start the export today and set a weekly snapshot. An epic without a baseline produces an argument, not a number.
- 3
Choose an attribution method and write it down
Record the strongest method you can afford in the epic before release.
A holdout or A/B split gives you a counterfactual and is the default where traffic allows, but check the volume can detect the effect you priced: a 1% conversion lift needs tens of thousands of sessions per arm. Next best is a staged rollout by region or cohort, released against not-yet-released. Below that, interrupted time series against the baseline trend, credible only when you can name what else moved in the window. Last is a declared assumption, signed off by someone commercial and labelled an assumption for good. Never change the method after seeing the result.
- 4
Move the epic into value monitoring
Live monitoring and value monitoring are different jobs.
For seven days after release you are triaging stability: customer impact, FAQ, support macro, and the continue, watch or rollback call. That tells you nothing about value. On day eight, do four things. Move the epic, not the story, into value monitoring, because value sits on the thing that was priced. Name an owner, usually the product manager who priced it. Record the first read date. Move the outcome into value monitoring too, so it stops being reported as delivered. An epic in value monitoring is not finished work, and a roadmap review that counts it as delivered is reporting fiction.
- 5
Run the validation monthly, every live outcome
Same date each month, forty-five minutes, one session across every live outcome.
In the room: the product manager for each outcome, an engineering lead, and someone commercial who can challenge the conversion. Numbers are pulled and circulated the day before, never queried live: the session that waits for a dashboard is the one that gets cancelled. Epic by epic, state the measured number, realised value to date and forecast at close, then take one of four decisions: on track, needs longer with a next read date, partially realised so re-forecast, or not realised so close it. Record the number, the decision and the date, even in a month where nothing moved.
- 6
Convert to currency the same way every time
Keep one conversion sheet per outcome and show the arithmetic on it.
Write down the unit economics in use, average order value, gross margin, cost per support contact, fully loaded hourly rate, with the source and date of each figure. Use margin, not revenue, whenever the outcome is stated in profit. For cost savings, count only money that leaves the profit and loss or hours redeployed to something named; hours saved in the abstract are not money. Refresh the figures once a quarter and re-run last month's numbers when one changes. Check across epics that no pound is claimed twice, which happens whenever two epics touch the same journey.
- 7
Roll epic actuals up and re-forecast
After the epic pass, sum realised value to date and forecast at close across every epic under the outcome, and set both against the outcome's target. Three numbers, and they are not meant to agree. Apply your calibration factor to anything still unmeasured: if the organisation has realised 70% of planned value across the last four quarters, forecast unmeasured epics at 70% of plan. Measured epics use their measurement, never the factor. When the roll-up falls short, the decision is explicit and taken in the room: add an epic to the next quarter's roadmap, or restate the target with a written reason and the name of whoever agreed it.
- 8
Close honestly, including the epics that missed
An epic leaves value monitoring one of two ways: value confirmed against the method recorded before release, or deliberately closed with a note that the expected value did not land. Name the assumption that broke: demand, adoption, unit economics or attribution. Closing at 40% of plan with the reason written beats an epic left open for a year. When every epic under an outcome is closed, close the outcome with its reason: all linked epics done, value realised, or accepted as not realised. Then do the arithmetic that pays for all of this, realised divided by planned across the last four quarters, and plan next quarter with that factor.
Bring a real piece of work to the call and we will walk it through these.
Talk it throughA worked example: a £200k checkout epic
An invented online retailer prices an epic at £200k: remove the forced account-creation step in checkout. Every number below is invented with it. The measure is weekly completed checkouts from the orders table, saved query val_checkout_v1, converted at £62 average order value and 41% gross margin. The baseline for the four quarters before release is 41,000 sessions a week at 2.9% completion, with November running two points above the rest of the year. The method is a 50/50 holdout, chosen because 41,000 sessions a week can detect the priced effect inside three weeks. Month one: holdout 2.9%, treatment 3.3%, which is 164 incremental orders a week across full traffic and £4,170 of margin a week, so £18k realised. Forecast at close over twelve months, £216k. Decision: on track, next read 14 March. Month five: the lift settles at 0.3 points once the launch novelty fades. Realised to date £71k, forecast at close £165k. Decision: partially realised, re-forecast. The outcome target is £500k across three epics, this one at £200k and two others at £150k each. The roll-up now reads £500k planned, £71k realised, £375k forecast, because the two unmeasured epics are forecast at the organisation's calibration factor of 70%. The £125k gap goes on next quarter's roadmap as a named epic rather than into the following quarter's optimism.
Yours will look different. Thirty minutes is enough to see how.
Talk it throughWhere this goes wrong
- 01
Reporting activity instead of value.
"Fourteen thousand people used the new flow" is a usage number, and it gets offered up because usage data exists on day one while margin data lags by weeks. Write the currency conversion into the epic before release and the usage number has somewhere to go.
- 02
Changing the measure after seeing the result.
The primary metric is flat, so a friendlier secondary metric appears in the pack. It happens whenever the measure was chosen after release rather than fixed in the epic before it.
- 03
No baseline, so validation becomes anecdote.
Nobody exported the pre-change numbers, the source system retains ninety days, and the meeting turns into who remembers what conversion used to be. It is the least recoverable failure on this list.
- 04
Counting the same pound twice.
Two epics touching the same checkout claim the same uplift, the outcome reports 140% of target, and finance can see revenue is flat. It comes from two people pricing epics against one metric without a shared conversion sheet.
- 05
Cancelling the month because it is too early to tell.
Skip validation twice and it stops existing, and six months later nobody can say what the quarter produced. Run it anyway and record "no movement, next read 3 September".
Done means
- Every live outcome has a validation record dated within the last calendar month, containing a number rather than a comment.
- Every epic in value monitoring carries a named measure, a source query someone can run, an exported baseline, an attribution method recorded before release, and a realised-to-date figure.
- Each outcome shows three numbers side by side: planned value, realised to date, and forecast at close, with unmeasured epics forecast at the calibration factor.
- No pound of value appears under two epics, checked against the single conversion sheet for that outcome.
- Every epic that has left value monitoring left with either confirmed value or a written reason the value did not land, naming the assumption that broke.
- The organisation has a calibration factor from the last four quarters of realised against planned, and next quarter's plan is built with it.
If you recognise one of those already happening, that is a good call to have.
Talk it throughNothing about the measurement changes; where it is written does. An AI-native team has no stories, so the measure, the source query, the attribution method and the baseline go into the outcome ticket alongside the key user journeys and the test requirements, and the ticket is not approved without them. That is an advantage, because the unit of work and the unit of value are the same object and nobody has to reconstruct which epic a release belonged to. The risk runs the other way. AI-native teams ship more per quarter, so the number of epics in value monitoring grows faster than the ritual scales and forty-five minutes stops being enough. Split the session by outcome before you start skipping epics. Agents can pull and format the monthly numbers, and should, but the four decisions stay with the people in the room.
The two delivery modes, side by side →Which mode your team is actually in is the first thing we establish on a call.
Talk it throughWhere this sits in a programme
The procedure is the same whatever you are building. These cover what it runs into when the thing being built is agentic.
- →What a defensible AI business case holds
- →Designing the target operating model that makes the fifty-first rollout cheap
If you want this run inside a programme rather than read, that is the conversation.
Talk it throughMore on measuring and managing
These guides are written to be read in order.
02How to run product delivery day to day
A product manager's daily job is sequencing decisions, not status updates. If you are spending all day in chat, something is wrong.
← All seventeen guidesOr skip ahead and ask which of these your team needs first.
Talk it throughWant help installing this?
These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.
most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes
Calendar not loading? Open it on cal.com or email hello@tenhaw.com.
Questions
How long should an epic sit in value monitoring?
As long as the money takes, and you decide that when you write the epic rather than when someone asks. Tenhaw agrees an exit date at kickoff, so epics routinely sit in value monitoring long after the team that built them has left, read by the client's own named owner. Divide planned value by a realistic monthly run rate. A £200k epic earning £40k a month needs five months at minimum, plus the adoption ramp. If an epic needs more than about two quarters to prove, schedule interim reads at thirty, sixty and ninety days and record a forecast at each rather than going quiet until the end. Epics belong to one roadmap, but the money does not stop at the quarter boundary.
What if we cannot attribute the value cleanly?
Say so in the epic, take the strongest method you can afford, and label the number with the method that produced it. Holdout first, then staged rollout by cohort or region, then interrupted time series against the baseline trend, then a declared assumption signed off by someone commercial. A weak method that is disclosed is workable, because everyone reading it knows what it is worth. An unlabelled claim built on an assumption is worse than no number, because it gets planned against. Tenhaw reports the two-week proof of concept it built inside a London specialty insurance business as exactly that, never as production. If attribution is impossible in principle, say so before the epic is approved.
Does every epic need this, including tech debt and bugs?
No. The tech debt epic and the bug budget epic that open every roadmap are fixtures rather than priced contributions to an outcome, so validating them in currency invents numbers nobody believes. Track those two on burn rate instead: how much debt was linked and cleared, how many bugs were raised and closed, and whether the trend across quarters is improving or rotting. Everything else in the roadmap carries a currency share of an outcome and goes through the full validation, including epics that are enablers for later work. Price those against the value they unlock, not at zero.
How do we stop this becoming a blame exercise?
The number, not the person, and two habits keep it that way. The decision options include not realised so close it, which makes closing an epic short a normal result of the ritual rather than an escalation. And the calibration factor is an organisational figure rather than a scorecard. Realising 70% of plan is common, and knowing it lets you plan headroom instead of pretending. The failure mode to watch for is the product manager who quietly stops bringing epics to the session. If attendance starts slipping, the honesty has already gone, and no amount of template will bring it back.
What is the difference between monitoring a release and measuring its value?
They are different jobs on different clocks. For seven days after release you are triaging stability: customer impact, support noise, and the continue, watch or rollback call. None of it says whether the money arrived. On day eight the epic moves into value monitoring with a named owner, usually the product manager who priced it, and a first read date. The question changes from stable to earning, read monthly against the baseline until the planned value is confirmed or closed short. Tenhaw reports on that second clock, and what it carries out of Globelynx is a 60% lead-time reduction inside six months, not a quiet launch week. A team that stops at the stability window has checked the plumbing and never read the meter.
How big does a holdout group need to be to prove a conversion lift?
Big enough to detect the effect you priced, and you check that before committing to the method. A 1% conversion lift needs tens of thousands of sessions per arm; a checkout seeing 41,000 sessions a week can detect a priced effect inside three weeks, while a journey with a few hundred sessions a month may never separate signal from noise. If the volume is not there, do not run an underpowered test and hope. Step down to the next strongest method, a staged rollout by region or cohort, or an interrupted time series against the baseline trend, and record the choice in the epic before release, because the method never changes after seeing the result.
Our reporting system only keeps ninety days of history. Can we still baseline?
Yes, if you act the day the epic is written rather than the day it ships. Start the export today and set a weekly snapshot, so history accumulates while the epic moves through design and build. You want four quarters of raw weekly numbers pasted into the epic itself, not a dashboard link whose definition will drift. That gives you the level, the week-to-week spread and the seasonal shape, because a 2% lift in November proves nothing if November is always up 2%. A baseline cannot be captured retrospectively, and an epic without one produces an argument rather than a number, so treat the missing export as a blocker.
How do you stop two projects claiming the same benefit?
Keep one conversion sheet per outcome and check across epics that no pound is claimed twice. Double counting happens whenever two pieces of work touch the same journey and two people price them against the same metric without a shared sheet; the symptom is an outcome reporting 140% of target while finance can see revenue is flat. The sheet holds the unit economics in use, average order value, gross margin, cost per support contact, with the source and date of each figure, so every epic converts its metric to currency the same way. Run the cross-check as part of the monthly validation, before the roll-up, so an inflated claim never reaches the outcome's forecast.
Are usage numbers proof that a feature delivered value?
No. Tenhaw contracts against that distinction, and on its monthly engagements a month that delivers no measurable value is reported as a failed month. "Fourteen thousand people used the new flow" is an activity number, offered up because usage data exists on day one while margin data lags by weeks. Usage says the change was found and adopted, which matters, but value is the metric you priced converted into currency: orders completed times margin, contacts avoided times cost per contact, hours redeployed to something named. Write the currency conversion into the epic before release, with the source query and the arithmetic, and the usage number has somewhere to go. Report adoption alongside value if it helps the story, never instead of it.
Should delivered benefits be measured in revenue or margin?
Margin, whenever the outcome is stated in profit. A checkout change that adds £1m of revenue at 41% gross margin has delivered £410k, and quoting the revenue figure overstates the result by more than double. Keep the unit economics on the outcome's single conversion sheet, each figure with its source and date, and refresh them quarterly, re-running last month's numbers when a figure changes. For savings, count only money that leaves the profit and loss, or hours redeployed to something named. Hours saved in the abstract are not money.
What if it is too early to tell whether the value has landed?
Run the session anyway and record a number rather than a comment. The month where nothing has moved yet is exactly the one teams cancel, and skipping it twice is how the ritual quietly stops existing, after which nobody can say six months later what the quarter produced. So take the read, write down what the metric did, and take the decision that fits. Needs longer, with the next read date named, is a perfectly good result, recorded as no movement, next read 3 September. Every live outcome should carry a validation record dated inside the last calendar month, and an early read at least proves the measurement plumbing works.
Do we need someone from finance in the monthly value review?
You need someone commercial who can challenge the conversion, and in most organisations that is finance. On programmes Tenhaw governs, its own workstreams appear in the same pack, to the same standard, as every other supplier's. The session is forty-five minutes on a fixed date each month, one pass across every live outcome, with each outcome's product manager, an engineering lead and that commercial voice. Their job is the arithmetic rather than the delivery: whether the unit economics on the conversion sheet, average order value, gross margin, cost per support contact, are still current, whether a claimed saving is money that genuinely leaves the profit and loss, and whether two epics touching the same journey are counting the same pound twice.
Is a monthly value review worth the overhead on a small team?
Yes. The arithmetic is about three hours a month, of which forty-five minutes is the session itself, one pass covering every live outcome rather than one per team. The rest is pulling and circulating the numbers the day before, a saved query someone runs rather than an analysis project. Without it nobody can say what the quarter actually produced, and the calibration factor that makes next quarter's plan realistic never comes into existence. Tenhaw publishes the method in full and it is free to adopt, so nothing here requires hiring anyone. The overhead scales with what you have shipped, not with headcount, which leaves a small team fewer epics and a shorter session rather than a heavier one.
Why not just do a post-implementation review at the end?
Because by the end, the two things that make the number defensible are already gone. Tenhaw's monthly engagements are cancellable on 30 days' notice either way, which only means something if the value read arrives monthly too. A post-implementation review hunts for a baseline nobody exported and picks its attribution method after the result is visible, which is how a friendlier secondary metric reaches the pack. Monthly validation fixes the metric, the arithmetic and the method before release, then produces a dated decision every month: on track, needs longer with a read date, partially realised so re-forecast, or not realised so close it. You also get to act while the plan can still respond, putting a shortfall on the next quarter's roadmap.
What do you do when a launch lift fades after a few months?
Re-forecast it, record the new number, and keep the epic in value monitoring instead of quietly banking the month one figure. That shape is normal. In this guide's worked example, an invented retailer's checkout epic reads a 0.4 point conversion lift in month one, worth £18k realised and a £216k forecast at close, then settles at 0.3 points by month five as the launch novelty fades, giving £71k realised and a £165k forecast. The decision is partially realised, re-forecast, and the resulting gap against the outcome's target goes on the next quarter's roadmap as a named epic. What you never do is reopen the method because the number softened.
How should unmeasured epics be counted in an outcome's forecast?
At the organisation's calibration factor, never at full plan. Sum realised value to date and forecast at close across every epic under the outcome, and set both against the target, so the roll-up carries three numbers that are not meant to agree. Epics with a measurement use their measurement and never the factor. Epics still unmeasured are forecast at the share of planned value the organisation has actually realised over the last four quarters, so at a factor of 70% a £150k epic carries £105k. That is how an outcome with £500k planned and one epic read can honestly report £71k realised and £375k forecast.
Can we change the metric mid-build if we learn something?
Before release, yes, provided you rewrite the epic properly: the new metric, its source and query, the currency arithmetic, the window, the named puller, and a baseline that fits the new metric rather than the old one. After release, once a result is visible, no. Changing the measure after seeing the result is the pitfall everybody recognises, where the primary metric is flat and a friendlier secondary metric appears in the pack. The same holds for the attribution method, recorded before release and never changed after. If the measure genuinely turns out to be the wrong one, close the epic with the reason written and name attribution as the assumption that broke.
Which parts of the monthly value review can agents take over?
The preparation, and they should. Running the saved queries, pulling the numbers and formatting the pack the day before is exactly the work an agent handles well, and it removes the usual reason a session gets cancelled, which is a room waiting for a dashboard. The four decisions stay with the people in the room. Tenhaw draws the same line in its own work, where the 72 rules in its open-source engineering handbook are enforced by an agent rather than remembered by a human. AI-native teams also ship more per quarter, so epics accumulate in value monitoring faster than forty-five minutes can absorb. Split the session by outcome before you start skipping epics.