How we deliver, in the FAQ

Measuring value, answered in full.

Validating that the outcome actually landed, in the currency the business uses rather than in story points. Written for the person running a quarter rather than for a buyer, and free to use with us or without us.
questions in this group, each answered in full
18
pages the answers are written on, every one linked
1
questions across the whole FAQ
1424

18 questions on measuring value, answered by Tenhaw, a UK AI consultancy and AI delivery partner based in London. Nothing here is a summary: each answer is the exact text from the page that owns it, and every group links back to that page for the context around it.

Elsewhere in the FAQ
18 questions

How to measure the value a project delivered

Answered on How to measure the value a project delivered, and rendered here in the same words.

Read the page these answers live on →

How long should an epic sit in value monitoring?

As long as the money takes, and you decide that when you write the epic rather than when someone asks. Tenhaw agrees an exit date at kickoff, so epics routinely sit in value monitoring long after the team that built them has left, read by the client's own named owner. Divide planned value by a realistic monthly run rate. A £200k epic earning £40k a month needs five months at minimum, plus the adoption ramp. If an epic needs more than about two quarters to prove, schedule interim reads at thirty, sixty and ninety days and record a forecast at each rather than going quiet until the end. Epics belong to one roadmap, but the money does not stop at the quarter boundary.

What if we cannot attribute the value cleanly?

Say so in the epic, take the strongest method you can afford, and label the number with the method that produced it. Holdout first, then staged rollout by cohort or region, then interrupted time series against the baseline trend, then a declared assumption signed off by someone commercial. A weak method that is disclosed is workable, because everyone reading it knows what it is worth. An unlabelled claim built on an assumption is worse than no number, because it gets planned against. Tenhaw reports the two-week proof of concept it built inside a London specialty insurance business as exactly that, never as production. If attribution is impossible in principle, say so before the epic is approved.

Does every epic need this, including tech debt and bugs?

No. The tech debt epic and the bug budget epic that open every roadmap are fixtures rather than priced contributions to an outcome, so validating them in currency invents numbers nobody believes. Track those two on burn rate instead: how much debt was linked and cleared, how many bugs were raised and closed, and whether the trend across quarters is improving or rotting. Everything else in the roadmap carries a currency share of an outcome and goes through the full validation, including epics that are enablers for later work. Price those against the value they unlock, not at zero.

How do we stop this becoming a blame exercise?

The number, not the person, and two habits keep it that way. The decision options include not realised so close it, which makes closing an epic short a normal result of the ritual rather than an escalation. And the calibration factor is an organisational figure rather than a scorecard. Realising 70% of plan is common, and knowing it lets you plan headroom instead of pretending. The failure mode to watch for is the product manager who quietly stops bringing epics to the session. If attendance starts slipping, the honesty has already gone, and no amount of template will bring it back.

What is the difference between monitoring a release and measuring its value?

They are different jobs on different clocks. For seven days after release you are triaging stability: customer impact, support noise, and the continue, watch or rollback call. None of it says whether the money arrived. On day eight the epic moves into value monitoring with a named owner, usually the product manager who priced it, and a first read date. The question changes from stable to earning, read monthly against the baseline until the planned value is confirmed or closed short. Tenhaw reports on that second clock, and what it carries out of Globelynx is a 60% lead-time reduction inside six months, not a quiet launch week. A team that stops at the stability window has checked the plumbing and never read the meter.

How big does a holdout group need to be to prove a conversion lift?

Big enough to detect the effect you priced, and you check that before committing to the method. A 1% conversion lift needs tens of thousands of sessions per arm; a checkout seeing 41,000 sessions a week can detect a priced effect inside three weeks, while a journey with a few hundred sessions a month may never separate signal from noise. If the volume is not there, do not run an underpowered test and hope. Step down to the next strongest method, a staged rollout by region or cohort, or an interrupted time series against the baseline trend, and record the choice in the epic before release, because the method never changes after seeing the result.

Our reporting system only keeps ninety days of history. Can we still baseline?

Yes, if you act the day the epic is written rather than the day it ships. Start the export today and set a weekly snapshot, so history accumulates while the epic moves through design and build. You want four quarters of raw weekly numbers pasted into the epic itself, not a dashboard link whose definition will drift. That gives you the level, the week-to-week spread and the seasonal shape, because a 2% lift in November proves nothing if November is always up 2%. A baseline cannot be captured retrospectively, and an epic without one produces an argument rather than a number, so treat the missing export as a blocker.

How do you stop two projects claiming the same benefit?

Keep one conversion sheet per outcome and check across epics that no pound is claimed twice. Double counting happens whenever two pieces of work touch the same journey and two people price them against the same metric without a shared sheet; the symptom is an outcome reporting 140% of target while finance can see revenue is flat. The sheet holds the unit economics in use, average order value, gross margin, cost per support contact, with the source and date of each figure, so every epic converts its metric to currency the same way. Run the cross-check as part of the monthly validation, before the roll-up, so an inflated claim never reaches the outcome's forecast.

Are usage numbers proof that a feature delivered value?

No. Tenhaw contracts against that distinction, and on its monthly engagements a month that delivers no measurable value is reported as a failed month. "Fourteen thousand people used the new flow" is an activity number, offered up because usage data exists on day one while margin data lags by weeks. Usage says the change was found and adopted, which matters, but value is the metric you priced converted into currency: orders completed times margin, contacts avoided times cost per contact, hours redeployed to something named. Write the currency conversion into the epic before release, with the source query and the arithmetic, and the usage number has somewhere to go. Report adoption alongside value if it helps the story, never instead of it.

Should delivered benefits be measured in revenue or margin?

Margin, whenever the outcome is stated in profit. A checkout change that adds £1m of revenue at 41% gross margin has delivered £410k, and quoting the revenue figure overstates the result by more than double. Keep the unit economics on the outcome's single conversion sheet, each figure with its source and date, and refresh them quarterly, re-running last month's numbers when a figure changes. For savings, count only money that leaves the profit and loss, or hours redeployed to something named. Hours saved in the abstract are not money.

What if it is too early to tell whether the value has landed?

Run the session anyway and record a number rather than a comment. The month where nothing has moved yet is exactly the one teams cancel, and skipping it twice is how the ritual quietly stops existing, after which nobody can say six months later what the quarter produced. So take the read, write down what the metric did, and take the decision that fits. Needs longer, with the next read date named, is a perfectly good result, recorded as no movement, next read 3 September. Every live outcome should carry a validation record dated inside the last calendar month, and an early read at least proves the measurement plumbing works.

Do we need someone from finance in the monthly value review?

You need someone commercial who can challenge the conversion, and in most organisations that is finance. On programmes Tenhaw governs, its own workstreams appear in the same pack, to the same standard, as every other supplier's. The session is forty-five minutes on a fixed date each month, one pass across every live outcome, with each outcome's product manager, an engineering lead and that commercial voice. Their job is the arithmetic rather than the delivery: whether the unit economics on the conversion sheet, average order value, gross margin, cost per support contact, are still current, whether a claimed saving is money that genuinely leaves the profit and loss, and whether two epics touching the same journey are counting the same pound twice.

Is a monthly value review worth the overhead on a small team?

Yes. The arithmetic is about three hours a month, of which forty-five minutes is the session itself, one pass covering every live outcome rather than one per team. The rest is pulling and circulating the numbers the day before, a saved query someone runs rather than an analysis project. Without it nobody can say what the quarter actually produced, and the calibration factor that makes next quarter's plan realistic never comes into existence. Tenhaw publishes the method in full and it is free to adopt, so nothing here requires hiring anyone. The overhead scales with what you have shipped, not with headcount, which leaves a small team fewer epics and a shorter session rather than a heavier one.

Why not just do a post-implementation review at the end?

Because by the end, the two things that make the number defensible are already gone. Tenhaw's monthly engagements are cancellable on 30 days' notice either way, which only means something if the value read arrives monthly too. A post-implementation review hunts for a baseline nobody exported and picks its attribution method after the result is visible, which is how a friendlier secondary metric reaches the pack. Monthly validation fixes the metric, the arithmetic and the method before release, then produces a dated decision every month: on track, needs longer with a read date, partially realised so re-forecast, or not realised so close it. You also get to act while the plan can still respond, putting a shortfall on the next quarter's roadmap.

What do you do when a launch lift fades after a few months?

Re-forecast it, record the new number, and keep the epic in value monitoring instead of quietly banking the month one figure. That shape is normal. In this guide's worked example, an invented retailer's checkout epic reads a 0.4 point conversion lift in month one, worth £18k realised and a £216k forecast at close, then settles at 0.3 points by month five as the launch novelty fades, giving £71k realised and a £165k forecast. The decision is partially realised, re-forecast, and the resulting gap against the outcome's target goes on the next quarter's roadmap as a named epic. What you never do is reopen the method because the number softened.

How should unmeasured epics be counted in an outcome's forecast?

At the organisation's calibration factor, never at full plan. Sum realised value to date and forecast at close across every epic under the outcome, and set both against the target, so the roll-up carries three numbers that are not meant to agree. Epics with a measurement use their measurement and never the factor. Epics still unmeasured are forecast at the share of planned value the organisation has actually realised over the last four quarters, so at a factor of 70% a £150k epic carries £105k. That is how an outcome with £500k planned and one epic read can honestly report £71k realised and £375k forecast.

Can we change the metric mid-build if we learn something?

Before release, yes, provided you rewrite the epic properly: the new metric, its source and query, the currency arithmetic, the window, the named puller, and a baseline that fits the new metric rather than the old one. After release, once a result is visible, no. Changing the measure after seeing the result is the pitfall everybody recognises, where the primary metric is flat and a friendlier secondary metric appears in the pack. The same holds for the attribution method, recorded before release and never changed after. If the measure genuinely turns out to be the wrong one, close the epic with the reason written and name attribution as the assumption that broke.

Which parts of the monthly value review can agents take over?

The preparation, and they should. Running the saved queries, pulling the numbers and formatting the pack the day before is exactly the work an agent handles well, and it removes the usual reason a session gets cancelled, which is a room waiting for a dashboard. The four decisions stay with the people in the room. Tenhaw draws the same line in its own work, where the 72 rules in its open-source engineering handbook are enforced by an agent rather than remembered by a human. AI-native teams also ship more per quarter, so epics accumulate in value monitoring faster than forty-five minutes can absorb. Split the session by outcome before you start skipping epics.

All how-to guides

If the sources do not answer it, a call will.

Talk it through
book a call

Still have a question?

A 30-minute discovery call with James Rooney. Bring the question this page did not answer. You'll leave with a rough scope whether you engage us or not.

most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes

// pick a slot · cal.com/tenhaw/professional-servicesLIVE CALENDAR

Calendar not loading? Open it on cal.com or email hello@tenhaw.com.