Pattern guide

The business case for an agentic programme: ROI, payback and what to measure

What the number is actually made of, what a finance function should refuse to accept, and the ways the case falls apart.

ways these programmes stall
6
steps in how we would address it
9
questions answered in full
8
In one paragraph

The return on an agentic programme is made of three different things, and only one of them is money without a further decision being taken: cost avoided, capacity released, and revenue or loss changed. Hours saved is not a saving until a headcount, a contract or a capacity constraint actually moves, which is the single most common defect in an AI business case. A defensible case states each benefit in currency with its assumptions exposed, prices the run cost as well as the build, adjusts explicitly for optimism bias, treats programme duration as a priced risk, and names the person who will be asked in twelve months whether the benefit arrived. Anyone offering you a return figure before they have looked at your process is quoting somebody else's.

James Rooney, Founder

Updated

Our approach

We have not delivered this at enterprise scale, and here is exactly what we are drawing on

We produce an investment case as part of every Agent-Readiness Audit, with the value stated in currency, a confidence attached to each figure, and the arithmetic laid out so your finance function can rework it with their own numbers. What we cannot give you is a realised return from an agentic system we have run, because we have not yet taken one into production for a client. So every figure on this page is either one of our own published prices, which you can check against the rate card, or a third-party finding with its source attached. There is no Tenhaw client return number here and there will not be one until there is a real one. If a supplier shows you theirs, ask whose process it came from, who verified it, and what the denominator was.

So what should you do about it

Build the cost side before the benefit side, because it is the half you can actually know. The build price is quotable from published rates, and ours are on the pricing page. The run cost is estimable per workflow before anything is committed, and it is the line most cases are missing entirely. Only then argue about the benefit, and argue about it with the process owner rather than with a supplier, because the supplier does not know whether your headcount will actually change. If you want the whole thing done properly that is what the audit is: six to eight weeks, a fixed fee of £30,000 to £90,000, a costed plan in three-month increments capped at twelve months, a do-not-do list, and an investment case your board can act on. If you already have a case and want it stress-tested rather than written, say so on the call, because that is a shorter and considerably cheaper conversation.

Why this matters now

The UK government's own appraisal guidance is the most useful public document a finance director can read on this, and it is free. The Green Book requires appraisers to account for optimism bias, which it defines as the proven tendency for appraisals to be over-optimistic about key assumptions, by increasing estimates of cost and duration and decreasing estimates of benefit. It sets a business case out as five interconnected perspectives, strategic, economic, commercial, financial and management, and it is explicit that these are not five separate documents. And it puts benefits realisation, the plan for monitoring costs and actually realising the stated benefits, inside the management case rather than leaving it as something that happens afterwards. Almost no AI business case we have been shown does any of those three things.

Ask a technical questionanswers from all 12 guides
Ask anything technical about the business case for an agentic programme: roi, payback and what to measure and I will answer from this guide, and tell you first whether this is work we have delivered or an approach we would be taking.

Prefer to talk it through? Ask us on a discovery call →

Diagnosis

Where these programmes actually stall

Not the risks a vendor lists. The ones that stop the work.

01

The case is built on hours saved, and hours saved are not money

An hours figure is a statement about effort, not about cash, and it becomes money only when a headcount changes, a contract is renegotiated, or a capacity constraint that was costing you revenue is released. A concrete example from our own history: the AI Voice Insights proof of concept our founder led at HSBC was projected to remove more than 1.5 million hours of manual administration a year. That figure is a projection, it was never realised, and the work was a proof of concept rather than a rollout. It is exactly the shape of number a finance function should refuse to accept on its own, including from us. The follow-up question is the whole job: whose budget line moves, by how much, and in which quarter.

02

Optimism bias is in the numbers and not in the arithmetic

Every estimate in an AI case is made by people who want the programme to happen, which is normal and is precisely why public-sector appraisal guidance requires an explicit adjustment for it. The Green Book defines optimism bias as the proven tendency for appraisals to be over-optimistic about key assumptions, and requires appraisers to increase their estimates of cost and duration and reduce their estimates of benefit. A case with no such adjustment is not neutral, it is optimistic by default, and the reviewers who eventually find that out will discount everything else in the document too.

03

The run cost was never estimated, so the payback is measured against half the cost

The build is the visible number and the run is the one that decides whether the thing survives its second year: per-run inference and retrieval, human review of everything the system routes to a person, monitoring and evaluation, and a full re-evaluation each time a model version changes underneath you. Cases that price only the build produce a payback period that is wrong in the direction the sponsor wanted, and the correction arrives during the year when the pilot is supposed to be scaling.

04

It was priced as a build and it lives as an operation

A successful pilot with no named owner for its run budget becomes an orphan: the sponsor who funded an experiment is rarely the person who will carry a live system, its incidents and its audit trail, and in most organisations nobody has been asked to. That is a finance problem before it is a technology problem, because the fix is a line in next year's plan with a name against it, agreed before the build rather than negotiated after the demo.

05

Duration is the largest risk in the case and it is not priced

Analysis of 1,355 public-sector IT projects, averaging $130m and 35 months, found that every additional year of duration added about 4.2 percentage points to expected cost overrun, and that 18% of projects were outliers with cost overruns above 25%. A two-year agentic programme carries that risk whoever delivers it, and that includes us. The implication for the case is not that long programmes are forbidden, it is that duration belongs in the risk line with a number attached rather than in the plan as a neutral fact.

06

Nobody agreed who books the benefit

A benefit with no owner does not appear in anybody's budget, so it is never checked, and the programme's actual result is decided by whoever tells the story most confidently at the end. Name the person whose numbers move, get them to agree the measurement and the date before the build starts, and accept that a benefit nobody is willing to own is usually a benefit nobody believes in.

Prescription

How we would address it

Grounded in The Tenhaw Way and the engagements written up in our case studies.

  1. 01

    Start with one process and its cost in currency, before any technology

    Volume, elapsed time, the people involved, the error and rework rate, and what the whole thing costs to run today. That is a fortnight of work and it is the foundation of everything else, because a benefit is a difference between two numbers and most cases only ever produce the second one. It also frequently changes which process you would attack, which is worth knowing before you fund the wrong one.

  2. 02

    Split the benefit into three kinds and treat only one of them as money

    Cost avoided is money when a headcount, a contract or a licence actually changes, and not before. Capacity released is money only if the released capacity is redeployed to something with a value attached, which is a decision somebody has to take and record. Revenue or loss changed, better conversion, fewer claims leaking, lower fraud losses, is the strongest kind and the hardest to attribute. Label every line as one of the three and the case becomes readable by a sceptic, which is the only kind of reader worth writing for.

  3. 03

    Attach a confidence to every line, and hand the model to finance

    Each figure carries the assumption it rests on and how confident we are in it, and the whole thing is handed over in a form your finance function can rework with their own numbers rather than as a PDF. This is what the investment case in our audit deliverable is: the value in currency, the assumptions exposed, the confidence stated, the arithmetic open. A case nobody can rework is a case nobody can check.

  4. 04

    Cost the run before the build is approved

    An estimated run cost per candidate workflow, produced during the audit rather than discovered in year two, covering inference and retrieval per run, the human review the routing design actually implies, monitoring, and re-evaluation on model change. Where that number is uncomfortable it usually redirects the design rather than killing the case, which is exactly what it is for.

  5. 05

    Adjust for optimism bias on the record, and say by how much

    Increase the cost and duration estimates, reduce the benefit estimates, and state the adjustment as a visible line rather than quietly baking it in. Two reasons. It is what serious appraisal guidance requires, so your reviewers recognise it. And a case that has already discounted itself is far harder to argue with than one that has not, which is a negotiating advantage rather than a concession.

  6. 06

    Keep the unit of work small and the clock short, because duration is priced risk

    Sequence the programme as short fixed-price stages with real decision points between them, rather than as a two-year commitment with milestones. Ours are built that way for this reason: an audit fixed over six to eight weeks, a proof of concept fixed over two to four, and productionisation scoped at four to six weeks with a dedicated team. Whoever you buy from, a case built on stages you can stop is worth more than a case built on a plan you have to believe.

  7. 07

    Price the first year against a published rate card rather than a range

    One worked sum, on our own published prices, because a finance director reading this deserves at least one. An Agent-Readiness Audit is a fixed £30,000 to £90,000 over six to eight weeks. An Agentic Proof of Concept is a fixed £20,000 to £55,000 over two to four weeks. Buying both, a diagnosis and one working thing, is therefore £50,000 to £145,000 before any production commitment is made. If the answer is then a build team of three at £70,000 to £85,000 a month, that is £840,000 to £1.02m across a full year, and productionising a single proof of concept is scoped at four to six weeks rather than a year of that. Those are published numbers and the arithmetic is checkable. What none of them tells you is the benefit, which is the entire argument for putting the audit first: it exists to put a defensible figure on the other side of the sum before the larger commitment is made.

  8. 08

    Name the benefit owner and the measurement date before the build starts

    One named person whose numbers move, one agreed measurement, one date in the diary, all fixed before anyone writes code. This is the same discipline as putting a named owner on the release gate, applied to the money instead of the risk, and it is what separates a programme that can prove its result from one that has to assert it.

  9. 09

    Write the do-not-do list and count it as return

    The things not worth doing here, and why, are usually the most valuable page in an audit readout, and they belong in the business case rather than in an appendix. A programme that avoids a seven-figure commitment to the wrong workflow has produced a return that is real, immediate and considerably more certain than any of the benefits above. Finance functions understand avoided spend better than anyone else in the building, so write it in their language.

An Agent-Readiness Audit is a fixed £30,000 to £90,000 over six to eight weeks.
Where this usually starts

Agent-Readiness Audit

Where agents add value, where they don't, and what to do first. Fixed price · £30k–£90k · 6–8 weeks.

What that engagement covers

The business case for an agentic programme: ROI, payback and what to measure: your questions

What is the ROI of an agentic AI programme?

Nobody can tell you without looking at your process, and a supplier who quotes a figure before doing so is quoting somebody else's business. What can be said generally is the shape of the answer. The return is made of three components: cost avoided, which becomes money only when a headcount, contract or licence actually changes; capacity released, which becomes money only if someone decides where the released capacity goes; and revenue or loss changed, which is the strongest and the hardest to attribute. Against that sits a build cost, which is quotable from published rates, and a run cost, which is per-run inference and retrieval, human review, monitoring, and re-evaluation on every model change. A case that names all three benefit types, prices both cost types, and adjusts for optimism bias is defensible. A case built on hours saved is not.

How do you build a business case for AI agents?

Six steps and none of them starts with technology. Measure the current process in currency: volume, elapsed time, people, error and rework, total cost to run today. Estimate the future state and split the difference into cost avoided, capacity released, and revenue or loss changed, labelling each. Attach a confidence and the underlying assumption to every line. Price the build from a published rate card and the run per workflow, before approval rather than after. Adjust explicitly for optimism bias by increasing costs and duration and reducing benefits, and show the adjustment. Then name the person who owns the benefit and the date it gets measured. Our Agent-Readiness Audit produces exactly this, over six to eight weeks at a fixed £30,000 to £90,000, alongside the sequenced costed plan and the do-not-do list.

What is the payback period on an agentic proof of concept?

A proof of concept is bought to remove uncertainty rather than to pay back, and treating it as an investment with a return is how organisations end up defending a two-week build as though it were a product. The arithmetic that is checkable is the cost side: ours is a fixed £20,000 to £55,000 over two to four weeks, and you keep the working code and the requirement corpus whichever way the decision goes. What it buys is a measured answer on whether the workflow can be done at all, an accuracy and confidence read on your real data, and an estimated run cost, which together determine whether the far larger production commitment is worth making. If the answer is no, the proof of concept has paid for itself several times over by preventing that commitment, and that is the return worth counting.

What does it cost to run an agentic system once it is built?

Any invented number would describe our workloads rather than yours. The lines to estimate are consistent though, and you can build the figure yourself. Per-run model cost, driven by how many turns the agent takes and how many tokens go into its context on each one. Retrieval and index maintenance, including re-embedding whenever the corpus or the embedding model changes. Human review, which is set by your routing design rather than by the technology, and is usually the largest line in a regulated process. Monitoring and observability. And re-evaluation on every model version change, which is a recurring cost most plans treat as a one-off. Our audits produce an estimated run cost per candidate workflow before anything is committed to, precisely because this is the number that decides between designs.

Should we count hours saved as savings?

Not as savings, no. Count them as a measurement of the change and then do the second piece of work, which is establishing what happens to those hours. If a fixed-term contract ends or a vacancy goes unfilled, that is cash and it belongs in the case with the date it lands. If the time returns to people who then do something else, it is capacity released, and it is only worth money once someone decides what that something else is and attaches a value to it. If nothing changes, the hours are real and the money is not. Writing that distinction into the case protects you twice: it survives scrutiny, and it stops the programme being judged in a year against a saving that was never going to appear in a ledger.

How do you stop the business case being over-optimistic?

By adjusting for it deliberately and visibly, which is what public-sector appraisal guidance has required for years. The Green Book defines optimism bias as the proven tendency for appraisals to be over-optimistic about key assumptions and instructs appraisers to increase their estimates of cost and duration and decrease their estimates of benefit. Three practical habits follow. Show the adjustment as its own line, so reviewers can see it rather than hunt for it. Have someone outside the programme write the pessimistic case, not the sponsor. And treat duration as a priced risk rather than a scheduling detail, because the published evidence on IT programmes is that cost risk rises with every additional year of runtime.

What should finance measure after go-live?

Four things, monthly, and only the first is about the technology. Run cost per unit of work, against the estimate made before approval, because that is where an unpleasant surprise shows up first. Volume actually processed by the system rather than volume it is capable of processing, since adoption is usually the constraint. The exception rate and therefore the human review hours the system is really consuming, which is the line that quietly eats the benefit. And the benefit itself, measured on the agreed date by the named owner, against the figure in the approved case rather than against a revised one. If the fourth measurement has no owner and no date, the first three will be reported and the programme's result will never actually be established.

Who should own benefits realisation?

The person whose budget or performance numbers move, which is almost never the person who sponsored the technology. Get them to agree the measurement and the date before the build starts, because agreeing it afterwards is a negotiation and agreeing it beforehand is a specification. The same person should also own the run budget, since a benefit owner with no cost exposure has an incentive to be generous. Where no such person exists, that is worth surfacing immediately: a process nobody owns end to end is a programme risk long before it is a measurement problem, and it usually means the first piece of work is establishing the ownership rather than building anything.

Sources

Where the checkable claims came from

Every source below was opened and read before it was attached. Where nothing survived that check, the claim on this page was softened rather than given a plausible-looking link.

  1. 01
    HM Treasury, The Green Book (2026): appraisal and evaluation in central government

    Source of the optimism bias definition and the instruction to increase cost and duration estimates and decrease benefit estimates, of the five case model, and of the requirement that the management case sets out plans for monitoring costs and realising benefits. Written for public-sector appraisal, and the discipline transfers.

  2. 02
    Budzier and Flyvbjerg, Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector (2013)

    1,355 public-sector IT projects, average budget $130m over 35 months. Each additional year of duration added about 4.2 percentage points to average cost risk, and 18% were outliers with cost overruns above 25%.

Talk to us about the business case for an agentic programme: roi, payback and what to measure.

A 30-minute call with James Rooney. We'll tell you honestly which parts of this we have done before and which we would be doing for the first time, and you'll leave with a rough scope either way.

30 minutesWith James personallyNo obligation

Most organisations start with a fixed-price Agent-Readiness Audit · £30k–£90k · 6–8 weeks