The business case for an agentic programme: ROI, payback and what to measure
What the number is actually made of, what a finance function should refuse to accept, and the ways the case falls apart.
The return on an agentic programme is made of three different things, and only one of them is money without a further decision being taken: cost avoided, capacity released, and revenue or loss changed. Hours saved is not a saving until a headcount, a contract or a capacity constraint actually moves, which is the single most common defect in an AI business case. A defensible case states each benefit in currency with its assumptions exposed, prices the run cost as well as the build, adjusts explicitly for optimism bias, treats programme duration as a priced risk, and names the person who will be asked in twelve months whether the benefit arrived. Anyone offering you a return figure before they have looked at your process is quoting somebody else's.
This is our approach, not a programme we have already run.
We produce an investment case as part of every AI Readiness Audit, with the value stated in currency, a confidence attached to each figure, and the arithmetic laid out so your finance function can rework it with their own numbers. What we cannot give you is a realised return from an agentic system we have run, because we have not yet taken one into production for a client. So every figure on this page is either one of our own published prices, which you can check against the rate card, or a third-party finding with its source attached. There is no Tenhaw client return number here and there will not be one until there is a real one. If a supplier shows you theirs, ask whose process it came from, who verified it, and what the denominator was.
The demand signal
The UK government's own appraisal guidance is the most useful public document a finance director can read on this, and it is free. The Green Book requires appraisers to account for optimism bias, which it defines as the proven tendency for appraisals to be over-optimistic about key assumptions, by increasing estimates of cost and duration and decreasing estimates of benefit. It sets a business case out as five interconnected perspectives, strategic, economic, commercial, financial and management, and it is explicit that these are not five separate documents. And it puts benefits realisation, the plan for monitoring costs and actually realising the stated benefits, inside the management case rather than leaving it as something that happens afterwards. Almost no AI business case we have been shown does any of those three things.
Why it stalls
6 failure modes we keep meeting
The case is built on hours saved, and hours saved are not money
An hours figure is a statement about effort, not about cash, and it becomes money only when a headcount changes, a contract is renegotiated, or a capacity constraint that was costing you revenue is released. From our own history, the AI Voice Insights proof of concept our founder led at HSBC was projected to remove more than 1.5 million hours of manual administration a year. That figure is a projection, it was never realised, and the work was a proof of concept rather than a rollout. It is exactly the shape of number a finance function should refuse to accept on its own, including from us. The follow-up question is the whole job: whose budget line moves, by how much, and in which quarter.
Optimism bias is in the numbers and not in the arithmetic
Every estimate in an AI case is made by people who want the programme to happen, which is normal and is precisely why public-sector appraisal guidance requires an explicit adjustment for it. The Green Book defines optimism bias as the proven tendency for appraisals to be over-optimistic about key assumptions, and requires appraisers to increase their estimates of cost and duration and reduce their estimates of benefit. A case with no such adjustment is not neutral, it is optimistic by default, and the reviewers who eventually find that out will discount everything else in the document too.
The run cost was never estimated, so the payback is measured against half the cost
The build is the visible number and the run is the one that decides whether the thing survives its second year: per-run inference and retrieval, human review of everything the system routes to a person, monitoring and evaluation, and a full re-evaluation each time a model version changes underneath you. Cases that price only the build produce a payback period that is wrong in the direction the sponsor wanted, and the correction arrives during the year when the pilot is supposed to be scaling.
It was priced as a build and it lives as an operation
The sponsor who funded an experiment is rarely the person who will carry a live system, its incidents and its audit trail, and in most organisations nobody has been asked to, so a successful pilot with no named owner for its run budget becomes an orphan. That is a finance problem before it is a technology problem, because the fix is a line in next year's plan with a name against it, agreed before the build rather than negotiated after the demo.
Duration is the largest risk in the case and it is not priced
Analysis of 1,355 public-sector IT projects, averaging $130m and 35 months, found that every additional year of duration added about 4.2 percentage points to expected cost overrun, and that 18% of projects were outliers with cost overruns above 25%. A two-year agentic programme carries that risk whoever delivers it, and that includes us. The implication for the case is not that long programmes are forbidden, it is that duration belongs in the risk line with a number attached rather than in the plan as a neutral fact.
Nobody agreed who books the benefit
A benefit with no owner does not appear in anybody's budget, so it is never checked, and the programme's actual result is decided by whoever tells the story most confidently at the end. Name the person whose numbers move, get them to agree the measurement and the date before the build starts, and accept that a benefit nobody is willing to own is usually a benefit nobody believes in.
How we approach it
9 moves, in order
- 01
Start with one process and its cost in currency, before any technology
Volume, elapsed time, the people involved, the error and rework rate, and what the whole thing costs to run today. That is a fortnight of work and it is the foundation of everything else, because a benefit is a difference between two numbers and most cases only ever produce the second one. It also frequently changes which process you would attack, which is worth knowing before you fund the wrong one.
- 02
Split the benefit into three kinds and treat only one of them as money
Cost avoided is money when a headcount, a contract or a licence actually changes, and not before. Capacity released is money only if the released capacity is redeployed to something with a value attached, which is a decision somebody has to take and record. Revenue or loss changed, better conversion, fewer claims leaking, lower fraud losses, is the strongest kind and the hardest to attribute. Label every line as one of the three and the case becomes readable by a sceptic, which is the only kind of reader worth writing for.
- 03
Attach a confidence to every line, and hand the model to finance
Each figure carries the assumption it rests on and how confident we are in it, and the whole thing is handed over in a form your finance function can rework with their own numbers rather than as a PDF. This is what the investment case in our audit deliverable is: the value in currency, the assumptions exposed, the confidence stated, the arithmetic open. A case nobody can rework is a case nobody can check.
- 04
Cost the run before the build is approved
An estimated run cost per candidate workflow, produced during the audit rather than discovered in year two, covering inference and retrieval per run, the human review the routing design actually implies, monitoring, and re-evaluation on model change. Where that number is uncomfortable it usually redirects the design rather than killing the case, which is exactly what it is for.
- 05
Adjust for optimism bias on the record, and say by how much
Increase the cost and duration estimates, reduce the benefit estimates, and state the adjustment as a visible line rather than quietly baking it in. Two reasons. It is what serious appraisal guidance requires, so your reviewers recognise it. And a case that has already discounted itself is far harder to argue with than one that has not, which is a negotiating advantage rather than a concession.
- 06
Keep the unit of work small and the clock short, because duration is priced risk
Sequence the programme as short fixed-price stages with real decision points between them, rather than as a two-year commitment with milestones. Ours are built that way for this reason: an audit fixed over four weeks, a proof of concept fixed over two to four, and productionisation scoped at four to six weeks with a dedicated team. Whoever you buy from, a case built on stages you can stop is worth more than a case built on a plan you have to believe.
- 07
Price the first year against a published rate card rather than a range
One worked sum, on our own published prices, because a finance director reading this deserves at least one. An AI Readiness Audit is a fixed £44,000 over four weeks. An Agentic Proof of Concept is a fixed £20,000 to £55,000 over two to four weeks. Buying both, a diagnosis and one working thing, is therefore £50,000 to £145,000 before any production commitment is made. If the answer is then a build team of three at £70,000 to £85,000 a month, that is £840,000 to £1.02m across a full year, and productionising a single proof of concept is scoped at four to six weeks rather than a year of that. Those are published numbers and the arithmetic is checkable. What none of them tells you is the benefit. The audit goes first for exactly that reason, to put a defensible figure on the other side of the sum before the larger commitment is made.
- 08
Name the benefit owner and the measurement date before the build starts
One named person whose numbers move, one agreed measurement, one date in the diary, all fixed before anyone writes code. This is the same discipline as putting a named owner on the release gate, applied to the money instead of the risk, and it is what separates a programme that can prove its result from one that has to assert it.
- 09
Write the do-not-do list and count it as return
The things not worth doing here, and why, are usually the most valuable page in an audit readout, and they belong in the business case rather than in an appendix. A programme that avoids a seven-figure commitment to the wrong workflow has produced a return that is real, immediate and considerably more certain than any of the benefits above. Finance functions understand avoided spend better than anyone else in the building, so write it in their language.
An Agentic Proof of Concept is a fixed £20,000 to £55,000 over two to four weeks.
Build the cost side before the benefit side, because it is the half you can actually know. The build price is quotable from published rates, and ours are on the pricing page. The run cost is estimable per workflow before anything is committed, and it is the line most cases are missing entirely. Only then argue about the benefit, and argue about it with the process owner rather than with a supplier, because the supplier does not know whether your headcount will actually change. If you want the whole thing done properly that is what the audit is: four weeks, a fixed fee of £44,000, a costed plan in three-month increments capped at twelve months, a do-not-do list, and an investment case your board can act on. If you already have a case and want it stress-tested rather than written, say so on the call, because that is a shorter and considerably cheaper conversation.
Prefer to talk it through? Ask us on a discovery call →
Sources
Every source below was opened and read before it was attached. Where nothing survived that check, the claim on this page was softened rather than given a plausible-looking link.
HM Treasury, The Green Book (2026): appraisal and evaluation in central governmentSource of the optimism bias definition and the instruction to increase cost and duration estimates and decrease benefit estimates, of the five case model, and of the requirement that the management case sets out plans for monitoring costs and realising benefits. Written for public-sector appraisal, and the discipline transfers.Budzier and Flyvbjerg, Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector (2013)1,355 public-sector IT projects, average budget $130m over 35 months. Each additional year of duration added about 4.2 percentage points to average cost risk, and 18% were outliers with cost overruns above 25%.The closest guides to this one
Nearest first, then the rest. Each one carries the same label: written from delivery, or the method we would bring. All 13 are on the hub.
Document intelligence to business intelligence
Getting information out of PDFs and into something the business can decide with.
Our approachAgent evaluation and assurance
Why AI pilots never reach production, and what it takes to get one through the gate.
Our approachAI governance and regulatory evidence
Building the evidence as a by-product of the work, rather than assembling it under deadline.
Our approachAgent identity and access
Agents are not users, and giving them a service account is how this goes wrong.
DeliveredAI-native SDLC and product delivery lifecycle
Changing how software gets specified, built and shipped once AI is in the room.
DeliveredTarget operating model for an AI-native organisation
What changes in structure, roles and decision rights once agents do a share of the work.
Or bring the problem to a call instead of reading three more of these.
Talk it throughTalk to us about the business case for an agentic programme: ROI, payback and what to measure.
A 30-minute call with James Rooney. We'll tell you honestly which parts of this we have done before and which we would be doing for the first time, and you'll leave with a rough scope either way.
most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes
Calendar not loading? Open it on cal.com or email hello@tenhaw.com.
Questions this guide answers
What is the ROI of an agentic AI programme?
Nobody can tell you without looking at your process, and a supplier quoting a figure beforehand is quoting somebody else's business. Cost avoided becomes money only when a headcount, contract or licence changes. Capacity released becomes money only once someone decides where it goes. Revenue or loss changed is strongest and hardest to attribute. Against those sit a build cost, quotable from published rates, and a run cost: inference, retrieval, human review, monitoring and re-evaluation on every model change. Tenhaw publishes no return figure from an agentic system, having taken none into production; its 60% lead-time reduction at Globelynx was delivery-transformation work. A case naming all three benefit types, pricing both and adjusting for optimism bias is defensible; one built on hours saved is not.
How do you build a business case for AI agents?
Six steps and none of them starts with technology. Measure the current process in currency: volume, elapsed time, people, error and rework, total cost to run today. Estimate the future state and split the difference into cost avoided, capacity released, and revenue or loss changed, labelling each. Attach a confidence and the underlying assumption to every line. Price the build from a published rate card and the run per workflow, before approval rather than after. Adjust explicitly for optimism bias by increasing costs and duration and reducing benefits, and show the adjustment. Then name the benefit's owner and the date it gets measured. Tenhaw publishes the method behind these steps in full, free to adopt, so a team can work through them without hiring anyone.
What is the payback period on an agentic proof of concept?
A proof of concept is bought to remove uncertainty rather than to pay back; treating it as an investment with a return is how organisations end up defending a two-week build as though it were a product. It buys a measured answer on whether the workflow can be done, an accuracy read on real data, and an estimated run cost, which together decide whether the larger production commitment is worth making. If the answer is no, it has paid for itself several times by preventing that commitment. Tenhaw's proof of concept for a London specialty insurance business took PDFs to business intelligence on Azure in two weeks, ground the business had circled for roughly a year, and the client owned the work product either way.
Which running costs belong in an agentic business case?
Any invented number would describe Tenhaw's workloads rather than yours, so it estimates a run cost per candidate workflow before anything is committed to. The lines are consistent and you can build the figure yourself. Per-run model cost, driven by the number of turns the agent takes and the tokens in its context each time. Retrieval and index maintenance, including re-embedding when the corpus or embedding model changes. Human review, set by your routing design rather than the technology, and usually the largest line in a regulated process. Monitoring and observability. And re-evaluation on every model version change, a recurring cost most plans treat as a one-off, and the line that decides between designs.
Should we count hours saved as savings?
Not as savings, no. Count them as a measurement of the change, then do the second piece of work and establish what happens to those hours. If a fixed-term contract ends or a vacancy goes unfilled, that is cash and it belongs in the case with the date it lands. If the time returns to people who then do something else, it is capacity released, and it is worth money only once someone decides what that something else is and attaches a value. If nothing changes, the hours are real and the money is not. Writing that distinction into the case protects you twice, since it survives scrutiny and stops the programme being judged in a year against a saving that was never going to appear in a ledger.
How do you stop the business case being over-optimistic?
By adjusting for it deliberately and visibly, which is what public-sector appraisal guidance has required for years. The Green Book defines optimism bias as the proven tendency for appraisals to be over-optimistic about key assumptions and instructs appraisers to increase their estimates of cost and duration and decrease their estimates of benefit. Three practical habits follow. Show the adjustment as its own line, so reviewers can see it rather than hunt for it. Have someone outside the programme write the pessimistic case, not the sponsor. And treat duration as a priced risk rather than a scheduling detail, because the published evidence on IT programmes is that cost risk rises with every additional year of runtime.
What should finance measure after go-live?
Four things, monthly, and only the first is about the technology. Run cost per unit of work, against the estimate made before approval, because that is where an unpleasant surprise shows up first. Volume actually processed by the system rather than volume it is capable of processing, since adoption is usually the constraint. The exception rate and therefore the human review hours the system is really consuming, which is the line that quietly eats the benefit. And the benefit itself, measured on the agreed date by the named owner, against the figure in the approved case rather than against a revised one. If the fourth measurement has no owner and no date, the first three will be reported and the programme's result will never actually be established.
Who should own benefits realisation?
The person whose budget or performance numbers move, which is almost never the person who sponsored the technology. Get them to agree the measurement and the date before the build starts, because agreeing it afterwards is a negotiation and agreeing it beforehand is a specification. The same person should also own the run budget, since a benefit owner with no cost exposure has an incentive to be generous. Where no such person exists, surface that immediately. A process nobody owns end to end is a programme risk long before it is a measurement problem, and it usually means the first piece of work is establishing the ownership rather than building anything.
Should we fund an AI programme in stages or as one multi-year commitment?
In stages, with a real decision point between them, because duration is the largest unpriced risk in most AI business cases. Short fixed-price stages turn the commitment into a sequence of small ones. Tenhaw's founder co-designed a target operating model at HSBC Global Payment Solutions, 500 teams against a $450M portfolio, piloted first, with a global rollout due in 2026 and not yet begun. Analysis of 1,355 public-sector IT projects, averaging $130m over 35 months, found that every additional year of duration added about 4.2 percentage points to expected cost overrun, and that 18% overran by more than 25%. That risk applies whoever delivers, Tenhaw included. A case built on stages you can stop is worth more than a plan you have to believe.
What does the cost side of an agentic business case look like?
Build the cost side first, because it is the half you can actually know. It is made of three lines you can know before approval: a fixed-price diagnosis, a fixed-price proof of concept if the diagnosis warrants one, and a build team charged by the month if the proof of concept earns it. Tenhaw fixes each stage separately, agrees the exit date at kickoff and works on 30 days' notice either way, so the commitment can be stopped between stages rather than only forecast. Estimate the run cost per workflow separately, then argue about the benefit.
What should we ask a supplier who quotes an AI ROI figure?
Three questions, in this order: whose process the figure came from, who verified it, and what the denominator was. A return quoted before anyone has looked at your workflow is describing somebody else's business, because it is the shape of your process, its volume, its error rate and the review it needs, that decides whether any of it transfers. Then ask whether the figure was realised or only projected, and if it was realised, what baseline it was measured against and who agreed that baseline before the work started. Arithmetic you can rework beats a headline you cannot, whoever is presenting it.
Can you stress-test a business case we have already written?
Yes. A stress test asks what a sceptical reviewer will ask. Is every benefit labelled as cost avoided, capacity released, or revenue or loss changed? Does anything rest on hours saved without a headcount, a contract or a capacity constraint actually moving? Is the run cost priced beside the build, or is the payback measured against half the cost? James Rooney runs that review personally, as he does every Tenhaw engagement, and reviewing an existing case is shorter work than writing one from scratch. What usually changes is not the conclusion but the size of the number and the confidence attached to it, which is what makes a case survive its second reading.
Does the Green Book five case model apply to an AI business case?
The discipline transfers, and the Green Book is the most useful free document a finance director can read on this even though it was written for public-sector appraisal. It sets a business case out as five interconnected perspectives, strategic, economic, commercial, financial and management, and is explicit that these are not five separate documents. It also puts benefits realisation, the plan for monitoring costs and actually realising the stated benefits, inside the management case rather than treating it as something that happens afterwards. For an agentic programme that is exactly where the run cost owner and the measurement date belong, and it is the section Tenhaw most often finds empty.
How long does it take to put an AI business case together?
Measuring one process properly, in currency and as it actually runs today, is about a fortnight of work. That fortnight is the foundation of everything else, because a benefit is the difference between two numbers and most cases only ever produce the second one. Tenhaw's own version of that job runs four weeks and covers several candidate processes with an estimated run cost each, a costed plan in three-month increments capped at twelve months and a do-not-do list. What stretches the timeline is rarely the arithmetic. It is getting the person whose numbers move to agree the measurement.
Can our finance team rework the numbers in an AI business case?
They should be able to, because a case nobody can rework is a case nobody can check. Every figure should carry the assumption it rests on and a stated confidence, and the whole thing should arrive in a form your finance function can put its own numbers into rather than as a PDF. Tenhaw hands that model over as work product the client owns outright, so your team can take it apart without asking anyone's permission. That matters because the assumptions most worth changing are yours: the fully loaded cost of an hour, the discount rate you apply, what the rework rate really is once somebody measures it. Finance revising the figures downwards is a good outcome rather than an argument.
Does deciding not to build something count as a return?
Yes, and it is usually the most certain return in the document. A programme that avoids a seven-figure commitment to the wrong workflow has produced something real and immediate, which no projected benefit sitting two years out can match. So write the do-not-do list, the things not worth doing here and why, into the business case itself rather than into an appendix, and put the avoided commitment in currency beside the benefits you are claiming. Finance functions understand avoided spend better than anyone else in the building, so it is worth writing in their language. In audit readouts that list is often the most valuable page.
How do we choose which process to build the first business case around?
Measure the candidates before you choose, rather than choosing and then justifying. For each one: volume, elapsed time, the people involved, the error and rework rate, and what it costs to run today. That exercise frequently changes which process you would attack, which is worth knowing before you fund the wrong one, and it produces the first of the two numbers any benefit is made from. Favour a process where you can name the person whose budget or performance numbers move if it improves, because a benefit nobody is willing to own is usually a benefit nobody believes in. A process you cannot measure today is not a first candidate, it is a measurement job.
Should the firm building the system also write the business case?
They can build the cost side, because that half is quotable from a published rate card and estimable per workflow before anything is committed. The benefit side is different. A supplier does not know whether your headcount will actually change, which contracts can be renegotiated, or where released capacity would be redeployed, so those figures belong with the process owner rather than whoever is selling the build. Tenhaw reports any month that delivers no measurable value as a failed month, which is accountability a supplier can carry; a forecast of your budget is not. The practical split is the build price and run cost estimate from the firm that would deliver it, and the benefit numbers from the person inside your organisation whose budget moves if the workflow improves.