Measuring and managing

How to deliver a project on time

Predictable delivery is not about pushing harder. It is about seeing the slip early enough to make a real choice about it.
steps
9
named failure modes
5
definition-of-done criteria
5

How to deliver a project on time, in one paragraph

Managing delivery to be on time is a detection problem, not an effort problem. You cannot make a late epic early by pushing, but you can see the slip in week four instead of week eleven, while cutting scope, moving an epic or reallocating people are still real options. The method is four habits: commit to dates derived from your own throughput, watch approval gates as the early warning, re-forecast every fortnight, and turn every slip into a priced decision with a named owner and an edited roadmap rather than a status update.

That is the procedure. The call is where it meets your delivery structure.

Talk it through

What these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.

When to use this

Use it from the day a quarter's roadmap is committed until the day it closes. Reach for it in particular when someone asks whether a date will hold and the honest answer is that nobody in the room knows.

Time to run it once
One quarter, from the day the roadmap commits to the day it closes
What you need
  • The delivery tracker
  • An export of the last eight to twelve weeks of throughput
  • A spreadsheet for the forecast simulation

If you do not have those in place, the call is a good place to work out what comes first.

Talk it through
On this page
9 steps

Step by step

Each step is deep-linkable, so you can send a colleague the one that is in dispute.

  1. 1

    Define on time before the quarter starts

    Put every epic in the quarter on one table, including the standing Tech Debt and Bug Budget epics.

    Each row carries four things: the date it must clear Ready for Dev, the date it must be shipped and in value monitoring, its planned share of its outcome's currency target, and the name of the person who accepted that date. That table is the commitment, and on time is a property of the whole table rather than of whichever epic is being asked about loudest. Anything not on the table is not committed, and you say so the first day someone assumes otherwise.

  2. 2

    Build a throughput baseline from your own history

    Export the last two or three quarters of finished work from your tracker.

    Count completed items per team per week: stories for an AI-augmented team, outcome tickets for an AI-native one. Keep the weekly samples raw, bad weeks, holidays and incident weeks included, because those recur. Aim for at least twelve weekly samples. Record a start date and a done date for every item so you get cycle time as well. Then compute the number few teams have: your calibration ratio, validated currency value over planned currency value from outcome validation. Plan four million, validate two point eight, and every forecast you publish carries seventy per cent.

  3. 3

    Forecast in ranges, plan p50, commit p85

    Never hand the business a single date.

    Take the weekly throughput samples, draw from them at random ten thousand times against the remaining item count, and read the distribution. Plan the team's work against the p50. Give anyone outside the team the p85. Inflate the item count first by your historical scope growth: if last quarter's epics finished with twenty per cent more items than they opened with, forecast 1.2 times today's count and say that is what you did. Publish the count next to the dates, because it is the variable that moves them most. When the count changes, re-run the simulation rather than debating the old date.

  4. 4

    Sequence the portfolio, not the loudest epic

    Capacity is shared even when boards are not.

    Lay every epic in the quarter on one timeline, mark the named people or teams each one needs, and find the weeks where two epics want the same person. Stagger start dates until every epic clears its p85 inside the quarter, and cap epics in flight per team at two, because a third converts into cycle time rather than output. Then redo the arithmetic: if the moved epics no longer cover their outcomes' currency targets, that gap gets a named owner before the quarter opens, not after it ends. One safe epic and three sitting at p50 is not a schedule.

  5. 5

    Track gate dates as your earliest warning

    Ship dates slip because gates slip first, weeks earlier and in plain view.

    For an AI-augmented team the gate is product approval, engineering approval and at least one product-approved story attached before the epic leaves Ready for Dev. For an AI-native team it is the currency share, the key user journeys and the test requirements written down, with both approvals. Show days-to-gate on every epic and review it weekly. Treat a gate date that moves as a ship date that has already moved at least as far, and raise it the same day. Miss a gate early, react early.

  6. 6

    Run a five-minute daily lookahead

    One person, five minutes, before stand-up rather than during it.

    Check five things: any item untouched for longer than twice its usual cycle time, any team over its work-in-progress limit, any gate due inside ten working days that is not ready, the Bug Budget epic's burn rate against the same week last quarter, and anything blocked on a team that does not know it is blocking. Write only the exceptions, in a channel the team already reads, one line each with a name attached. The day it becomes a meeting with a round-the-room, it has stopped doing its job.

  7. 7

    Re-forecast fortnightly and raise slips the same day

    At refinement, re-point whatever has shifted in understanding, then re-run the simulation against updated throughput and the current item count. Compare the new p85 with the committed date on your table. If the p85 has crossed it, that is a slip, and it goes to the accountable person that day in one line: this epic was committed for 12 September, its p85 is now 3 October, and the value at risk is six hundred thousand pounds. Send it before you know what to do about it. Nobody thanks you for a slip announced in week eleven that the numbers showed in week four.

  8. 8

    Turn every slip into a priced decision

    A slip is a choice, so arrive with the options costed.

    There are four. Cut scope inside the epic, naming which stories or user journeys come out and what the planned value drops to. Move the epic to next quarter and move its currency share with it, so the outcome gap shows in the plan. Pull people off a named lower-value epic and state what that epic loses. Or accept the later date and say what it costs in months of value monitoring before the year ends. One named person picks, the decision and its date go in the RAID log, and the roadmap is edited to match the choice.

  9. 9

    Close the quarter honestly and recalibrate

    At roadmap close, record for every epic whether it shipped, moved or was dropped, and what value travelled with it.

    Compute two numbers: committed epics landed over committed epics, and validated value over planned value once outcome validation has caught up. Both feed the next quarter's baseline, and the second is your calibration ratio. Do this before you open the next roadmap, because a roadmap opened on last quarter's assumptions inherits last quarter's overrun. Teams that skip it forecast from optimism, and their p85 becomes decoration nobody outside the team believes twice.

Bring a real piece of work to the call and we will walk it through these.

Talk it through
Worked example

A slip found in week four instead of week eleven

Northbridge Retail is an invented mid-sized online retailer, and the numbers below are invented with it. It opens Q3 with one outcome, reduce checkout abandonment, target 2.4 million pounds, carried by four epics at 700k, 600k, 600k and 500k plus the standing Tech Debt and Bug Budget epics. Its calibration ratio last year was seventy per cent, so the plan expects about 1.68 million against a 2.4 million target. That gap sits on the table before the quarter opens rather than surfacing in October. Two teams have averaged nine completed stories a week over three quarters, with weekly samples ranging from four to fourteen. The 600k guest checkout epic holds 46 items, inflated to 55 for scope growth: p50 lands on 5 September, p85 on 12 September, and 12 September becomes the commitment. In week four the daily lookahead flags days-to-gate on that epic going negative, because engineering approval is stuck behind payment sandbox access. The gate moves two weeks. At the next refinement the item count is 58 and the p85 lands on 3 October, past the commitment and past quarter end. The delivery lead sends the sponsor one line that afternoon naming 600k as the value at risk, with three costed options. The sponsor cuts the saved-card journey, dropping the epic's planned value to 380k, and moves one engineer from the 500k epic, which slips its own p85 by nine days. Both changes go in the RAID log and onto the roadmap that week. Found in week eleven, the only surviving option would have been accepting 3 October.

Yours will look different. Thirty minutes is enough to see how.

Talk it through
Failure modes

Where this goes wrong

  1. 01

    Treating percentage complete as progress.

    It is self-reported, it converges on ninety per cent and stays there, and it is why teams discover slips in the final fortnight. A gate date that has moved tells you more in one second than a week of status percentages.

  2. 02

    Forecasting from capacity rather than throughput.

    Counting available developer days and dividing assumes nobody is interrupted, blocked, ill or on holiday. Historical throughput already has all of that baked in, which is why it is the only input worth trusting.

  3. 03

    Forecasting against a fixed item count.

    Backlogs grow as work is understood, so a forecast against today's count is optimistic by construction. Measure how much last quarter's epics grew between opening and closing, and carry that multiplier openly rather than absorbing the growth as a surprise in week nine.

  4. 04

    Protecting the loudest epic and quietly starving the rest.

    The escalated epic gets the people, three others drift unwatched, and the quarter still misses, because on time was always a property of the whole roadmap rather than of the epic with the most senior sponsor.

  5. 05

    Re-baselining the date instead of recording the slip.

    Quietly moving a committed date to match the current forecast makes every quarter look successful and destroys the calibration data the next forecast needs. Move the date if that is the decision, but record that it moved, when, and why.

Definition of done

Done means

  • Every epic in the quarter, Tech Debt and Bug Budget included, has a committed gate date, a committed ship date, a planned currency share and a named person who accepted it, in one place the business can read without asking.
  • Every date given outside the team is a p85 derived from your own throughput, published with the p50 and the item count the simulation assumed.
  • No slip reaches the accountable person later than the fortnight in which the forecast first showed it.
  • Every raised slip has a recorded decision, a named owner, a stated currency consequence and a roadmap edited to match.
  • The quarter closed with a landed-epic hit rate and a recomputed calibration ratio, and both are inputs to the next quarter's forecast.

If you recognise one of those already happening, that is a good call to have.

Talk it through
If your team is AI-native

For an AI-native team the arithmetic is the same but the samples are fewer and fatter: throughput is counted in outcome tickets, so a team may finish two or three a week rather than nine stories, and the gap between p50 and p85 will be wider for the same confidence. Collect more weeks before trusting it. The bigger risk is the definition of done. An outcome ticket counts as done only when its test requirements pass and its key user journeys are demonstrated working, never when the model reports it has finished. Count the model's word as done and throughput inflates, the forecast tightens, and the slip surfaces in value monitoring instead of week four.

The two delivery modes, side by side →

Which mode your team is actually in is the first thing we establish on a call.

Talk it through
The subject behind the procedure

Where this sits in a programme

The procedure is the same whatever you are building. These cover what it runs into when the thing being built is agentic.

If you want this run inside a programme rather than read, that is the conversation.

Talk it through
book a call

Want help installing this?

These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.

most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes

// pick a slot · cal.com/tenhaw/professional-servicesLIVE CALENDAR

Calendar not loading? Open it on cal.com or email hello@tenhaw.com.

Questions

We have no clean history to forecast from. Where do we start?

Start counting this week and forecast anyway. Four weekly throughput samples give a crude range that beats an invented date, and you widen the gap between p50 and p85 to reflect how thin the data is. The numbers have to come from this team's own flow, so do not borrow another team's velocity or an industry benchmark. Until the samples build up, lean on gate dates as the primary signal, because they are observable from day one and need no history to mean something.

The business will not accept a range. They want one date.

Give them the p85 as their one date. That is what the range is for. You plan the team against the p50, you commit externally to the p85, and you keep the p50 inside the team because outside it the earlier number is heard as the date. Say the p85 is a date you expect to beat five times in six, based on the last three quarters of this team's throughput and an item count you publish alongside it. That answer survives week nine, which a confident single date does not. Tenhaw gives its own clients p85 dates for exactly this reason, and publishes the method that produces them.

The date is fixed externally, by a regulator or a contract. What changes?

The date stops being the variable and scope becomes the variable, so run the simulation backwards. Ask how many items this team finishes by the fixed date at p85, compare that with the item count in the epic, and the difference is scope you cut now rather than in the final fortnight. Take it out explicitly, restate the epic's planned currency value at the reduced scope, and have that accepted by name. Tenhaw works to hard windows of its own, and a proof of concept for a London specialty insurance business took PDFs through to business intelligence on Azure inside two weeks, ground it had circled for roughly a year. A fixed date with unfixed scope is not a commitment, it is a countdown.

How large a forecast movement is worth escalating?

Any movement that puts the p85 past the committed date on the table, however small, and any gate date that moves at all. Everything else stays inside the team. The accountable person then hears from you rarely, and when they do it always means a decision is needed, which is what keeps escalation cheap and credible. Escalating every wobble in the p50 trains people to ignore you, which is how the week-eleven surprise reaches teams that were technically reporting all along.

Why do delivery slips only show up in the last few weeks?

Because most teams track percentage complete, which is self-reported, converges on ninety per cent and sits there until the final fortnight makes it unignorable. Ship dates slip because approval gates slip first, weeks ahead of the deadline, so the slip itself was there much earlier and in plain view. Tenhaw installs that early warning on the client programmes it runs. Show days-to-gate on every epic, review it weekly, and treat any gate date that moves as a ship date that has already moved at least as far. Add a fortnightly re-forecast from your own throughput and no slip needs to reach the accountable person later than the fortnight in which the numbers first showed it.

Can we get a late project back on time by pushing the team harder?

No. You cannot make a late epic early by pushing, which is why delivering on time is a detection problem rather than an effort problem. The teams that land their quarters are the ones that see the slip in week four instead of week eleven, while cutting scope, moving an epic or reallocating people are still real choices. That takes four habits: commit to dates derived from your own throughput, watch approval gates as the early warning, re-forecast every fortnight, and turn every slip into a priced decision with a named owner. Tenhaw runs those four habits with teams of two or three senior people, James Rooney on every engagement, rather than a surge at the end.

What are our options when an epic is going to miss its date?

Four, and the job is to arrive with all of them costed rather than with a status update. Tenhaw's delivery leads cost the same four for clients before a slip reaches a sponsor. Cut scope inside the epic, naming which stories or user journeys come out and what the planned value drops to. Move the epic to next quarter and move its currency share with it, so the outcome gap shows in the plan. Pull people off a named lower-value epic and state what that epic loses. Or accept the later date and say what it costs in months of value monitoring. One named person picks, the decision and its date go in the RAID log, and the roadmap is edited to match the choice.

Should we forecast delivery from team capacity or from throughput?

From throughput, every time. Counting available developer days and dividing assumes nobody is interrupted, blocked, ill or on holiday, and historical throughput already has all of that baked in. Export the last two or three quarters of finished work from your tracker and count completed items per team per week, stories for an AI-augmented team and outcome tickets for an AI-native one. Keep the weekly samples raw, bad weeks, holidays and incident weeks included, because those recur. Aim for at least twelve samples, record a start and a done date for every item so you get cycle time as well, then simulate against the remaining item count and read the distribution rather than averaging it.

Is it OK to move a committed date when the forecast slips?

Moving the date can be the right decision; moving it quietly never is. Re-baselining the commitment to match the current forecast makes every quarter look successful and destroys the calibration data the next forecast needs, so if the date moves, record that it moved, when, and why. The honest route is to raise the slip the same day the fortnightly re-forecast shows the p85 crossing the committed date, put the choice to the accountable person with the value at risk attached, and edit the roadmap to match whatever they decide. At quarter close, count how many committed epics actually landed, so the next plan inherits your record rather than your optimism.

Can we count work as done when the AI says it is finished?

No. An outcome ticket counts as done only when its test requirements pass and its key user journeys are demonstrated working, never when the model reports it has finished. Count the model's word as done and your throughput inflates, the forecast built on it tightens, and the slip you should have seen in week four surfaces in value monitoring instead. Tenhaw holds that line on its own client builds, against an open-source standard of 72 rules enforced by an agent. Forecasting in outcome tickets also needs patience. An AI-native team may finish two or three a week rather than nine stories, so the samples are fewer and fatter and the gap between p50 and p85 is wider for the same confidence.

Who should own a delivery date, the delivery lead or the sponsor?

Both, at different points, and the split is what makes the date mean anything. Tenhaw staffs the delivery lead's half of that split on client programmes it runs. Every committed date on the quarter's table carries the name of the person who accepted it, and that person owns the decision when the date comes under threat. The delivery lead owns detection: tracking days-to-gate, re-forecasting every fortnight, and putting the slip in front of the accountable person the day the numbers show it, with the options costed and the value at risk attached. What a delivery lead must not do is choose quietly on the sponsor's behalf, because a date moved without a named decision destroys the record the next forecast is built on.

How many epics should one team have in flight at once?

Two. A third converts into cycle time rather than output, so cap epics in flight per team and stagger start dates instead. What makes the cap real is a portfolio view: lay every epic in the quarter on one timeline, mark the named people or teams each one needs, and find the weeks where two epics want the same person, because capacity is shared even when boards are not. Move start dates until every epic clears its p85 inside the quarter, then redo the arithmetic. If the moved epics no longer cover their outcomes' currency targets, that gap gets a named owner before the quarter opens. One safe epic and three sitting at p50 is not a schedule.

What does on time actually mean for a delivery team?

It means the whole commitment table landed, not that one epic landed. Before the quarter starts, put every epic on one table and give each row four things: the date it must clear Ready for Dev, the date it must ship and be in value monitoring, its planned share of its outcome's currency target, and the name of the person who accepted that date. The standing Tech Debt and Bug Budget epics sit on it too, because they consume the same capacity. That table is the commitment, it lives where the business can read it without asking, and on time is a property of all of it rather than of the epic being asked about loudest.

Why does the quarter still miss after we rescue the escalated epic?

Because on time was always a property of the whole roadmap rather than of the epic with the most senior sponsor. The rescue takes people from somewhere, three other epics drift unwatched, and the arithmetic that mattered was never redone. If people move, name the lower-value epic they come off, state what it loses and re-forecast it, because its p85 moves too. Keep days-to-gate visible on every epic in the quarter rather than on the escalated one alone, and let the fortnightly re-forecast cover all of them. Then reallocation is a priced trade with a named owner, and the epics nobody rescued stop being a surprise in week eleven.

We plan more value than we deliver every quarter. What should we do?

Measure the ratio and carry it into the next plan instead of promising the full number again. Tenhaw reports a month that delivers no measurable value as a failed month for the same reason. At quarter close, once outcome validation has caught up, compute validated currency value over planned currency value. Plan four million, validate two point eight, and every forecast you publish afterwards carries seventy per cent. Compute the landed hit rate alongside it, committed epics that shipped over committed epics. Do both before you open the next roadmap, which otherwise inherits last quarter's overrun. That ratio puts the shortfall on the table in week one, where a sponsor can add an epic or accept a lower target, rather than in October.

Someone assumed their project was committed. How do we handle it?

Anything not on the quarter's table is not committed, and you say so the first day the assumption surfaces. That is blunt, and it is far cheaper than a business planning around a date nobody ever accepted. Then treat the request as what it is, a change to the commitment rather than a favour. Putting it on the table means a named person accepts its dates, and it means something else moves, so arrive with the options costed: cut scope inside a named epic, move one to next quarter, or accept a later date and say what that costs in months of value monitoring. One person decides, the decision goes in the RAID log, and the roadmap is edited to match.

What does catching a slip in week four actually look like?

Take an invented retailer and invented numbers. A 600k guest checkout epic holds 46 items, inflated to 55 for historical scope growth, so the p50 lands on 5 September and the p85 on 12 September, and 12 September becomes the commitment. In week four the daily lookahead flags days-to-gate going negative, because engineering approval is stuck behind payment sandbox access, and the gate moves two weeks. At the next refinement the count is 58 and the p85 is 3 October. The sponsor cuts the saved-card journey, dropping planned value to 380k, and moves one engineer off a 500k epic, which slips nine days. Found in week eleven, the only surviving option is 3 October.

Is a rising bug count an early warning that a date will slip?

It is, and it is worth reading weekly rather than at the end. The Bug Budget epic sits on the same commitment table as everything else, so time spent over its budget is time taken from committed work. The daily lookahead compares its burn rate with the same week last quarter, which beats arguing about whether things feel noisier than usual, and a burn running ahead of that comparison is capacity quietly leaving the plan. It will surface soon enough in the fortnightly re-forecast as a p85 drifting past a committed date. Raise it the same day, exactly as you would a gate that has moved, and name the epic it is costing.