Measuring and managing

How to manage delivery to be on time

Predictable delivery is not about pushing harder. It is about seeing the slip early enough to make a real choice about it.

steps
9
named failure modes
5
definition-of-done criteria
5

How to manage delivery to be on time, in one paragraph

Managing delivery to be on time is a detection problem, not an effort problem. You cannot make a late epic early by pushing, but you can see the slip in week four instead of week eleven, while cutting scope, moving an epic or reallocating people are still real options. The method is four habits: commit to dates derived from your own throughput, watch approval gates as the early warning, re-forecast every fortnight, and turn every slip into a priced decision with a named owner and an edited roadmap rather than a status update.

What these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.

When to use this

Use it from the day a quarter's roadmap is committed until the day it closes. Reach for it in particular when someone asks whether a date will hold and the honest answer is that nobody in the room knows.

Time to run it once
One quarter, from the day the roadmap commits to the day it closes
What you need
  • The delivery tracker
  • An export of the last eight to twelve weeks of throughput
  • A spreadsheet for the forecast simulation
9 steps

Step by step

Each step is deep-linkable, so you can send a colleague the one that is in dispute.

  1. 1

    Define on time before the quarter starts

    Put every epic in the quarter on one table, including the standing Tech Debt and Bug Budget epics.

    Each row carries four things: the date it must clear Ready for Dev, the date it must be shipped and in value monitoring, its planned share of its outcome's currency target, and the name of the person who accepted that date. That table is the commitment, and on time is a property of the whole table rather than of whichever epic is being asked about loudest. Anything not on the table is not committed, and you say so the first day someone assumes otherwise.

  2. 2

    Build a throughput baseline from your own history

    Export the last two or three quarters of finished work from your tracker.

    Count completed items per team per week: stories for an AI-augmented team, outcome tickets for an AI-native one. Keep the weekly samples raw, bad weeks, holidays and incident weeks included, because those recur. Aim for at least twelve weekly samples. Record a start date and a done date for every item so you get cycle time as well. Then compute the number few teams have: your calibration ratio, validated currency value over planned currency value from outcome validation. Plan four million, validate two point eight, and every forecast you publish carries seventy per cent.

  3. 3

    Forecast in ranges, plan p50, commit p85

    Never hand the business a single date.

    Take the weekly throughput samples, draw from them at random ten thousand times against the remaining item count, and read the distribution. Plan the team's work against the p50. Give anyone outside the team the p85. Inflate the item count first by your historical scope growth: if last quarter's epics finished with twenty per cent more items than they opened with, forecast 1.2 times today's count and say that is what you did. Publish the count next to the dates, because it is the variable that moves them most. When the count changes, re-run the simulation rather than debating the old date.

  4. 4

    Sequence the portfolio, not the loudest epic

    Capacity is shared even when boards are not.

    Lay every epic in the quarter on one timeline, mark the named people or teams each one needs, and find the weeks where two epics want the same person. Stagger start dates until every epic clears its p85 inside the quarter, and cap epics in flight per team at two, because a third converts into cycle time rather than output. Then redo the arithmetic: if the moved epics no longer cover their outcomes' currency targets, that gap gets a named owner before the quarter opens, not after it ends. One safe epic and three sitting at p50 is not a schedule.

  5. 5

    Track gate dates as your earliest warning

    Ship dates slip because gates slip first, weeks earlier and in plain view.

    For an AI-augmented team the gate is product approval, engineering approval and at least one product-approved story attached before the epic leaves Ready for Dev. For an AI-native team it is the currency share, the key user journeys and the test requirements written down, with both approvals. Show days-to-gate on every epic and review it weekly. Treat a gate date that moves as a ship date that has already moved at least as far, and raise it the same day. Miss a gate early, react early.

  6. 6

    Run a five-minute daily lookahead

    One person, five minutes, before stand-up rather than during it.

    Check five things: any item untouched for longer than twice its usual cycle time, any team over its work-in-progress limit, any gate due inside ten working days that is not ready, the Bug Budget epic's burn rate against the same week last quarter, and anything blocked on a team that does not know it is blocking. Write only the exceptions, in a channel the team already reads, one line each with a name attached. The day it becomes a meeting with a round-the-room, it has stopped doing its job.

  7. 7

    Re-forecast fortnightly and raise slips the same day

    At refinement, re-point whatever has shifted in understanding, then re-run the simulation against updated throughput and the current item count. Compare the new p85 with the committed date on your table. If the p85 has crossed it, that is a slip, and it goes to the accountable person that day in one line: this epic was committed for 12 September, its p85 is now 3 October, and the value at risk is six hundred thousand pounds. Send it before you know what to do about it. Nobody thanks you for a slip announced in week eleven that the numbers showed in week four.

  8. 8

    Turn every slip into a priced decision

    A slip is a choice, so arrive with the options costed.

    There are four. Cut scope inside the epic, naming which stories or user journeys come out and what the planned value drops to. Move the epic to next quarter and move its currency share with it, so the outcome gap shows in the plan. Pull people off a named lower-value epic and state what that epic loses. Or accept the later date and say what it costs in months of value monitoring before the year ends. One named person picks, the decision and its date go in the RAID log, and the roadmap is edited to match the choice.

  9. 9

    Close the quarter honestly and recalibrate

    At roadmap close, record for every epic whether it shipped, moved or was dropped, and what value travelled with it.

    Compute two numbers: committed epics landed over committed epics, and validated value over planned value once outcome validation has caught up. Both feed the next quarter's baseline, and the second is your calibration ratio. Do this before you open the next roadmap, because a roadmap opened on last quarter's assumptions inherits last quarter's overrun. Teams that skip it forecast from optimism, and their p85 becomes decoration nobody outside the team believes twice.

Worked example

A slip found in week four instead of week eleven

Northbridge Retail is an invented mid-sized online retailer, and the numbers below are invented with it. It opens Q3 with one outcome, reduce checkout abandonment, target 2.4 million pounds, carried by four epics at 700k, 600k, 600k and 500k plus the standing Tech Debt and Bug Budget epics. Its calibration ratio last year was seventy per cent, so the plan expects about 1.68 million against a 2.4 million target. That gap sits on the table before the quarter opens rather than surfacing in October. Two teams have averaged nine completed stories a week over three quarters, with weekly samples ranging from four to fourteen. The 600k guest checkout epic holds 46 items, inflated to 55 for scope growth: p50 lands on 5 September, p85 on 12 September, and 12 September becomes the commitment. In week four the daily lookahead flags days-to-gate on that epic going negative, because engineering approval is stuck behind payment sandbox access. The gate moves two weeks. At the next refinement the item count is 58 and the p85 lands on 3 October, past the commitment and past quarter end. The delivery lead sends the sponsor one line that afternoon naming 600k as the value at risk, with three costed options. The sponsor cuts the saved-card journey, dropping the epic's planned value to 380k, and moves one engineer from the 500k epic, which slips its own p85 by nine days. Both changes go in the RAID log and onto the roadmap that week. Found in week eleven, the only surviving option would have been accepting 3 October.

Failure modes

Where this goes wrong

  1. 01

    Treating percentage complete as progress.

    It is self-reported, it converges on ninety per cent and stays there, and it is why teams discover slips in the final fortnight. A gate date that has moved tells you more in one second than a week of status percentages.

  2. 02

    Forecasting from capacity rather than throughput.

    Counting available developer days and dividing assumes nobody is interrupted, blocked, ill or on holiday. Historical throughput already has all of that baked in, which is why it is the only input worth trusting.

  3. 03

    Forecasting against a fixed item count.

    Backlogs grow as work is understood, so a forecast against today's count is optimistic by construction. Measure how much last quarter's epics grew between opening and closing, and carry that multiplier openly rather than absorbing the growth as a surprise in week nine.

  4. 04

    Protecting the loudest epic and quietly starving the rest.

    The escalated epic gets the people, three others drift unwatched, and the quarter still misses, because on time was always a property of the whole roadmap rather than of the epic with the most senior sponsor.

  5. 05

    Re-baselining the date instead of recording the slip.

    Quietly moving a committed date to match the current forecast makes every quarter look successful and destroys the calibration data the next forecast needs. Move the date if that is the decision, but record that it moved, when, and why.

Definition of done

Done means

  • Every epic in the quarter, Tech Debt and Bug Budget included, has a committed gate date, a committed ship date, a planned currency share and a named person who accepted it, in one place the business can read without asking.
  • Every date given outside the team is a p85 derived from your own throughput, published with the p50 and the item count the simulation assumed.
  • No slip reaches the accountable person later than the fortnight in which the forecast first showed it.
  • Every raised slip has a recorded decision, a named owner, a stated currency consequence and a roadmap edited to match.
  • The quarter closed with a landed-epic hit rate and a recomputed calibration ratio, and both are inputs to the next quarter's forecast.
If your team is AI-native

For an AI-native team the arithmetic is the same but the samples are fewer and fatter: throughput is counted in outcome tickets, so a team may finish two or three a week rather than nine stories, and the gap between p50 and p85 will be wider for the same confidence. Collect more weeks before trusting it. The bigger risk is the definition of done. An outcome ticket counts as done only when its test requirements pass and its key user journeys are demonstrated working, never when the model reports it has finished. Count the model's word as done and throughput inflates, the forecast tightens, and the slip surfaces in value monitoring instead of week four.

The two delivery modes, side by side →

Questions

We have no clean history to forecast from. Where do we start?

Start counting this week and forecast anyway. Four weekly throughput samples give a crude range that beats an invented date, and you widen the gap between p50 and p85 to reflect how thin the data is. Do not borrow another team's velocity or an industry benchmark: the point is that the numbers come from this team's own flow. Until the samples build up, lean on gate dates as the primary signal, because they are observable from day one and need no history to mean something.

The business will not accept a range. They want one date.

Give them one date: the p85. That is what the range is for. You plan the team against the p50, you commit externally to the p85, and you keep the p50 inside the team because outside it the earlier number is heard as the date. Say the p85 is a date you expect to beat five times in six, based on the last three quarters of this team's throughput and an item count you publish alongside it. That answer survives week nine, which a confident single date does not.

The date is fixed externally, by a regulator or a contract. What changes?

The date stops being the variable and scope becomes the variable, so run the simulation backwards. Ask how many items this team finishes by the fixed date at p85, compare that with the item count in the epic, and the difference is scope you cut now rather than in the final fortnight. Take it out explicitly, restate the epic's planned currency value at the reduced scope, and have that accepted by name. A fixed date with unfixed scope is not a commitment, it is a countdown.

How large a forecast movement is worth escalating?

Any movement that puts the p85 past the committed date on the table, however small, and any gate date that moves at all. Everything else stays inside the team. That rule keeps escalation cheap and credible: the accountable person hears from you rarely, and when they do it always means a decision is needed. Escalating every wobble in the p50 trains people to ignore you, which is how the week-eleven surprise reaches teams that were technically reporting all along.

Want help installing this?

These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.

30 minutesWith James personallyNo obligation

Most organisations start with a fixed-price Agent-Readiness Audit · £30k–£90k · 6–8 weeks