Forecasting and team health

How to run a team health check

Monthly, per team, reviewed by both the team and management. The trend is the signal; any single month is noise.

steps
8
named failure modes
5
definition-of-done criteria
6

How to run a team health check, in one paragraph

A team health check is a monthly forty-five minute session where everyone on one team privately scores the same fixed statements about the work, with the tracker data on the table so predictability and quality are not scored on mood. Good looks like this: ten cards whose wording never changes, scores revealed at once, a colour and a direction arrow on each, at most two actions that exist as real tickets, the card published unedited to management the same day, and a decision rule that acts on three consecutive bad months rather than on one.

What these are. The delivery operating model our engagements install alongside client teams: the method underneath the agentic work rather than the agentic work itself, published in full and free to use. It is written for the person running a quarter, not for a buyer, so if you are evaluating us, read the five priced engagements or the case studies instead.

When to use this

Once a month, per team, in the same week every month, alongside the fortnightly retrospective rather than instead of it. Run an extra one as a baseline in your first fortnight with a new team, and again when a team switches between AI-augmented and AI-native delivery.

Time to run it once
About two hours a month, of which the session is forty-five minutes
What you need
  • The ten health-check cards
  • One page of the month's delivery data
  • A way to score privately before the reveal
8 steps

Step by step

Each step is deep-linkable, so you can send a colleague the one that is in dispute.

  1. 1

    Fix ten cards and name an owner

    A health check is ten statements the team scores, not an open conversation about how things are going.

    Use these ten: line of sight, value confidence, predictability, flow, work arriving ready, tech debt, defects, tooling, safety, pace. Word each so a person can agree or disagree with it: "I know which outcome my current work rolls up to and what it is worth", "we landed what our p85 forecast said we would", "I can change this codebase without fear". The delivery lead owns the ritual, books the same slot monthly, and freezes the wording so month three compares to month one. Ten cards, no additions.

  2. 2

    Pull the delivery data before the room opens

    Half these cards have evidence in the tracker, so bring it.

    For the month just gone: throughput, what the p85 forecast said against what landed, the quarter's Bug Budget epic burn as raised against closed, the Tech Debt epic burn, how many epics bounced back out of Ready for Dev for a missing approval or a missing approved story, and the outcome validation result for every live outcome. One page, circulated the day before so nobody meets it for the first time in the room. The scores stay subjective and should. The data stops predictability and quality being scored on last Tuesday.

  3. 3

    Score privately, then reveal at once

    Everyone scores all ten cards alone before any discussion: a form, a spreadsheet column, sticky notes face down.

    Reveal the lot together. If the delivery lead or the engineering manager scores out loud first, every other score drifts towards theirs and you have spent forty-five minutes measuring the most senior person in the room. Testers, designers, contractors and the product manager all score, because they hit different failure modes. Anyone on the team less than two weeks abstains and says so out loud rather than guessing, which tells you something about onboarding on its own.

  4. 4

    Score two things per card: state and direction

    Every card gets a colour and an arrow.

    Green is good, amber is workable with friction, red is hurting us. Record 2, 1 and 0 beside the colours. Take the arrow from the team median against last month's median, then let the team overturn it out loud with a reason. The pair carries more than either half on its own: amber-improving is a team already fixing something and needs no intervention, green-worsening is the card worth an hour while it still looks fine. Same sheet every month, so twelve months fit on one chart at year end.

  5. 5

    Discuss divergence and movement, nothing else

    Ten cards will not fit in forty-five minutes, so do not try.

    Five minutes on last month's actions, five on the reveal, thirty on six cards, five agreeing new ones. Pick the six: the three with the widest spread across the team, then the three that moved a step since last month. Spread first, because when half the team scores green and half red on the same statement, the gap is the finding, and it usually means two groups are living in different parts of the system. Ask what someone specifically saw that made them score that. Five minutes a card, facilitator cutting it. Everything else is logged, not debated.

  6. 6

    Leave with two actions, each a real ticket

    Two actions that land beat ten that do not.

    Each needs a named owner, a date inside the next month, and a ticket in the same tracker as everything else: tech debt goes on the quarter's Tech Debt epic, anything that changes the product becomes a story or an outcome ticket under a real epic, anything the team cannot fix itself becomes a RAID entry with an owner outside the team. Nothing lives only in the health check document. Open last month's two at the start of the next session and mark each done or not done, out loud, before anyone proposes a new one.

  7. 7

    Publish the card unedited the same day

    The card is read by the team and by management, and it is worth nothing if the second audience gets a softened version.

    Send the scores, the arrows, the two actions and the list of things the team cannot fix alone, in full, on the day. Agree the protective rule in writing before the first session and hold management to it: the card exists to remove obstacles, it never enters an individual's performance review or a league table of teams, and nobody outside the team changes a number. The first time a red quietly becomes an amber before it reaches a director, every score after it is decoration.

  8. 8

    Read the trend, act on three in a row

    One month is noise: somebody had a bad sprint, a release went sideways, half the team was on leave.

    At quarterly roadmap close, put the quarter's three cards side by side with the Tech Debt and Bug Budget epics you are closing out and read them together. Hold this rule: a card red or worsening three months running is not a retro item, it is structural, and it gets an owner at management level plus a named change in next quarter's roadmap. Re-word cards once a year at most, or when a team changes delivery mode, and mark the break on the chart so nobody compares across it.

Worked example

Three months of cards at an invented retailer

Take an invented mid-market retailer, call it Retailer A, and its nine-person payments team running AI-augmented. The live outcome is reducing failed-payment churn, target £1.4m, with five epics planned at £1.2m combined, so there is a visible £200k gap. Month one: line of sight comes back red, four of nine people scoring zero, and all four are developers. Predictability is amber-worsening, because the p85 forecast said six epics would land by this point in the quarter and four did. Quality is amber-worsening: 61 bugs raised against a 40-bug budget. Safety is green-flat. The widest spread is on work arriving ready, product green and engineering red, and the discussion surfaces that seven of eleven epics left Ready for Dev without an approved story attached, so developers were starting on faith. Two actions: the product manager puts the outcome name and its currency share at the top of every epic, and the engineering manager blocks the Ready for Dev transition until both approvals and one approved story exist. Both become tickets with dates inside the month. Month two: line of sight amber-improving, one developer still scoring zero. Quality still amber-worsening, 78 raised against 41 closed. Predictability flat. Month three: line of sight green. Quality red, and worsening for the third month running. That trips the rule, so it stops being a retro item: the engineering director owns it, and next quarter's roadmap carries a named change, the bug budget cut to 25 with a defect triage owner, rather than a fourth conversation about testing more.

Failure modes

Where this goes wrong

  1. 01

    The manager scores in the room, or scores at all on a team of six.

    Anchoring is fast and quiet: two people watch the boss put green on predictability and their own amber starts to feel unfair. Three months later every card is green and the instrument confirms what you already believed.

  2. 02

    Scoring predictability and quality with no numbers in front of you.

    Feelings track the last two days, so a team that missed its p85 forecast by three epics scores green because this week went well. Put throughput, forecast against actual, and bug budget burn on the table first, then let people score against reality.

  3. 03

    Reacting to one red month with a reorg, a new process and a working group, all landing the month the score would have recovered on its own.

    The rule is three consecutive months red or worsening, and holding that line when a director wants visible action is most of the discipline.

  4. 04

    Actions that live only in the health check document.

    If it is not a ticket with an owner and a date in the normal tracker, it competes with committed epic work and loses every time, the same card scores the same next month, and the team learns the ritual changes nothing.

  5. 05

    Averaging the cards into an organisation-wide health percentage for a board slide.

    It destroys the two properties that make the instrument useful, per team and per card, and it rewards generous scoring. Report the cards as cards, and report which ones have been red three months running.

Definition of done

Done means

  • Every person on the team scored all ten cards privately, and the scores were revealed simultaneously.
  • Each card carries a colour and a direction arrow, recorded as 2, 1 or 0 beside the previous two months.
  • Last month's actions were opened and marked done or not done before any new action was agreed.
  • At most two new actions exist, each with a named owner, a date inside the month and a ticket in the normal tracker.
  • The unedited card reached management the same day, including everything the team cannot fix alone.
  • Any card red or worsening for three consecutive months has a named owner outside the team and a named change in the next roadmap.
If your team is AI-native

The instrument is the same and three cards change wording. On an AI-native team stories and chapters do not exist, so work arriving ready is scored on the outcome ticket instead: does it carry the outcome's currency share, the key user journeys and the test requirements, or are you filling the gaps by guessing? Replace the codebase-fear card with a verification card, "I can tell whether what the model produced works, not just that it ran", because that is where AI-native teams fail quietly. Predictability and value confidence stay as they are, since outcomes, epics and the quarterly roadmap do not change with the mode. When a team moves between modes, re-word once, mark the break on the chart, and treat the next month as a fresh baseline rather than a drop.

The two delivery modes, side by side →

Questions

How is this different from a retrospective?

Different scope and different audience. The retrospective runs fortnightly, belongs to the team, and works on the last two weeks: what happened, what to try next. The health check is monthly, scored, and read by management as well as the team, and it works on the system the team sits inside: line of sight, tech debt, defects, safety, pace. Run both. Fold the health check into the retro and the structural problems get traded away for the nearest process tweak, while management never sees the card.

Should the scores be anonymous?

Private until the reveal, not anonymous after it. Anonymous scores kill the only question worth asking, which is what did you specifically see that made you score it that way. Collect scores individually so nobody anchors, reveal them together, then discuss them attributed. If people will not put a red on the board with their name against it, that is your safety card answering itself, and it is a bigger finding than anything else in the session.

What if every card comes back green?

Assume a measurement problem before you assume a healthy team. Check three things: whether a manager scored or spoke first, whether the data agrees (forecast against actual, bug budget burn, epics bouncing out of Ready for Dev), and whether the wording is soft enough that agreeing costs nothing. A card everyone can agree with in a bad month is a badly worded card. Fix it at the annual re-word, not mid-year, and mark the break on the chart.

Who runs it and who attends?

The delivery lead owns and facilitates it, one team at a time: engineers, testers, designers, the product manager, and any contractor who has been there more than two weeks. Line managers of the people in the room do not attend, and no score is taken from anyone outside the team. If the delivery lead also line-manages half the room, borrow a facilitator from another team and have the lead abstain from scoring, because a score from the person who writes your review is not a score.

Want help installing this?

These guides are free and you owe us nothing for using them. If you would rather have operators install the operating model alongside your teams and stay until it sticks, that is what our engagements do.

30 minutesWith James personallyNo obligation

Most organisations start with a fixed-price Agent-Readiness Audit · £30k–£90k · 6–8 weeks