How we deliver, in the FAQ

Bugs and root cause, answered in full.

What happens when something is wrong: raising a bug somebody can act on, and finding the cause rather than the symptom. Written for the person running a quarter rather than for a buyer, and free to use with us or without us.
questions in this group, each answered in full
36
pages the answers are written on, every one linked
2
questions across the whole FAQ
1424

36 questions on bugs and root cause, answered by Tenhaw, a UK AI consultancy and AI delivery partner based in London. Nothing here is a summary: each answer is the exact text from the page that owns it, and every group links back to that page for the context around it.

On this page
18 questions

How to write a bug report

Answered on How to write a bug report, and rendered here in the same words.

Read the page these answers live on →

Is a missed requirement a bug?

No. If nobody approved the behaviour being asked for, nothing is defective, the product is doing what was agreed. That is new scope, either a story under an existing epic or a new epic carrying its own share of an outcome's value. The test is whether you can link to an acceptance criterion, user journey or test requirement that the live behaviour contradicts. If you cannot produce that link, you are raising a change request wearing a bug's clothes, and the bug budget will lie about quality all quarter.

Do defects found before release count against the bug budget?

No. A story in progress or in review that does not meet its acceptance criteria is not done, so send it back rather than opening a bug. The same applies in an AI-native team when a test requirement in the outcome ticket fails before release. The bug budget forecasts what escapes into production, and polluting it with in-flight rework destroys the one signal it carries. Track pre-release rework, if you want it, as its own measure of how well work is being written and reviewed.

Who sets severity, and can it be changed?

The reporter sets an opening severity against the published rubric, because they have seen the impact. One named person, usually the product manager who owns the parent outcome, is allowed to change it, and the reason goes in the ticket. Set the rubric on user and revenue impact, never on how loudly the request arrived. Severity decides three things: whether the bug interrupts the current sprint, whether it triggers a rollback conversation, and whether it earns a root cause analysis, so an inflated S1 costs real capacity.

What do we do when the bug budget runs out mid-quarter?

Raise it at the next monthly health check and treat it as a forecast that has broken, not a cap you have breached. Tenhaw reports its own months on those terms, and a month that moves no number the client agreed is written up as a failed month. S1s still interrupt the sprint, because a budget does not make revenue exposure acceptable. What changes is the planned work. Something in the quarter gives way, and product decides which epic slips rather than letting the team absorb it silently. Then look at where the defects came from using the introduced-by links, because an overrun concentrated in one epic is a different problem from one spread evenly.

What makes a good bug report?

One a stranger can act on without finding you. It reproduces on someone else's machine, with the environment, build version, account type, data state and exact input written down. It states expected against actual in two sentences, linking the acceptance criterion or approved journey the live behaviour contradicts, which is what stops triage becoming an argument about whether anyone agreed the behaviour. Severity is set against a published rubric, and anything serious carries a weekly cost in currency. The title names the area, what breaks, for whom and under what condition, so "Checkout: card payments over £500 fail for guest users" beats "payments broken". Keep your theory of the cause in the body, labelled as a theory.

How do you decide between rolling back and fixing forward?

Compare two numbers you can write down, the exposure per day in currency and the hours to a fix you would actually trust. Roll back when the daily exposure exceeds what the change earns per day and the trusted fix is more than a day out, or when you cannot yet bound the blast radius. The call is made explicitly by a named person, on the record, not by drift. On Tenhaw engagements that person is James Rooney, who leads every engagement personally and stays the named escalation route throughout. Record the decision in the ticket with both numbers and the reasoning, including a decision not to roll back, because the root cause analysis will ask for it.

What should we do with a bug nobody can reproduce?

Say so in the first line rather than leaving it implied, then list what you tried, how many users hit it each day, and where you stopped. Attach whatever evidence exists: a trace ID, the log lines either side of the failure, a screen recording. Then give it a decision date fourteen days out with three ways to resolve: reproduce it, ship instrumentation that will catch it next time, or close it with the reasoning written down. The failure mode to avoid is the unreproducible bug that ages for a quarter because nobody can prove it has gone, sitting in the backlog radiating uncertainty.

How do you put a cost on a bug?

Multiply the affected users per week by the value lost per affected user, using the same conversion and basket numbers the parent work was priced with, so both sides of the ledger match. A checkout defect hitting 30 guest payments a day at an average £610 basket, where roughly half recover by retrying, costs about £9k a day. A weekly figure in currency does two jobs. It makes the bug rankable against feature work in planning, and it gives the rollback decision a number to weigh against the hours to a trusted fix. Without it, priority gets set by whoever is most annoyed. Tenhaw reports its own engagements as a measured movement, with delivery lead times at Globelynx falling 60% within six months.

Should you fix every bug?

No. Fix by severity and by value. A defect that blocks a key journey with no workaround, or carries revenue, safety or regulatory exposure, interrupts the sprint immediately. A degraded journey with a workaround gets scheduled inside the quarter. Everything else, cosmetic defects and edge cases with negligible reach, competes on value with the rest of the backlog, and some of it should be closed unfixed with a note saying so. Closing a low-value bug deliberately, with the reasoning recorded, is healthier than letting it age in the backlog pretending it will ever be scheduled.

How do you stop the same bug coming back?

Write the regression test before the fix and watch it fail on the affected build. Tenhaw closes defects on that failing-then-passing test, against an open-source standard of 72 rules with RFC 2119 severities. A fix with no test that failed beforehand is a claim, and it is why the same defect returns two releases later. Name the test and the failing build so the next person can rerun it, and close only when the behaviour is demonstrated in the environment where it was reported, not staging. When a model wrote it, correct the requirement first, the user journey or test requirement on the ticket, then have the model implement against it; patching the symptom leaves the requirement wrong, and the next rebuild reintroduces the defect.

How long should it take to raise and triage a bug?

About ninety minutes, covering both the writing and the triage. That pass classifies it, reproduces it three times, states expected against actual with the approved source linked, sets a severity and a weekly cost, links it to the quarter's bug budget and to the release that introduced it, and records the continue, watch or roll back call. Against dropping a message into a channel that sounds heavy, and it is meant to. The time comes back twice. Nobody re-argues at triage whether the behaviour was ever agreed, and the tickets never bounce back to you for the detail you left out.

How many times should you reproduce a bug before raising it?

Three, and record the hit rate rather than just the fact that it happened. Intermittency changes both the fix and the test that has to prove it, so three of three and one of five are different defects, and an unstated hit rate hides the hardest part of the work. Note the timestamp of an occurrence you watched happen, and attach a request or trace ID, the log lines either side of the failure, and a short screen recording. Then have someone other than the reporter follow the steps successfully before the ticket moves, which is the only real test of whether they were written for a stranger.

When is a defect tech debt rather than a bug?

When no user can see it and it is painful only to engineers. Tenhaw made that split pay at Greggs, where throughput data proved that paying down tech debt made delivery faster. A defect users experience goes under the quarter's Bug Budget epic, and one that only slows the team down goes under the Tech Debt epic. Keeping the line clean matters because the bug budget is meant to describe what customers actually met in production, so filing debt against it makes the burn rate say something it does not mean, and filing a customer-visible defect as debt hides it from the quality signal entirely. Decide before you write a word, because everything downstream trusts the classification.

How do you track which release introduced a bug?

Keep an "introduced by" field on every bug, traced from the release that first shipped the broken behaviour, and leave it empty rather than guessing at it. It is the link most teams skip, and it is what turns the bug budget from a bucket into three numbers at quarter close: burn against forecast, defects per epic, and the severity mix quarter on quarter. Without it you have a count of bugs, which measures how busy the team was rather than where quality is leaking. The quarter close report reads that field, so a bug raised without it goes missing exactly when the pattern matters most.

Who has to sign off before a bug can be closed?

Four conditions, rather than one person's opinion. Tenhaw holds the same four inside client repositories, where the code and documentation are the client's from day one. The named regression test passes, having failed first on the affected build. A human has reviewed the change. Product has approved it. And the behaviour has been demonstrated in the environment where the defect was reported, not only in staging, because staging is not where the report came from. Anything set as S1 also has its root cause analysis booked before the bug closes, and that session investigates the system that let the defect through rather than the person who wrote the line. A model reporting that it has fixed something is not evidence that it has.

How do you stop teams relabelling bugs to protect the numbers?

Make the classification a test anyone can apply, and apply it at triage rather than at quarter close. Tenhaw publishes that test in full and free to adopt, so a delivery team can be held to a public document rather than a private convention. A bug has to cite the acceptance criterion, journey or test requirement that the live behaviour contradicts, so moving one to a story is a claim that no such citation exists, which is checkable, rather than a quiet edit that makes a number look better. It matters because relabelling happens under pressure near the end of a quarter, and it leaves the burn rate worthless. The number stops describing quality and starts describing how motivated the team was to protect it.

What do you check a bug against when there are no acceptance criteria?

The user journeys and test requirements written on the outcome ticket. An AI-native team writes no stories and no acceptance criteria, so those are the approved source, and a bug is live behaviour that contradicts one of them. Cite it by link on the expected line exactly as you would an acceptance criterion, because the citation is what stops triage becoming an argument about whether anyone ever agreed the behaviour. Where the ticket is silent on the behaviour being demanded, you are not looking at a defect at all, you are looking at scope, and it belongs in the ticket as a new requirement.

Do bug tickets need refining and sizing like user stories?

Yes. Bugs go through refinement and sizing like everything else, and through the same phases: todo, in progress, in review, done. The reason is capacity. Unsized bug work still consumes the team's week, it just consumes it somewhere the plan cannot see, and a budget tracked in ticket counts rather than in sized work tells you how busy people were rather than how much of the quarter has gone. Sizing also puts a defect and a feature in the same units, which makes the choice between them a planning decision instead of an improvisation. Past half the budget before half the quarter, raise it at the next health check.

If the sources do not answer it, a call will.

Talk it through
18 questions

How to run a root cause analysis

Answered on How to run a root cause analysis, and rendered here in the same words.

Read the page these answers live on →

How is an RCA different from live monitoring?

Live monitoring is the scheduled watch over a change after it ships: customer impact, support volume, the FAQ and the support macro, and the continue, watch or rollback call. It runs on every release, whether or not anything is wrong. An RCA is triggered by what live monitoring finds, or by an incident that arrives with no warning at all. Live monitoring asks whether this release is behaving. An RCA asks why the system allowed it not to, and it ends in funded tickets rather than a call.

Who should run the session?

Someone who did not build the thing that failed and does not manage the people who did. Their job is the method rather than the investigation: holding the timeline until it is agreed, applying the counterfactual test to every candidate cause, and stopping the room whenever an answer names a person instead of a control. Tenhaw does this as an outside facilitator, brought in because it did not build the system under discussion. Keep the room to the people who were there plus that facilitator. Once it becomes a stakeholder audience, people start performing rather than remembering, and you lose the detail you came for.

What if the cause sits with a supplier we do not control?

Then you have found the trigger and you still owe the conditions. You cannot action a third party's deploy schedule, but you can action the timeout you did not set, the fallback you did not build, the contract test you did not write, and the alert that would have told you their response had changed shape. Raise the commercial conversation separately, and log the dependency in the RAID log as an accepted risk with a review date, but do not let a supplier's name become the reason no ticket was raised.

Is this supposed to be blameless?

Blameless means the cause is never a person, not that names vanish from the timeline. Write plainly that an engineer ran the deploy at 09:12, because the timeline is worthless without it. What you never write is that the engineer was the cause. If the honest finding is that one person's memory was the only thing standing between a change and production, then the missing control is a gate, and the action is to build it.

Which incidents actually need a root cause analysis?

Agree the trigger list in advance and apply it without debate: anything that reached users, any rollback called during live monitoring, any severity-one defect, any near miss caught by a person rather than by a control firing, and the third occurrence of the same failure shape. The delivery lead makes the call within one working day and records which trigger fired; everything else becomes a bug ticket with no ceremony. Deciding case by case, in the mood straight after an outage, means the embarrassing incidents get investigated and the boring recurring ones, which cost more in aggregate, never do.

How soon after an incident should you run a root cause analysis?

Freeze the evidence within 24 hours and hold the session within five working days. One named owner copies logs, alert payloads, deploy records and the incident chat into a record that will not rotate away, and everyone involved writes down what they saw before any group meeting. The ninety-minute session then builds one agreed timeline before causes are discussed, and the write-up is published within five working days. Wait two weeks and logs have rotated and everyone has rehearsed a version of events, so what you get is a coherent document that is partly fiction.

What is the difference between a trigger and a root cause?

The trigger is the change that lit the fuse: a deploy, a traffic spike, an expiring certificate, a third party altering a response shape. The root causes almost always sit in the conditions that made the system flammable: no timeout, no fallback, an alert routed to a channel with no rota, a test suite that never covered that journey. A rollback removes the trigger and leaves every condition standing, which is why the same incident returns wearing different clothes. Nearly every action you raise should come out of the conditions column, not the trigger.

How do you decide if something is a real cause or just context?

Run a counterfactual test on every candidate. If this alone had been different and nothing else had changed, would the incident still have happened? If yes, it is context and belongs in a background paragraph. If no, it is a contributing cause and it earns exactly one action. Run the test out loud so the room hears which items fail it. Expect two to four contributing causes from a typical incident; a list of eleven is a wish list of everything anyone dislikes about the codebase, and nobody funds a wish list.

How many whys should a five whys analysis actually take?

Usually three or four. Stop when you reach a control the team can change: a standard, a test, an alert, a gate, a default, a runbook. If an answer names a person, you are not finished; ask why the system let one person's attention be the control. If it names something outside your control, the control you own is whatever should have contained it, such as a timeout, a circuit breaker or a contract test. Chains of seven whys tend to be creative writing rather than analysis.

How do you stop RCA actions being forgotten?

Raise every action as a ticket during the session, on screen, before anyone leaves; an action that is not a ticket is a sentence. Each one gets a named owner, not a team, and a home where it competes for capacity, which in The Tenhaw Way means the quarter's Bug Budget epic for defects, the Tech Debt epic for missing tests and alerting, or a new epic under an outcome for anything larger. If an action cannot get an owner and a quarter in the room, record it as deliberately declined, with the reason. Then put the actions on the agendas of the next two retrospectives and close them out loud.

Is a root cause analysis the same as a post-mortem?

The names get used interchangeably, and the label matters far less than the bar you hold the session to. Four things decide whether it was worth running, whatever you call it: evidence frozen inside the first 24 hours, one timeline in UTC that names the source of every entry and states both time to detect and time to restore, every candidate cause put through a counterfactual so you finish with two to four rather than eleven, and at least one action that shortens detection. Every action leaves the room as a ticket with a named owner. If it ends as a document, it did not happen.

How do we introduce root cause analysis to a team that has never done one?

Set the mechanics up before the next incident rather than during it. Tenhaw publishes this method in full, and a team can adopt every step of it without hiring anyone. Write the trigger list down and get it agreed, and name the delivery lead for that team as the person who applies it within one working day. Decide where the incident record lives and who owns freezing it inside 24 hours, because logs and dashboards roll at seven or thirty days. Agree who facilitates, which has to be someone who did not build the thing that failed. Then book ninety minutes within five working days, open with the timeline and nothing else, and raise every action as a ticket on screen before anyone leaves.

How do you build an incident timeline everyone agrees on?

Collect separate accounts before the room ever meets. Ask everyone involved to write down what they saw and when they saw it, in their own words, ahead of any group discussion, because ten minutes into a meeting the room has converged on the loudest person's version and the original recollections cannot be recovered. Then open the session with the timeline and nothing else. Timestamp to the minute in UTC, name the source of every entry, and mark four moments: when the condition entered the system, when the failure began, when a person or a control first noticed, and when customers stopped being affected. Circulate it for correction before anyone discusses causes.

Do we still need an alert once the bug itself is fixed?

Yes. Every RCA raises at least one detection action, because fixing the defect closes one hole and leaves you exactly as blind to the next one. Write three things down: the current time to detect taken from the timeline, the target you are committing to, and the specific signal that would have fired. Then name who is rostered to answer it, because an alert routed to a channel nobody owns is not detection. In the invented worked example on this page, card declines jump from 2% to 31% and a customer email is the first signal 41 minutes later, so the action is a decline-rate alert at 5% sustained over five minutes.

What should a root cause analysis report include?

A timeline agreed by everyone who was there, with a named source on every entry and both time to detect and time to restore stated. Then two to four contributing causes, each having survived the counterfactual test and each naming a control the team can change rather than a person. Then the detection action, written as the current number, the target and the signal that would have fired. Then every other action as a linked ticket with a named owner and a quarter. Anything that failed the counterfactual sits in a background paragraph and gets no ticket. Publish within five working days, where anyone in the organisation can read it without asking.

What changes in a root cause analysis when AI wrote the code?

On AI-native teams two conditions recur, and both are fixed in the outcome ticket rather than in the codebase. The first is a user journey the ticket never listed, so nothing tested it, and the fix is to add it to the standing list. The second is verification that rested on the model reporting it had finished, and the fix is to require a demonstrated run instead of a claimed one. On AI-augmented teams the failing control is usually a story's acceptance criteria or the human review that waved the change through, so the actions land in the definition of done and the test suite. Tenhaw's engineers build to a 72-rule engineering handbook, open source and enforced by an agent rather than by memory.

How should past incidents shape the next quarter's plan?

Reread the quarter's RCAs together before you open the next quarter's roadmap. A failure shape that appears twice is not bad luck, it is the tech debt you should be funding, and it will keep recurring until it is carried by something that competes for capacity. Missing tests and alerting go to the Tech Debt epic; anything large enough to need product prioritisation becomes an epic under an outcome, with a quarter and a planned currency share. Those epic types come from The Tenhaw Way, the operating model behind every Tenhaw engagement. Where an incident exposed a risk you are choosing to live with, put it in the RAID log as an accepted risk with a named owner and a review date.

Is a formal root cause analysis overkill for a small team?

Not at the size it actually runs, one person spending a day freezing evidence and then a ninety-minute session. The written trigger list is what keeps it proportionate, because most incidents never clear the bar and get a bug ticket with no ceremony at all. In small teams the real failure mode is the opposite of overkill. Only the embarrassing outages get investigated, while the small recurring failure that costs more in aggregate never earns anyone's attention, because the call is made in the mood straight after an outage rather than against a threshold agreed in advance. Ninety minutes beats paying for the same incident a third time.

All how-to guides

If the sources do not answer it, a call will.

Talk it through
book a call

Still have a question?

A 30-minute discovery call with James Rooney. Bring the question this page did not answer. You'll leave with a rough scope whether you engage us or not.

most start with a fixed-price AI Readiness Audit · £44,000 · 4 weeks · working prototypes

// pick a slot · cal.com/tenhaw/professional-servicesLIVE CALENDAR

Calendar not loading? Open it on cal.com or email hello@tenhaw.com.