Root Cause Analysis in Software Testing
A practical, evidence-led guide to investigating escaped software defects, understanding why tests missed them, and preventing recurrence.
Root cause analysis (RCA) in software testing is a structured, evidence-led investigation into why a defect occurred, how it escaped detection, and what changes will reduce the chance of recurrence. Start by defining the observed failure precisely, reconstruct the timeline, examine the test escape, map causes and contributing factors, then assign and verify corrective actions. The goal is an explanation supported by evidence and a change that addresses the conditions behind the defect—not simply a code patch or a count of how many times the team asked “why.”
NASA describes RCA as a structured evaluation to identify causes and actions adequate to prevent recurrence, and recommends clearly defining the problem, building a timeline, distinguishing root causes from other causal factors, and establishing causal relationships. AWS likewise recommends assessing why existing tests did not find an issue and adding tests for the case when they are missing. NASA Software Engineering Handbook, SWE-204; AWS Well-Architected, REL12-BP02.
1. Define the failure before explaining it
Write down what happened, what should have happened, who or what was affected, and under what conditions. Keep this description separate from hypotheses about why it happened. “Checkout failed” is too vague to guide an investigation; “A retry after a payment timeout created a second order when the first authorization had succeeded” is more useful if the evidence supports those details.
Capture the following, marking unknowns explicitly:
- Observed behavior: exact output, state transition, error, or missing response.
- Expected behavior: the relevant requirement, contract, or user-visible outcome.
- Impact and severity: affected users, data, workflows, safety or business consequences, duration, and scope.
- Operating context: release/build, configuration, dependencies, platform, inputs, traffic pattern, and triggering conditions.
- Evidence sources: logs, traces, metrics, test results, deployment records, reports, and reproducible examples.
Use neutral, specific language. “The input was not validated before the service used it” describes a condition. “A developer forgot validation” assigns blame before the team has established why the validation gap existed.
2. Preserve evidence and build an event timeline
Preserve relevant logs, traces, test reports, artifacts, configuration, and deployment metadata before they expire or are overwritten. Follow access-control and privacy rules when collecting production data; redact secrets and personal information from shared investigation records.
Build a timeline from the last known-good behavior through detection, response, and recovery. Include both technical events and decisions that could matter:
| Time or sequence | Event | Evidence | Confidence / open question |
|---|---|---|---|
| Before release | Requirement, design, dependency, or configuration changed | Change record, review, artifact | Was this change present in the affected build? |
| Verification | Tests were designed, selected, run, skipped, or failed | CI logs, test reports, environment record | Did the relevant case run against this build? |
| Occurrence | Trigger and faulty behavior occurred | Trace, logs, user report, reproduction | Can the trigger be reproduced? |
| Detection and response | Alert, report, mitigation, rollback, or fix | Incident record, deployment history | Which action changed the behavior? |
Sort by timestamp where clocks are reliable; otherwise record sequence and uncertainty. Note gaps, time-zone differences, sampling, delayed telemetry, and any changes made during mitigation. A timeline is a reconstruction, not proof that adjacent events caused one another. NASA recommends tracing from normal operation to failure and annotating milestones, tests, contributing events, and decision points.
3. Investigate why the tests missed the defect
“There was no test” is only a starting point. Identify what test condition could have exposed the failure and why the verification system did not reveal it. Consider the full chain from test basis to feedback:
| Area | Questions to ask | Possible corrective direction |
|---|---|---|
| Requirements and test basis | Was the behavior specified? Was a relevant risk or boundary case omitted? | Clarify the requirement or add an acceptance criterion. |
| Test design and oracle | Was the triggering sequence represented? Would the assertion detect the wrong result? | Add a scenario and assert the consequential state, not merely that a request returned. |
| Test data | Did data include the empty, duplicate, malformed, expired, or high-volume case? | Add controlled boundary and stateful data. |
| Environment and dependencies | Did the test use the relevant browser, service version, configuration, timing, or failure mode? | Make the environment representative or explicitly test the difference. |
| Execution and selection | Did the test run on the affected path and build? Was it skipped, quarantined, or excluded? | Fix selection, gating, scheduling, or ownership. |
| Feedback and triage | Was a failure masked, flaky, ignored, or hard to diagnose? | Improve signal, failure reporting, and flaky-test handling. |
Then add the smallest reliable regression test that would fail before the fix and pass after it. Prefer a test at the lowest level that can faithfully reproduce the causal behavior, with a higher-level test if integration or user-visible behavior is part of the failure. Do not add a brittle test that merely encodes incidental implementation details.
4. Separate root causes, direct causes, and contributing factors
A direct or proximate cause is close to the failure in time or execution, such as a retry issuing a non-idempotent operation. A contributing factor increases the likelihood or impact, such as an unusually slow dependency. A root cause is a deeper condition whose removal or modification would prevent recurrence of the defined problem. There can be multiple root causes; do not force a single explanation when the evidence supports interacting causes.
For each proposed cause, record:
- The supporting observation or artifact.
- The causal mechanism connecting it to the failure.
- Whether it is confirmed, plausible, or still unverified.
- What evidence would disprove it.
- Whether changing the condition would prevent the failure or only reduce its impact.
Use a causal graph or a cause-and-effect tree when several conditions interact. For example: a retry policy resubmits a request; the endpoint lacks idempotency protection; the test suite covers timeout errors but not a timeout after successful processing; the duplicated state reaches users. Treat each arrow as a claim to validate, not as established simply because it appears in a diagram.
5. Choose an RCA technique that fits the evidence
| Technique | Useful when | Limit |
|---|---|---|
| Five Whys | A well-defined issue appears to have a short causal chain and the team can test each answer. | Can create a misleading single chain when causes branch. Five questions are not a required stopping point. |
| Fishbone / Ishikawa | The group needs to organize candidate causes across requirements, design, code, test data, environment, and process. | Organizes hypotheses; it does not prove a branch caused the defect. |
| Causal graph / cause-effect tree | Multiple events or conditions combine, and the team needs to show dependencies. | Distinguish observed facts from inferred links. |
| Counterfactual causal testing | Execution-level evidence is available and the team can compare conditions or executions to test which changes alter buggy behavior. | Requires suitable executions and expertise; published results are bounded to the study’s benchmark and experiment. |
In a 2018 study, researchers reported that Causal Testing applied to 71% of real-world defects in the Defects4J benchmark; among those applicable defects, it helped identify the root cause for 77%. In a controlled experiment with 37 developers, cause identification was reported in 86% of cases with Causal Testing and 80% with standard testing tools. These figures describe that paper’s benchmark and experiment, not expected results for any particular team or defect. Johnson, Brun, and Meliou, “Causal Testing: Finding Defects’ Root Causes” (2018).
6. Turn findings into corrective actions
Each action should address a cause found in the analysis, have a named owner and due date, and include completion evidence and an effectiveness check. Fixing the defective line may be necessary, but may not address why the defect entered or escaped the system.
| Finding | Action | Owner / due date | Completion evidence | Effectiveness check |
|---|---|---|---|---|
| Retry after ambiguous timeout can duplicate processing. | Enforce idempotency and add a regression test for timeout after successful processing. | Named owner / date | Reviewed change and passing test in CI. | Test remains in the required pipeline; monitor duplicate outcome signal. |
| Relevant browser state was absent from verification. | Add the missing state to the test fixture and assert the user-visible outcome. | Named owner / date | Fixture, test, and recorded run. | Confirm the test catches a deliberately reintroduced failure in a safe development environment. |
Actions may include a regression test, a clearer requirement, improved environment control, review guidance, an automated guardrail, monitoring, or a response procedure. Select based on the causal evidence. Avoid actions such as “be more careful” unless they are paired with a concrete change to the workflow and a way to assess whether it works. AWS cautions against stopping at contributing factors or identifying only human error without mitigation or automation. NASA calls for recording findings and tracking corrective actions to closure.
7. Keep the review blame-free and close the loop
Ask participants what they observed, what information they had at the time, and what constraints shaped their decisions. Describe actions and system conditions without turning the review into a judgment about an individual. Blame can discourage people from sharing evidence and uncertainty; a blameless review still expects precise explanations and follow-through.
Publish a concise record that includes the problem statement, impact, timeline, evidence, confirmed causes and contributing factors, test escape analysis, actions, owners, and follow-up date. Share relevant lessons with teams that may have the same exposure. Revisit the actions to check both completion and effectiveness; “merged” does not prove the risk was reduced. AWS recommends documenting findings, reviewing why tests missed the issue, and sharing corrective actions so other workloads can mitigate similar factors.
What should a software root cause analysis include?
Use this practical report outline as a starting point. Scale its formality to impact and risk; high-severity or safety-critical defects warrant more rigorous evidence and review.
- Summary: concise description and current status.
- Impact: affected functions, users, data, time window, and severity.
- Expected and observed behavior: precise reproduction conditions where possible.
- Timeline: relevant changes, tests, occurrence, detection, and response.
- Evidence: links to artifacts, logs, traces, test runs, and known gaps.
- Cause analysis: direct cause, root cause(s), contributing factors, causal reasoning, and confidence.
- Test escape: which condition was absent, missed, unexecuted, or poorly asserted.
- Corrective actions: action, owner, due date, status, evidence, and effectiveness measure.
- Sharing and follow-up: affected teams or components to review and review date.
Standards context
ISO/IEC/IEEE 29119-1:2022 presents general concepts in software testing. It provides testing-process context, not a dedicated RCA procedure. ISO/IEC 30130:2016 concerns categorizing software test entities and testing-tool capabilities; it can inform tool assessment, but it does not prescribe an RCA workflow. Follow your organization’s applicable quality, safety, incident, and regulatory requirements where those impose additional process.
Or skip the browser setup
If visual behavior is part of the defect, a screenshot can preserve what a page looked like during a reproduction or help document a rendering regression. ScreenshotNeo is a website screenshot API and MCP server. Its API call can capture a page without setting up browser automation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo API documentation
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.
Common problems and fixes
| Problem | Likely cause | Next step |
|---|---|---|
| The team agrees on a cause but cannot show supporting evidence. | A plausible story has been treated as a finding. | Label it a hypothesis, seek a reproduction or artifact, and record what would falsify it. |
| The analysis ends at “the test was missing.” | The test gap is identified but the conditions that created the gap are not. | Ask why the case was absent from the requirement, design, test selection, or review, then choose a systemic action. |
| A regression test passes but the bug still occurs in production. | The test does not match the trigger, state, dependency, environment, or assertion that mattered. | Compare production evidence with the test setup and improve fidelity or add a focused integration test. |
| Many possible causes, no clear priority. | Candidate explanations are mixed with evidence and impact. | Map causal links, rank by evidence and ability to explain the observed behavior, and test the highest-value uncertainty first. |
| Actions remain open after the report is published. | Actions lack owners, due dates, closure evidence, or review cadence. | Track them in the team’s normal work system and schedule an effectiveness review. |
| Five Whys produces “human error” as the answer. | The chain stops at an individual action instead of investigating conditions and controls. | Ask what information, interface, safeguards, workload, or process allowed the action to produce an undetected defect. |
| The failure cannot be reproduced. | Trigger data may be missing, nondeterministic, environment-specific, or lost. | Preserve remaining telemetry, compare affected and unaffected runs, document uncertainty, and add observability or controlled reproduction before asserting a cause. |
Performance, reliability, and cost of RCA
RCA consumes engineering time, so match investigation depth to severity, recurrence risk, and uncertainty. A small, low-impact defect with a direct, reproducible cause may need a short record and regression test. A severe or recurring failure may justify broader log and trace analysis, cross-team review, causal mapping, and follow-up audits. NASA’s SWE-204 guidance is particularly framed around high-severity non-conformances and closed-loop process assessment.
Improve investigation reliability by preserving evidence early, recording uncertainty, checking causal claims against a reproduction or counterexample, and verifying corrective actions later. Avoid collecting every possible log indefinitely: define the time window and signals that can distinguish hypotheses, protect sensitive data, and retain only what policy permits. Treat investigation tools as aids; a dashboard, diagram, standard, or automated test cannot establish a causal relationship without evidence.
FAQ
Is root cause analysis the same as debugging?
No. Debugging locates and fixes faulty behavior. RCA also investigates why the defect was introduced or escaped and what will reduce recurrence.
How many “whys” should we ask?
There is no required number. Stop when the explanation is evidence-supported, addresses relevant process conditions, and leads to actions whose effectiveness can be checked. Follow additional branches when causes interact.
Should every bug get a formal RCA?
Not necessarily. Use a level of investigation proportionate to impact, recurrence, uncertainty, and risk. Formal requirements may apply to high-severity or regulated systems.
Does adding a regression test prevent recurrence?
It prevents the same observable behavior from escaping that test when the relevant conditions and assertions are represented. It may not address a deeper requirement, environment, or process weakness by itself.
What is the difference between a root cause and a contributing factor?
A root cause is a deeper condition whose modification would prevent recurrence of the defined failure. A contributing factor influences the likelihood or impact but may not explain the underlying weakness on its own.
References
- NASA Software Engineering Handbook, SWE-204: Process Assessments.
- AWS Well-Architected Framework, REL12-BP02: Perform post-incident analysis.
- ISO/IEC/IEEE 29119-1:2022: Software testing—General concepts.
- ISO/IEC 30130:2016: Capabilities of software testing tools.
- Johnson, Brun, and Meliou, “Causal Testing: Finding Defects’ Root Causes” (2018).


