Anomaly Reports in Software Testing: How to Find and Fix Test Issues
Turn failed or intermittent tests into actionable reports: preserve evidence, trace likely causes, assign a remedy, and verify the fix.
A failed test is a signal to investigate, not proof that the product code is defective. A useful anomaly report preserves the execution context, shows whether the failure is new or recurring, records the evidence and likely cause, assigns a next action, and links the result to the defect or work item when appropriate. Then rerun the affected test and review later executions to confirm the fix and watch for recurrence.
This guide uses “anomaly report” to mean a report or investigation record for an unexpected test result, including consistent failures and intermittent (flaky) failures. The same workflow applies whether your team records results in a test management system, an issue tracker, or a CI pipeline.
1. Capture the evidence before changing anything
Start with the exact execution that failed. Preserve enough detail for someone else to reproduce or investigate it without relying on your memory.
- Test identity and outcome: test name or ID, suite, expected result, actual result, and whether it failed, was skipped, timed out, or was marked inconclusive.
- Execution context: build or release, branch, commit or change list, pipeline or run ID, timestamp, runner or agent, and relevant environment details.
- Failure evidence: complete error message and stack trace, failed step, assertion values, logs, screenshots or video when available, and attachments that help explain the state.
- Reproduction details: setup, test data, steps to reproduce, prerequisites, and whether the failure occurred on a rerun.
- Initial analysis: what changed recently, what has already been ruled out, and the next diagnostic action.
Keep raw evidence where the team can retrieve it, and link the test result to the related bug or work item. For example, Azure DevOps test-run summaries can include linked work items, step outcomes, automated-run stack traces, analysis information, and attachments. Those are product-specific capabilities, but the underlying need to preserve context applies to any workflow. See Microsoft’s test runs documentation.
Example report template
Title: [test or behavior] fails [consistently/intermittently] in [context]
Test ID/name:
Outcome: Failed / Timed out / Skipped / Other
Expected:
Actual:
Build or release:
Branch and commit/change:
Pipeline or run ID:
Timestamp and time zone:
Runner, OS, browser/device, and relevant environment:
Test data and prerequisites:
Steps to reproduce:
Failure message and stack trace:
Logs, screenshots, video, and attachments:
Recent related changes:
Runs reviewed and observed pattern:
Current hypothesis (label as unconfirmed until supported):
What has been ruled out:
Owner and next action:
Linked bug or work item:
Verification plan:
Separate observations from hypotheses. “The test failed after waiting 10 seconds for the confirmation element” is evidence; “the confirmation service is broken” is a hypothesis until the investigation supports it.
2. Decide whether the issue is new, recurring, or intermittent
A single failure cannot establish a trend. Inspect multiple executions across a useful time window, including successful runs, and keep the test’s history visible. Ask:
- When did the first failure occur? Which build, branch, or change was running then?
- Does it fail every time, only on particular runners or environments, or unpredictably under similar conditions?
- Do failures share a message, stack frame, failed step, time of day, test data, or dependency?
- Did nearby tests fail too, suggesting a shared setup, service, or resource issue?
- Does the failure disappear when run alone, or follow another test in the suite?
Use result history and execution-instance details where your test platform provides them. Azure DevOps Test Analytics, for example, supports views of top failing tests and drill-down into execution instances; see Test Analytics. Treat its reporting ranges and interface details as specific to that product and verify current settings in the tool.
3. Trace the likely cause before editing the test
Investigate four broad cause groups. They overlap, so follow evidence rather than assigning blame based on the red/green result alone.
| Cause group | Questions to ask | Evidence to look for |
|---|---|---|
| Product code | Did the behavior or contract change? Is the expected result still correct? | First failing change, application logs, reproducible behavior outside the test, related defects. |
| Test code or data | Are assertions, fixtures, selectors, mocks, or data assumptions wrong or stale? | Failure at the same assertion, invalid or reused data, brittle assumptions, changed UI/API contract. |
| Runner, dependencies, or environment | Is the runner healthy and correctly configured? Did a service, network, OS, browser, or resource limit change? | Correlated failures, runner logs, dependency status, resource pressure, environment-specific patterns. |
| Nondeterministic behavior (flakiness) | Can the same code pass and fail depending on timing, order, state, or external conditions? | Mixed results at the same code, order-sensitive failures, races, variable dependency responses. |
Microsoft’s troubleshooting guidance lists source-under-test errors, bad test code, environment problems, and flaky tests among possible causes. Its traceability guidance also describes following persistent failures back to changes where they began. See test failure troubleshooting and traceability and test failures.
4. Find the root cause of a flaky test
A flaky test passes and fails with the same code under conditions that appear equivalent. Do not treat rerunning until green as proof that the issue is fixed. Use reruns as diagnostic evidence and preserve each result.
Check state, setup, and cleanup
- Run the test independently. If it passes alone but fails in the suite, look for order dependence, shared state, leaked resources, or a previous test that changes data.
- Check initialization and teardown paths, including what happens after an assertion fails or a timeout occurs.
- Verify test data is created deliberately, unique where necessary, and reset or cleaned up reliably.
- Review hidden assumptions about test order, cached state, shared accounts, ports, files, queues, and external services.
Check timing and synchronization
- Determine what event the test is actually waiting for: a request completing, a state transition, an element becoming actionable, or a durable result being written.
- Prefer waiting for a meaningful condition or application state over sleeping for an arbitrary duration.
- Record timestamps around the action and the observed state. This helps distinguish a race from a slow or unavailable dependency.
- Check whether asynchronous work continues after the test thinks it has finished.
Check resources and dependencies
- Compare failures across runner machines, OS or browser versions, network paths, and dependency versions.
- Look for resource exhaustion, contention, rate limits, service instability, and shared infrastructure failures.
- Review test framework and runner logs as well as application logs. The test harness itself can introduce failures.
Google’s guidance groups potential sources across the test, test-running framework, application and dependencies, and OS, hardware, or network. It recommends examining setup and teardown, data assumptions, independent execution, and access timing. It also cautions that arbitrary delays can make tests both flaky and slow. Read Google’s flaky-test guidance. Its historical observation that about 1.5% of Google test runs were flaky is specific to Google and is not a general industry rate.
5. Choose a remedy and assign ownership
Once the evidence supports a cause, record the decision, owner, and next action. A report should lead to work that can be completed and checked.
- Product defect: create or link a defect with impact, severity, owner, and reproduction evidence. Keep the failed test result associated with it.
- Test defect: correct invalid assertions, stale selectors or contracts, bad fixtures, shared state, or unsafe data assumptions.
- Race or timing issue: synchronize on application state and make completion conditions explicit. Avoid masking uncertainty with a longer fixed sleep.
- Environment or runner issue: fix configuration, resource limits, dependency availability, or environmental assumptions; record affected environments.
- Known intermittent issue awaiting repair: mark or annotate it only through the team’s documented process, retain its history, and make the follow-up owner and review point visible.
- Duplicate reports: link reports that share a supported root cause to one underlying defect, while retaining each run’s evidence and impact.
Do not “fix” a product regression by weakening an assertion unless the expected behavior has actually changed and the requirement or contract is updated. Likewise, suppressing or quarantining a flaky test may help a team manage a pipeline, but it does not repair the underlying cause; retain a route to investigate and resolve it.
6. Verify the fix and keep recurrence visible
- Run the affected test after the change under the conditions that previously exposed the failure.
- Run it independently and in its normal suite if order or shared state was a possibility.
- Review the new result’s logs and evidence; a pass alone may not prove the intended behavior was checked.
- Review subsequent executions for recurrence across relevant runners or environments.
- Update the linked report with the fix, verification runs, remaining limitations, and closure decision.
Azure DevOps documents workflows for detecting, marking, reporting, and later unmarking flaky tests after resolution or manual review. Its documentation notes that changing a flaky designation affects future executions rather than retroactively changing the current pipeline result. Keep such tool behavior distinct from the general principle: historical results remain important evidence. See Microsoft’s flaky test management documentation.
7. Use screenshots as evidence when the failure is visual
For browser and UI tests, a screenshot can show the rendered state at the failure point: a missing control, unexpected dialog, layout shift, or consent overlay. Capture at the failed step and preserve the URL, viewport or device, timestamp, and relevant test/run identifier alongside the image. A screenshot supplements logs and assertions; it does not establish why the state occurred by itself.
If a capture needs to show a specific control or state, capture the relevant element or page at the failure point, and avoid including unrelated personal or secret data. Where cookie banners or chat widgets obscure the interface, note whether they are part of the expected test state or noise to remove from diagnostic evidence.
8. Common reporting and diagnosis mistakes
| Problem | Why it hurts | Better action |
|---|---|---|
| “Test failed” with no run details | Investigators cannot reliably find the execution or its context. | Include test ID, pipeline/run, build, commit, timestamp, environment, and evidence links. |
| Calling every failure a product bug | Test logic, data, infrastructure, and nondeterminism can also cause failures. | Label the suspected cause as a hypothesis and gather evidence across the cause groups. |
| Rerunning until the test passes | A later pass can hide an intermittent problem and does not explain the first failure. | Keep all executions, compare conditions, and use reruns to isolate a cause. |
| Adding a fixed delay immediately | It can slow the suite and still fail when the environment is slower than expected. | Wait on an observable condition and log timing around it. |
| Changing or disabling an assertion to restore green | A real regression may become invisible. | Confirm the intended contract with requirements or owners before changing expected behavior. |
| Closing without later history review | The issue may recur unnoticed in a different run or environment. | Record verification and review subsequent executions for recurrence. |
9. Troubleshooting: common test anomaly symptoms
| Symptom | Likely causes to investigate | What to do next |
|---|---|---|
| Same test fails on every run | Product regression, stale test expectation, invalid fixture, deterministic environment issue. | Reproduce on the same commit; inspect the assertion and first failing change; compare actual and expected behavior. |
| Test passes alone but fails in the suite | Order dependence, shared state/data, incomplete cleanup, resource contention. | Vary preceding tests, isolate shared resources, and inspect setup/teardown and data reset. |
| Timeout with no useful stack trace | Wait condition never became true, dependency stalled, runner overloaded, or timeout handling discarded context. | Capture runner and application logs, timestamps, dependency status, and the exact awaited condition; improve timeout diagnostics. |
| Failure only on one runner or environment | Configuration drift, browser/OS difference, resource limit, network or dependency path. | Compare runner images and settings; rerun on a known-good runner and inspect environmental differences. |
| Intermittent failure after an action | Race, asynchronous completion, variable network or service timing. | Wait for a meaningful state transition; log action and observation times; inspect whether work outlives the test. |
| Many tests fail together | Shared service, test data setup, deployment, credentials, network, or infrastructure issue. | Group by time and common dependency; check environment and setup before treating every test as a separate defect. |
| Failure disappears on rerun | Flakiness, transient infrastructure, changed state, or order/timing effect. | Preserve both runs, compare context, and investigate the condition that differed; do not close solely because of a pass. |
10. ScreenshotNeo: capture browser evidence without browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. A single GET request can return a PNG, JPEG, WebP, or PDF. For visual test investigations, it can capture a page or element, set a viewport or device, wait for a selector or network idle, apply custom CSS or JavaScript, and control cookies and headers. See ScreenshotNeo and its API documentation.
Use this after deciding what page state is useful evidence and ensuring the target URL does not expose private test data. Replace the example URL with an appropriate test page. Keep the response with the test run context.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For an automated evidence pipeline, use the documented options for full-page or selector capture, viewport and device, waiting, output format, and any required headers or cookies. ScreenshotNeo also supports asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, caching with a chosen TTL, and signed links for public image tags. Its response headers identify the page verdict and billing status, which can help a pipeline distinguish a clean capture from a bot check, blank page, failed load, or cache hit. Consult the docs for parameter names and configuration.
Or skip the browser setup
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include page-verdict and billing headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Sign up for 1,000 free screenshots a month, with no card required.
11. Reliability, performance, and cost considerations
- Evidence reliability: retain a run identifier and timestamp with each artifact. Keep screenshots, logs, and result records tied together so they are not mistaken for evidence from another execution.
- Capture timing: capture at the relevant failure step. Waiting for the right state can improve evidence quality; arbitrary delay adds time and may still capture the wrong state.
- Test suite cost: repeated reruns can consume CI capacity and mask a pattern. Choose reruns to answer a diagnostic question, then make the test deterministic or fix the confirmed cause.
- Screenshot API cost: ScreenshotNeo bills only clean shots. Failed loads and other listed non-clean outcomes, as well as cache hits, are not billed. Its monthly options are Free (1,000), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free. Every feature is on every plan.
- Privacy: screenshots and logs can contain account details, tokens, or personal data. Use test accounts and non-sensitive data, control access to artifacts, and avoid placing secrets in URLs or captured content.
12. A practical quality checklist
- [ ] The report identifies the exact test and execution.
- [ ] Build, branch/change, timestamp, runner, and relevant environment are recorded.
- [ ] Failure details and useful attachments are preserved.
- [ ] Multiple runs were reviewed before describing an issue as intermittent or fixed.
- [ ] Product, test/data, runner/dependency, and nondeterministic causes were considered.
- [ ] A cause is supported by evidence or clearly labeled as a hypothesis.
- [ ] The next action has an owner and is linked to a defect or work item when appropriate.
- [ ] Verification includes the conditions that exposed the issue and later history review.
Frequently asked questions
How do I write a good test failure report?
Identify the test and run, preserve expected and actual results, include the build and environment, attach logs or traces, explain how to reproduce it, and state the next action and owner. Link the report to the relevant defect or work item.
Why does my test pass sometimes and fail other times?
Likely causes include order or shared-state effects, timing races, variable test data, runner or dependency instability, and environmental differences. Compare passing and failing executions at the same code and look for the condition that changes.
Should I rerun a failed test?
Yes, when the rerun helps isolate state, order, or environmental effects. Keep both results and investigate a failure that disappears on rerun; a pass does not explain the original failure.
When should a test failure become a bug?
When evidence indicates product behavior violates its requirement or contract, record a defect and link the result. If the evidence points to test code or infrastructure, track that corrective work with its own owner and history.
Should flaky tests be disabled?
Follow the team’s pipeline policy if a test must be isolated temporarily, but preserve its history, record the reason and owner, and track a resolution. Disabling it alone does not identify or fix the cause.


