How to Improve Test Stability and Handle Flaky Tests
Diagnose flaky tests by controlling state, timing, dependencies, and runner conditions. Use retries and quarantine only as visible, temporary mitigations.
A flaky test passes and fails intermittently under effectively unchanged code and inputs. Treat that result as a symptom to investigate: uncontrolled state, dependencies, timing, or execution conditions can make a real regression harder to distinguish from noise. First reproduce the failure and restore determinism; use retries or quarantine only as visible, temporary controls.
This guide covers a tool-agnostic workflow for diagnosing intermittent failures in unit, integration, browser, and CI tests. It does not assume a particular framework or CI provider, so it avoids version-specific retry flags.
1. Confirm that the test is flaky
Before changing the test or enabling retries, preserve enough context to compare runs:
- Record the commit or code revision, test name, runner, environment, and relevant configuration.
- Keep the complete failure output, logs, timestamps, and any artifacts the test produces.
- Rerun the suspect test on the same revision, first independently and then in its usual suite context.
- Compare a passing and failing run for differences in order, test data, timing, environment, resource use, and external responses.
A pass on retry is evidence of intermittency, not proof that the failure was harmless. A flaky result weakens the signal that a failing test indicates a regression, and the same intermittent test can also hide a genuine defect.
2. Find the uncontrolled condition
Work through these causes before deciding that the test merely needs more time or another attempt.
Shared or stale state
Look for database rows, files, queues, caches, environment variables, static fields, singletons, and shared fixtures that survive between tests. Verify that each test starts from known data and that setup and teardown complete even when an assertion or action fails.
Where practical, rebuild a known starting state for each test. Cleanup can be cheaper for a large fixture, but it is harder to reason about if cleanup is incomplete or skipped. Choose based on the cost of setup and the confidence that each test remains isolated.
Order dependence
Run the suspect test by itself and in different positions or subsets of the suite. If its result changes with the order, identify which earlier test or shared resource changes its assumptions. Tests should not depend on another test having run first; isolation lets the suite run in different sequences without changing outcomes.
Uncontrolled time
Wall-clock reads can cross midnight, a month boundary, or an expiry threshold between runs. They can also disagree with fixture timestamps. Put time access behind a controllable seam and set a fixed clock in tests. If the behavior under test is specifically about elapsed time, advance a controlled clock rather than waiting for real time.
Asynchronous races
For asynchronous work, wait for a defined application state: a particular record, event, element, or completed job. Set a bounded timeout and include useful context in the timeout failure. A fixed sleep may pass on a fast run and fail on a slow one; increasing it can also lengthen every run without establishing that the required condition occurred.
External services and dependencies
A remote API, third-party service, or shared test service introduces behavior and timing outside the test’s control. For stable regression coverage, use a test double at the appropriate boundary. Add contract or integration checks when needed to compare the double’s important behavior with the real interaction. Doubles improve repeatability but do not provide the same end-to-end fidelity as the real dependency.
Runner resources and environment
Inspect runner logs and environment assumptions. Slow or overloaded machines, incomplete setup, differing configuration, and resource contention can change execution timing or starve the system under test. Make prerequisites explicit, use a consistent environment where practical, and allocate sufficient resources for the workload.
3. Repair the cause and verify the fix
- Write down the failure condition. State what differed between the passing and failing runs, or what hidden assumption the test made.
- Choose the smallest control that restores determinism. Reset the relevant state, isolate the fixture, control the clock, synchronize on a real condition, stabilize a dependency, or correct runner setup.
- Repeat the original reproduction. Run the test alone and in the suite context that exposed it, on the same revision and under the same conditions.
- Check that the test still catches the intended defect. A test that always passes because its assertion was weakened is stable but no longer useful.
- Record ownership and evidence. Capture the root cause, fix, and any remaining uncertainty so a future intermittent failure can be compared against this one.
Prefer hermetic, explicitly configured tests when feasible. The wider the test’s uncontrolled environment, the more likely it is that a change outside the code under test will alter its result.
4. Use retries and quarantine carefully
Retries can help identify an intermittent failure or keep a pipeline moving while an owner investigates. They do not make the underlying test deterministic: a passing retry does not establish correctness. Keep the initial failure visible in logs and reports, and track repeated intermittent failures as work.
If quarantine is necessary to protect the main suite’s signal, make the test visible, assign an owner, set a repair deadline, and schedule it to run outside the blocking path. Review the quarantine list regularly. Without ownership and a time limit, quarantine can become abandonment.
Choose a remedy based on its tradeoff:
| Choice | Benefit | Cost or limitation |
|---|---|---|
| Test double | More repeatable coverage of the code’s response | Less direct fidelity to the real dependency; supplement with contract or integration checks when warranted |
| Rebuild fixture | Known starting state that is easier to reason about | Can take longer than cleanup for a large fixture |
| Retry | Can reduce immediate pipeline disruption and expose intermittent behavior | Can conceal a real regression if the original failure is hidden |
| Quarantine | Keeps a known unstable test from obscuring the blocking suite signal | Coverage is outside the main path until repaired; requires visible ownership and follow-up |
5. Troubleshooting common flaky-test patterns
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Passes alone, fails in the suite | Shared state or order dependence | Run in different orders; inspect fixtures, static state, databases, files, and cleanup. |
| Fails only on slower CI runners | Timing assumption or insufficient resources | Wait for a specific condition with a timeout; inspect runner load and setup logs. |
| Fails near a date or time boundary | Wall-clock dependency | Control the clock and make fixture timestamps consistent. |
| Longer sleep appears to help | Race condition still present | Replace the delay with synchronization on the state the test actually needs. |
| Retry passes but first attempt fails | Intermittency remains | Keep the first failure visible and investigate data, timing, dependencies, and resources. |
| Works locally, fails in CI | Environment mismatch or runner constraint | Compare configuration, dependencies, setup, permissions, available resources, and logs. |
| Occasional third-party errors | External service behavior or availability | Use a controlled double for regression coverage and maintain integration coverage where real behavior matters. |
6. Keep browser-based checks observable
Browser tests can fail because the page state is not ready, an external resource changes, a consent dialog obscures content, or a capture depends on timing. Make the page and its inputs predictable where possible, and wait for a specific state before asserting or capturing. If you need screenshots to inspect a failure, capture the relevant state consistently and retain them with the test’s revision and failure details.
When the goal is simply to obtain a clean screenshot of a URL for a report or debugging artifact, you can also use a screenshot API. That is separate from fixing the test’s underlying nondeterminism.
7. Performance, reliability, and cost
Repeated setup, retries, and arbitrary waits all consume runtime. Rebuilding fixtures may improve isolation but cost more for large datasets; cleanup can be faster but risks leaving stale state. A specific condition wait can finish as soon as the condition is true, while a fixed delay always spends its full duration. Measure the suite before and after a change and keep reliability work focused on the source of variation.
Hermetic environments and controlled dependencies reduce the number of outside conditions that can change a result. They do not remove the need for integration coverage when behavior across a real boundary matters. Retries and quarantine should be tracked as operational costs because they consume CI capacity and weaken the failure signal if left unmanaged.
Or skip the browser setup
For a screenshot artifact, [ScreenshotNeo](https://screenshotneo.com) returns a PNG, JPEG, WebP, or PDF from one API request. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; the response reports the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000 screenshots. Sign up for free and get 1,000 screenshots a month with no card.
Frequently asked questions
Is a flaky test necessarily a bad test?
No. It may cover important behavior while relying on an uncontrolled condition. Preserve its intended assertion and make that condition explicit or controllable.
Should every test use a mock?
No. Use a double where repeatable regression coverage needs control over a dependency, and keep appropriate contract or integration checks for real interactions.
How do I know when to remove a retry or quarantine?
Remove the mitigation after the root cause is fixed and the test remains reliable in the context that previously exposed the failure. Keep the failure history available for comparison.
Are there reliable industry-wide flakiness rates?
The cited research includes historical figures from Google’s own test corpus, not a current industry-wide estimate. Avoid applying those figures to another team’s suite.


