How to Find and Fix Flaky Tests
Find the uncontrolled dependency behind intermittent test failures, fix it, and restore a reliable regression signal without hiding coverage.
A flaky test passes and fails on the same code and inputs because something relevant to its result is uncontrolled. A rerun can help confirm the symptom, but it does not fix the cause. Record the failure conditions, isolate the test, find the changing dependency, control it, and verify the repair both alone and in the suite.
Flakiness is a symptom, not evidence that the application is correct or the test can safely be ignored. Shared state, asynchronous timing, clocks, external services, browser behavior, and resource leaks can all change a test’s result. The reliable fix is to identify which condition changed and make the test independent of it where possible. Martin Fowler’s guide to eradicating nondeterminism and Mike Bland’s discussion of nondeterministic tests describe the same core problem.
1. Confirm that the failure is intermittent
Before calling a test flaky, compare runs on the same revision and relevant environment. A test that fails after a code, input, dependency, or environment change may be exposing a reproducible regression.
For each failure, save:
- The test name, assertion or error, and complete failure output.
- The commit or revision, runtime and browser versions where applicable, and relevant environment details.
- Whether it fails in isolation, in its suite, or only under parallel execution.
- Its position in the test order, random seed if the runner uses one, and whether a rerun of the same revision passed.
- Relevant logs and state, such as fixture identifiers, timestamps, network results, or browser console errors.
A pass on rerun is a useful clue that the outcome is nondeterministic; it is not a repair. A rerun after changing the environment or code does not establish that the original failure was flaky.
2. Reproduce the failure in a controlled way
- Run the failing test alone. Keep the revision, configuration, and inputs unchanged. Repeat enough to observe whether the failure can recur; there is no universal run count that proves a test is flaky.
- Run it in its normal suite. If it passes alone but fails in the suite, suspect order dependence, shared fixtures, leaked state, cleanup, or parallel collisions.
- Start from a known state. Reset or recreate database records and other mutable fixtures where practical. Check whether a prior test leaves static data, singleton state, files, or service configuration behind.
- Vary one factor at a time. For example, compare serial and parallel runs, a clean and reused database, or controlled random seeds. Avoid changing several conditions at once; that makes the result hard to interpret.
- Capture evidence around the failure. Log the awaited condition, timestamps, fixture keys, response status, or browser state that directly relates to the assertion. Keep logs focused enough to compare passing and failing runs.
Do not assume that repeatedly running a large suite is the best first experiment. Narrowing the test and controlling the surrounding conditions usually gives more useful evidence.
3. Check the most common sources of nondeterminism
Shared state and test order
Tests may read or modify the same database row, static variable, singleton, file, queue, or account. A test can then pass alone but fail after another test changes the shared resource. Incomplete setup or faulty teardown can have the same effect.
- Give each test unique records and resource names where possible.
- Build the starting state explicitly instead of relying on data left by earlier tests.
- Check teardown paths, including failures partway through setup or execution.
- Use transaction rollback when the test can run without committing and the database setup supports it.
- For parallel failures, look for collisions in fixed filenames, account IDs, ports, queues, or shared records.
Rebuilding fixture state is often easier to reason about when its cost is acceptable. Shared immutable fixtures or cleanup may be practical when setup is expensive, but cleanup itself must be reliable. A cleanup error can make the next test look like the source of the problem. Fowler discusses these isolation risks in Eradicating Non-Determinism in Tests.
Asynchronous work and fixed sleeps
A fixed sleep assumes that an operation will finish within a guessed interval. If it takes longer on a slower machine or under load, the test fails; if it usually finishes sooner, the test wastes time. Prefer an event callback when the system supports one, or bounded polling for the condition the assertion actually needs.
async function waitFor(condition, { timeoutMs = 5000, intervalMs = 50 } = {}) {
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
if (await condition()) return;
await new Promise(resolve => setTimeout(resolve, intervalMs));
}
throw new Error(`Condition was not met within ${timeoutMs} ms`);
}
// Example: wait for the observable result, not an assumed duration.
await waitFor(() => page.locator('[data-state="complete"]').isVisible());
Adapt the example to the test framework’s condition and locator APIs. Ensure the condition is safe to check repeatedly, set a timeout appropriate to the operation, and include enough context in the timeout error to diagnose a missing response. A bounded wait is not a cure if it polls the wrong condition or the operation has no reliable completion signal.
Time and date boundaries
Tests that read the wall clock can behave differently around midnight, month or year boundaries, daylight-saving transitions, or clock skew. Pass a clock or fixed time into the code under test when feasible. If a test must cover a time boundary, specify the instant and timezone explicitly and assert the intended boundary behavior.
Remote services and network conditions
A third-party service, DNS lookup, network connection, rate limit, or changing remote data can make a test depend on conditions outside the test. For most detailed behavior, use a controlled substitute at that boundary and test the integration separately. Stubbing improves repeatability but removes some end-to-end confidence, so preserve another verification method for the real boundary. See Fowler’s guidance on testing strategies in a microservice architecture.
Browser timing and visual state
Browser tests can race page rendering, lazy-loaded content, animations, popups, dialogs, and network requests. Wait for a meaningful state such as a visible element or completed navigation instead of a generic delay. Disable or finish animations when they are irrelevant to the behavior under test. Handle expected dialogs explicitly and make the test report the relevant browser state when an assertion fails.
Keep end-to-end tests focused on important user journeys. Put detailed business rules and edge cases in faster lower-level tests, while retaining integration coverage at meaningful boundaries. End-to-end tests provide realistic integration confidence, but browser timing and GUI behavior carry additional maintenance and reliability costs. Fowler describes this balance in The Practical Test Pyramid and the microservice testing strategies guide.
Leaked or exhausted resources
Unclosed database connections, files, browser contexts, processes, or other managed resources can cause later tests to fail unpredictably. Check acquisition and release paths, including exceptions and early returns. Where possible, use scoped cleanup constructs or framework fixtures that guarantee teardown.
4. Choose a repair that keeps the test useful
Compare candidate fixes on diagnostic confidence, stability under the observed failure conditions, regression coverage retained, runtime, maintenance burden, and fidelity to production behavior.
| Observed pattern | Likely area | Repair direction | Coverage tradeoff |
|---|---|---|---|
| Fails only after another test | Shared state, setup, or teardown | Reset state, isolate fixtures, repair cleanup, or avoid resource collisions | Usually retains the original behavior under test |
| Fails under load or on slow runs | Timing assumption | Wait for a callback or observable condition with a useful timeout | Retains coverage if the awaited condition represents the required behavior |
| Fails when a remote service changes or is unavailable | External dependency | Control the dependency for detailed tests; verify the real boundary separately | Stubbing reduces end-to-end confidence at that boundary |
| Fails around dates or timezones | Uncontrolled clock or timezone | Inject a clock and specify the timezone and instant | Can retain boundary coverage with explicit time cases |
| Fails on a visual or browser transition | Animation, dialog, rendering, or navigation timing | Wait for the intended state and handle browser events explicitly | Retains journey coverage if the state is part of the journey |
Do not make the assertion looser merely to make the test pass. First determine whether the test expectation is invalid, the product behavior is wrong, or the test setup is nondeterministic. Preserve an assertion for the original defect whenever possible.
5. Validate the fix where it failed
- Run the repaired test alone with the conditions that previously triggered failure.
- Run it in its ordinary suite and, when relevant, under parallel execution.
- Compare logs from passing and failing runs to verify the suspected dependency is controlled.
- Confirm that the assertion still detects the regression the test was meant to catch.
- Run the relevant neighboring tests to check that fixture or cleanup changes did not shift the failure elsewhere.
Repeated passes increase confidence only to the extent that the test was exercised under the conditions that caused the failure. They do not replace checking the repair’s logic or preserving the intended assertion.
6. Quarantine only as a temporary containment step
Quarantining a flaky test can protect the ordinary suite’s signal while investigation continues. But a quarantined test no longer acts as an ordinary regression check. Keep it visible in a separate queue or later pipeline stage, assign an owner, record the reason, and set a removal deadline. Fowler gives a one-week limit as an example, not a universal rule. Do not allow quarantine to become a permanent hiding place for an unresolved failure.
If an unstable external boundary must be excluded or stubbed in one test layer, retain another means of verifying that boundary. A small set of important end-to-end journeys combined with detailed lower-level tests often gives a more useful balance than relying on a large, timing-sensitive browser suite for every rule.
7. Troubleshooting checklist
| Symptom | Common cause | What to check or change |
|---|---|---|
| Passes alone, fails in the suite | Order dependence, shared state, or faulty teardown | Run after suspected predecessors; reset fixtures; inspect static state and cleanup. |
| Passes serially, fails in parallel | Two tests claim the same resource | Use unique records, names, files, ports, and accounts; inspect shared queues. |
| Fails more often on slower workers | Fixed sleep or timing race | Wait for an observable condition with a bounded timeout and useful error details. |
| Fails at certain times or dates | Wall-clock, timezone, or boundary assumptions | Inject a clock; set an explicit instant and timezone; cover boundary cases directly. |
| Fails when a remote service is slow or unavailable | External dependency or network variability | Control the dependency for detailed tests and verify the real integration separately. |
| Browser assertion fails despite the page appearing eventually | Rendering, navigation, animation, dialog, or lazy-load race | Wait for the behavior’s actual completion signal; capture browser state and logs. |
| A different test fails after a cleanup change | Cleanup masks or shifts the original state leak | Check teardown on both success and failure paths and confirm each test starts clean. |
| Reruns pass but the defect can still recur | Rerun treated as the fix, or weak assertion | Find the uncontrolled input and confirm the assertion still detects the intended regression. |
8. Practical reliability and runtime tradeoffs
Adding retries can reduce disruption from an intermittent failure, but a retry can also conceal a broken test and delay diagnosis. Keep first-failure evidence, make retries visible, and track repeated retry success as a problem to investigate rather than a stable pass.
Long fixed waits consume suite time without making the condition more deterministic. Polling or callbacks can avoid unnecessary waiting, but every wait needs a timeout so a missing response fails clearly. Isolating fixtures has setup cost; shared mutable fixtures may be faster initially but add order and cleanup risks. Choose based on the observed failure and retain the coverage the test is intended to provide.
9. FAQ
How many reruns prove that a test is flaky?
No fixed number proves it. A pass and failure on the same revision and relevant conditions establishes an intermittent symptom; the investigation still needs to identify the changing dependency.
Should I delete a test that keeps failing?
Only if its assertion is obsolete or invalid and the coverage decision is deliberate. If it covers important behavior, diagnose and repair the cause; temporary quarantine needs an owner and deadline.
Can a flaky test reveal a real product bug?
Yes. Intermittency may arise from a real race or timing-sensitive product behavior. Establish whether the changing result is a test defect, an uncontrolled environment, or a genuine product defect before changing the assertion.
Should every user journey be an end-to-end test?
No. Keep end-to-end coverage focused on important integrated journeys and test detailed rules at lower levels, while retaining another check for any boundary that is stubbed.
Or skip the browser setup
If your flaky test or debugging workflow needs a website screenshot, ScreenshotNeo is a website screenshot API and MCP server. Its API takes a URL and returns an image or PDF; its browser capture options can help inspect visual state during an investigation. It does not diagnose or repair test nondeterminism.
For test reliability, the same principle applies: capture only after the page reaches the intended state, and keep the underlying test assertion. ScreenshotNeo can wait for a selector, a delay, or network idle, capture a full page or selected element, and apply custom headers, cookies, or JavaScript when needed. See the ScreenshotNeo API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes known cookie and consent banners, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can use its MCP server to take screenshots, get page information, or capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.


