How to Find and Clean Up Dirty Automated Tests
Find tests that pass alone but fail in the suite, trace the cause, and make setup, cleanup, and retries reliable without hiding product defects.
A “dirty automated test” is an informal term for a test that depends on uncontrolled prior state, leaves state behind for later tests, or has unreliable setup or cleanup. Find candidates in CI history, reproduce them under the same conditions, then isolate the cause across test code, runner, application dependencies, and environment. Repair the root cause by making state explicit, synchronizing on real conditions, and guaranteeing cleanup. Retries and quarantine can help diagnose or contain a failure, but neither makes a test reliable.
1. Find candidate tests from evidence
Start with test history, not a hunch. Look for tests that alternate between pass and fail without an explanatory code change, fail repeatedly in the same area, or change behavior when suite order or parallel execution changes. A pass/fail change for the same code is evidence of flakiness; by itself it does not identify whether the cause is the test, product, framework, dependencies, or infrastructure.
For every candidate, preserve enough context to make comparisons meaningful:
- Test name or stable test identifier and the exact failure output.
- Commit, branch, and relevant code changes.
- Runner image, operating system, runtime, framework version, and configuration.
- Whether it ran alone or in a suite, its order, and whether execution was parallel.
- Logs, timestamps, resource use, and relevant environment setup.
pytest describes a flaky test as one with intermittent or sporadic failure, and notes that uncontrolled system state and cleanup leakage can affect other tests. See the pytest guide to flaky tests.
2. Reproduce and narrow the failure
- Rerun the suspect test by itself with the same commit and environment. Repeat enough times to establish whether the failure recurs; record every result.
- Run the relevant suite under the same conditions. Compare isolated and suite behavior.
- If evidence points to ordering, vary order or run the test after likely predecessors. If it points to concurrency, compare serial and parallel execution.
- Change one variable at a time. Preserve logs and environment details as you narrow the scope.
A test that passes alone but fails in the suite suggests interaction, ordering, or shared state. It is a useful lead, not proof of a particular cause. Google’s triage guidance recommends independently rerunning tests when investigating test-state assumptions.
Inspect each layer that participates in the run:
- Test code and data: shared fixtures, reused records, mutable globals, fixed identifiers, and assumptions about existing files or database rows.
- Framework and runner: fixture scope, teardown behavior, retries, test ordering, parallel workers, and runner configuration.
- Application and dependencies: asynchronous work, external services, queues, caches, eventual consistency, and dependency versions.
- Host environment: clocks, time zones, locale, network availability, CPU and memory pressure, ports, temporary storage, and shared services.
3. Check the usual sources of test dirt
Uncontrolled or shared state
Find every state change a test makes: database writes, files, environment variables, global configuration, singleton caches, browser storage, ports, and external resources. Ask whether the test establishes the state it needs or quietly relies on a previous test or run. Give each test unique data where possible, isolate resources, and restore shared state in teardown.
Setup and teardown gaps
Cleanup must run after a passing assertion, a failing assertion, and an unexpected exception. Prefer the framework’s fixture, teardown, or scope-guard mechanism so cleanup is tied to the test lifecycle. In GoogleTest, each fixture gets a fresh object, then SetUp(), the test body, and TearDown() run in sequence. A fatal assertion returns from the current function, so cleanup statements later in that function may be skipped. Put essential cleanup in TearDown() or another guaranteed cleanup mechanism.
class DatabaseTest : public ::testing::Test {
protected:
void SetUp() override {
db_ = StartTemporaryDatabase();
SeedRequiredRows(db_);
}
void TearDown() override {
if (db_) {
StopTemporaryDatabase(db_);
db_ = nullptr;
}
}
Database* db_ = nullptr;
};
This is a lifecycle example; replace the placeholder database helpers with the project’s actual setup and cleanup functions. See the GoogleTest Primer.
Timing and asynchronous work
Do not assume that a background task, network response, animation, or database update finishes within a fixed delay. Wait for a meaningful condition, such as a state transition or a bounded completion signal, and set a timeout that gives useful failure diagnostics. Google’s Testing Blog cautions: “Do NOT add arbitrary delays as these can become flaky again over time and slow down the test unnecessarily.” See its guidance on test flakiness and triage.
Order and parallel execution
If changing the order changes the result, locate the state left by a predecessor or assumed by the suspect. If parallel execution triggers it, identify shared files, ports, database records, global state, or external rate limits. Remove cross-test dependencies or isolate the contested resource. A test that mutates global state may not be safe to run concurrently.
External dependencies and environment
Where practical, replace unstable external dependencies with controlled test doubles, local fixtures, or isolated test instances. When the external behavior itself is under test, make the dependency and its failure modes explicit. Capture environment data alongside the failure so infrastructure variation is not misdiagnosed as a code defect.
4. Make the repair preserve the behavior under test
Make the test establish its own required state, use unique or isolated resources, synchronize on meaningful conditions, and guarantee teardown even on early failure. Then rerun it alone and in the suite, including the order or concurrency condition that originally exposed the problem.
If a broad end-to-end test is fragile and hard to diagnose, consider covering the core behavior with a smaller test at a lower level, while retaining end-to-end coverage for behavior that genuinely depends on system integration. Do not simply delete a meaningful test: first ensure equivalent behavior remains covered. Google’s test strategy guidance discusses the faster, more isolated feedback of smaller tests; pytest also recommends rewriting or deleting a test only when its functionality remains covered.
5. Use retries and quarantine as temporary controls
A retry can expose an intermittent outcome and some CI systems support retrying failed tests. But a retry that passes does not explain the cause, and repeated retries can delay discovery of a genuine regression. A retry policy should retain the first failure and make retry results visible.
Quarantine can keep a known unreliable test from blocking every change while work proceeds, but it can also mask a real race or product defect. Keep quarantined tests visible, link each one to an issue, assign an owner, and review it until the root cause is fixed and the test returns to normal gating. GitLab documents an issue-backed approach to unhealthy tests; it is GitLab practice, not a universal standard. Google’s historical discussion of quarantine likewise warns about hiding real races or bugs.
| Choice | Restores determinism? | Useful when | Main risk |
|---|---|---|---|
| Fix state, setup, timing, or isolation | Usually, if the root cause is found | The failure can be reproduced or traced to a condition | Misdiagnosing the cause can leave the issue unresolved |
| Retry | No; it repeats the execution | Collecting evidence or containing intermittent failures briefly | Hides or delays a real regression |
| Quarantine | No; it changes gating | A tracked repair needs time while failures remain visible | Unowned tests can stay excluded and defects can pass unnoticed |
| Refactor to a smaller test | Can improve isolation | A broad test is slow or difficult to localize | Behavior coverage can be lost if replacement coverage is not verified |
There is no universal acceptable flake rate, retry count, or quarantine duration established by the cited sources. Google reported historical figures for its own corpus in 2016; those figures are not current industry benchmarks and should not be used as a target for another team.
6. Troubleshooting common symptoms
| Symptom | Likely causes to investigate | Next action |
|---|---|---|
| Passes alone, fails in the suite | Leaked state, order dependency, shared fixture or resource | Reproduce in the suite; inspect predecessors and teardown; isolate data and resources. |
| Fails only in parallel | Concurrent mutation of globals, files, ports, records, or services | Identify the shared resource; use unique resources or serialize only the constrained operation while fixing isolation. |
| Fails after adding a sleep, then flakes again | The delay guesses at completion rather than observing it | Wait for an explicit condition with a bounded timeout and diagnostic output. |
| Cleanup does not run after assertion failure | Cleanup appears after a fatal assertion or exception path | Move cleanup into fixture teardown, a scope guard, or a guaranteed finalization construct. |
| Retry passes but original failure remains unexplained | Retry masked an intermittent condition | Retain the initial failure logs, compare environment and order, and continue root-cause analysis. |
| Failure occurs only in CI | Runner configuration, resource pressure, clock or locale differences, network, or dependency variation | Capture runner and environment metadata; reproduce in a matching environment before changing test logic. |
| Quarantined test is forgotten | No owner, issue, or review point | Attach an issue and owner, keep failure history visible, and review until the test is repaired or replaced. |
7. A repeatable cleanup checklist
- Capture the test identifier, commit, runner, environment, order, logs, and first failure.
- Compare repeated isolated runs with suite runs under matching conditions.
- Inspect data, shared state, lifecycle hooks, asynchronous work, external dependencies, and resource contention.
- Make required state explicit and cleanup reliable on all exit paths.
- Use condition-based synchronization and useful timeouts instead of arbitrary sleeps.
- Verify the fix under the original failure condition and confirm the suite still covers the intended behavior.
- If retrying or quarantining temporarily, preserve visibility, ownership, and a path back to normal gating.
Or skip the browser setup
For a browser-based end-to-end test, screenshots can preserve visual evidence of the rendered page at the failure point. You can capture with your own browser tooling, or request a screenshot directly with ScreenshotNeo, a website screenshot API and MCP server for developers. See the ScreenshotNeo API documentation for the available parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The API also accepts parameters used by other screenshot APIs, which can make switching easier. ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan. Sign up for 1,000 free screenshots a month with no card.
FAQ
Is “dirty automated test” a formal testing term?
No. This article uses it informally for tests that rely on uncontrolled prior state, contaminate later runs, or have unreliable setup and cleanup.
Does a flaky test mean the product is broken?
Not by itself. The variation can originate in the test, runner, application or dependencies, or the host environment. Investigation must distinguish those causes.
Should I delete a test that keeps flaking?
Only after deciding that its behavior is covered elsewhere or replacing it with coverage that preserves the intended check. Rewriting at a smaller scope can make failures easier to localize.
Can screenshots prove why a UI test failed?
A screenshot records visible page output at capture time. It can aid diagnosis, but it does not replace logs, state inspection, or reproducing the failure condition.


