Are Automated UI Tests Unstable? Common Causes and Fixes
UI tests can pass once and fail the next time. Learn how to find the cause, fix the test, and keep retries from hiding instability.
Yes. Automated UI tests can be unstable: the same test may pass in one run and fail in another even when the relevant application code has not changed. A retry that passes is evidence of flakiness, not proof that the test or application is healthy. First compare the failed and passing attempts, then fix the cause: commonly timing, shared state, brittle assertions, external dependencies, or CI environment differences.
1. What a flaky UI test is
A flaky test produces different results under apparently equivalent conditions. It can fail because the application is broken, because the test makes an assumption that is not always true, or because the environment changes the timing or inputs. A single failure does not identify which explanation is correct.
Playwright calls a test flaky when it fails initially and passes on retry. Retries are disabled by default in Playwright; Cypress also supports retries and reports retry outcomes. A retry pass should remain visible in reports so recurring instability is investigated. Playwright retries · Cypress test retries.
2. Common causes and the matching fix
| Cause | Typical symptom | Fix |
|---|---|---|
| Timing and asynchronous updates | Click or assertion runs before a request, animation, or render completes. | Wait for the specific element or application state, then assert the user-visible outcome. |
| Shared state or test data | Passes alone, fails in a suite or parallel run. | Give each test independent setup and unique records; avoid order dependencies. |
| Brittle checks | Fails after markup refactoring although the user behavior still works. | Assert meaningful rendered behavior instead of incidental implementation details. |
| External dependencies | Intermittent failures around a third-party API, network, or service. | Control the dependency response when testing your application, and use a stable environment. |
| CI variation | Stable locally, inconsistent on a runner. | Compare browser, OS, resource pressure, service availability, and data collisions before changing timeouts. |
Timing: wait for behavior, not a guessed duration
Fixed sleeps can be too short on a slow run and waste time on a fast one. Prefer a condition tied to the behavior being tested. Playwright checks that supported action targets resolve to one element and are visible, stable, unobscured, and enabled before interaction. Its assertions retry while waiting for the requested condition. Selenium likewise documents condition-based waits and warns that mixing implicit and explicit waits can produce unpredictable timeout durations.
// Playwright Test: wait for the outcome, not an arbitrary delay
import { test, expect } from '@playwright/test';
test('shows the saved state', async ({ page }) => {
await page.goto('http://localhost:3000/settings');
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByRole('status')).toHaveText('Settings saved');
});
Use accessible roles and names where they express the control a user interacts with. If a specific request drives the behavior, wait for that request or its resulting visible state. Avoid waiting for every network connection to become idle when the application legitimately keeps background connections open.
References: Playwright actionability, Playwright assertions, and Selenium waiting strategies.
Shared state: isolate setup, data, and cleanup
Separate browser contexts isolate cookies and browser storage, but they do not isolate a shared database record, file, or external account. Create unique test data, clean up deliberately, and do not make one test depend on another test’s side effects. If a shared resource cannot be isolated, limit concurrency explicitly and make that constraint clear.
import { test, expect } from '@playwright/test';
// A unique value prevents parallel tests from editing the same account.
test('updates a profile', async ({ page }) => {
const testId = `${Date.now()}-${Math.random().toString(16).slice(2)}`;
await page.goto(`/profile/new?testId=${testId}`);
await page.getByLabel('Display name').fill(`Test user ${testId}`);
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByRole('status')).toHaveText('Profile saved');
});
The setup route above is application-specific: implement it only if your test application supports such a route. Otherwise create the record through your test fixture or API and pass its unique identifier into the page. Playwright’s guidance emphasizes independent tests, and its parallelism documentation notes that state outside an isolated browser context can still cause flakiness. Best practices · Parallelism.
Brittle checks: assert what the user needs
A selector tied to a generated class or a deep DOM structure can break during a harmless refactor. Prefer a role, label, or other stable contract when possible. Assert the outcome of the interaction rather than private component state. Retrying assertions help with asynchronous rendering; they do not make an incorrect expectation correct.
3. Diagnose a failure in a repeatable sequence
- Preserve the first failure. Record the failing step, error, browser, and whether a retry passed. Do not let a green retry erase the initial failure.
- Compare one failed and one passing attempt. Check the DOM, relevant element state, network requests, console output, and timing around the action.
- Run it alone, then in context. If it only fails in a suite or parallel workers, investigate ordering, shared records, and cleanup.
- Compare local and CI conditions. Check service readiness, runner resource pressure, browser and OS versions, network access, and whether test data collides.
- Change the cause, then repeat under the failing conditions. Increasing a global timeout without identifying the slow or missing condition can conceal a race and slow every run.
Cypress Cloud’s replay context can help inspect DOM state, requests, console logs, and element state near a failure. The key diagnostic question is: “Are these failures real regressions, or known flakiness?” See Cypress Cloud flaky-test management.
4. Retries: useful diagnostic signal, poor permanent fix
Retries can reveal that a test is unstable and may absorb an occasional transient interruption. They also rerun tests and hooks, adding execution time. Keep retry counts limited, retain diagnostics from the first attempt, and track retry-passing tests as work to investigate rather than treating the final green status as a fix. Cypress recommends low retry counts and using flake data to find root causes. Cypress test performance.
5. What the published evidence does and does not show
A 2025 IEEE ICST empirical study examined 49 web projects and 123 DOM-event-related test cases. Within that study’s dataset and scope, the observed repair strategies were 50.4% DOM interaction synchronization, 38.2% conditional waits for event completion, and 11.4% ensuring consistent DOM state transitions. These are shares of strategies observed in that study, not the overall prevalence of all UI test flakiness or a guarantee that synchronization is the cause in a particular test. Read the IEEE ICST study.
6. Screenshot checks and test evidence
Visual comparisons can help locate changes in rendered output, but they need consistent browser, viewport, data, and page state. Wait for the page to reach the intended state before capture, and avoid comparing pages while animations, live content, or external requests are changing. A screenshot can show what appeared; it does not by itself explain whether the cause was a rendering defect, a late request, or unstable test data.
Or skip the browser setup
If you need a clean page capture while documenting or investigating a UI failure, ScreenshotNeo is a website screenshot API and MCP server. Make one GET request with a URL to return PNG, JPEG, WebP, or PDF. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. See the docs for parameters and configuration.
Sign up free for 1,000 screenshots a month, with no card required.
7. Reliability, performance, and cost
- Reliability: deterministic setup, isolated data, condition-based waits, and retained failure diagnostics make results easier to reproduce. A retry policy cannot compensate for uncontrolled inputs.
- Performance: arbitrary sleeps and repeated retries increase run time. Wait on the narrow condition that matters and avoid broad timeout increases that affect every test.
- Cost: retries consume CI time, and unnecessary waits add to it. For hosted capture tools, check what outcomes are billed and how cache behavior is reported; ScreenshotNeo states that only clean shots are billed.
- Visual stability: keep viewport, browser, data, and capture state consistent. Control third-party content when it is outside the behavior under test.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Element not found intermittently | Render or request has not completed, or selector is brittle. | Wait for the user-visible condition and use a stable role or label. |
| Click intercepted or target unstable | Overlay, animation, moving layout, or another element covers it. | Check the failure snapshot and wait for the intended state; do not force the click until the obstruction is understood. |
| Passes alone, fails in suite | Shared backend state, order dependency, or cleanup collision. | Use unique records, independent setup, and explicit cleanup; then run with parallel workers. |
| Passes locally, fails in CI | Service, resource, browser, OS, or network difference. | Compare the first failing attempt and CI conditions before tuning timeouts. |
| Retry passes repeatedly | Underlying nondeterminism remains. | Keep the flake visible, inspect attempt differences, fix the root cause, and rerun under the same conditions. |
| Timeout becomes longer after adding waits | Implicit and explicit Selenium waits are mixed, or a broad condition never becomes true. | Use one deliberate wait strategy, verify the expected condition, and avoid waiting for unrelated network activity. |
9. FAQ
Does a flaky test mean the application is defective?
No. It means the result is unreliable until you distinguish an application regression from test or environment instability.
Should I disable a test that flakes?
Do not silently discard its signal. If it blocks useful feedback, quarantine it visibly with an owner and investigation plan while preserving its failure data.
Will increasing the timeout fix flakiness?
Only when a legitimate operation needs more time and the test waits for the right condition. A larger timeout does not fix shared data, external failures, or a wrong assertion.
Can screenshots prove a test is stable?
No. Screenshots provide visual evidence for a particular attempt. Stability requires repeatable setup and consistent outcomes across relevant runs.


