Why Web Test Automation Fails and How to Fix It
Fix flaky web tests by waiting for meaningful states, isolating browser and data state, checking user-visible behavior, and diagnosing CI failures.
Web test automation usually fails because the test and application get out of sync, tests share browser or backend state, checks depend on fragile page internals, or CI behaves differently from a developer’s machine. Fix the cause: wait for the state the next step needs, assert user-visible behavior, isolate each test’s state and data, and inspect the first failure’s evidence before changing timeouts or retry counts.
A page reaching a browser load milestone does not prove that a modern application is ready for the next interaction. Nor does a passing rerun prove that an earlier failure was harmless. Treat failures as evidence to diagnose, not noise to hide.
1. Identify what kind of failure you have
Start by classifying the failure. The fix depends on whether the test observed the wrong state, selected the wrong element, inherited state from another test, encountered an application defect, or ran in a constrained or different environment.
| Failure signature | Likely cause | First useful check |
|---|---|---|
| Intermittent missing element or failed click | The action ran before the application reached the needed state | Inspect the failing step and wait for its specific condition |
| Failure changes when tests are reordered or run alone | Browser state or shared test data leaks between tests | Run the test alone and inspect setup, cleanup, and shared records |
| Element found, but assertion fails after a UI change | The selector or assertion relies on implementation details or stale expectations | Check the rendered behavior a user would see |
| Passes locally, fails in CI or another browser | Browser, configuration, service, network, or resource difference | Compare versions and environment; inspect trace, screenshots, logs, and runner load |
| Passes on retry | Nondeterminism remains, or the first failure was transient | Keep the first failure’s evidence and look for a repeatable signature |
2. Fix synchronization races
Browser automation and a changing application can race: sometimes the application reaches the required state first, and sometimes the next automation command runs first. Selenium describes this as a common challenge in browser automation. Navigation completing is not necessarily the state your next action needs.
Replace arbitrary sleeps with waits for the condition tied to the next interaction or assertion. For example, wait for a submit button to become enabled, a result row to appear, or a status message to reach its expected value. The right condition is application-specific.
Selenium: wait for a meaningful condition
This Python example uses Selenium’s explicit wait for a result element. Install Selenium and provide a compatible browser driver in your environment. The selectors and expected text should match your application.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
# Keep a fresh WebDriver instance per test.
driver = webdriver.Chrome()
try:
driver.get("http://localhost:3000/search")
driver.find_element(By.LABEL, "Search").send_keys("invoice")
driver.find_element(By.ROLE, "button", name="Search").click()
result = WebDriverWait(driver, 10).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "[data-testid='search-result']"))
)
assert "invoice" in result.text.lower()
finally:
driver.quit()
Note: Selenium Python does not provide the shown By.LABEL or By.ROLE APIs as written in current Selenium bindings. For a runnable locator, use a CSS selector or XPath that matches your page, such as By.CSS_SELECTOR, "label[for='search'] + input" and By.CSS_SELECTOR, "button[type='submit']". A corrected version of the interaction is:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
driver = webdriver.Chrome()
try:
driver.get("http://localhost:3000/search")
driver.find_element(By.CSS_SELECTOR, "label[for='search'] + input").send_keys("invoice")
driver.find_element(By.CSS_SELECTOR, "button[type='submit']").click()
result = WebDriverWait(driver, 10).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "[data-testid='search-result']"))
)
assert "invoice" in result.text.lower()
finally:
driver.quit()
Use the actual accessible or stable selectors supported by your page and Selenium version. Explicit waits can time out when the condition never occurs; that should prompt investigation of the application response and selector, rather than automatically increasing the limit.
Playwright: use actionability and retrying assertions
Playwright waits for documented actionability conditions before actions and retries its web-first assertions until they pass or time out. This reduces timing mistakes when used with an assertion about the state you expect.
import { test, expect } from '@playwright/test';
test('search shows matching results', async ({ page }) => {
await page.goto('http://localhost:3000/search');
await page.getByRole('textbox', { name: 'Search' }).fill('invoice');
await page.getByRole('button', { name: 'Search' }).click();
await expect(page.getByTestId('search-result')).toContainText('invoice');
});
Prefer a role, accessible name, label, or visible text when it accurately describes how the user finds the control. A team-owned test identifier can be appropriate when user-facing text is ambiguous or changes frequently. No selector style is a universal cure: pair a suitable locator with an assertion that checks the expected state.
Cypress: assert on the expected state
Cypress commands retry queries and assertions according to their documented behavior. Avoid fixed waits as a substitute for checking the state that matters.
describe('search', () => {
it('shows matching results', () => {
cy.visit('http://localhost:3000/search');
cy.findByRole('textbox', { name: 'Search' }).type('invoice');
cy.findByRole('button', { name: 'Search' }).click();
cy.get('[data-testid="search-result"]')
.should('be.visible')
.and('contain.text', 'invoice');
});
});
findByRole is supplied by the Cypress Testing Library integration; if your project does not use it, use a supported Cypress query for your page. Cypress’s guidance recommends avoiding arbitrary waits in favor of observing application state.
3. Make selectors and assertions reflect user behavior
A test can become brittle when it depends on incidental markup, such as a styling class that changes during a redesign. Prefer to verify rendered behavior that matters to users: the control can be found by its role or label, the expected message appears, or an action changes the visible state.
- Choose a locator that communicates what the user interacts with when that is stable and unambiguous.
- Use a deliberate test identifier when the interface has no stable user-facing locator for the behavior.
- Assert the state after the action, not merely that a selector matched something.
- Keep assertions specific enough to catch regressions without coupling them to unrelated layout or styling details.
A locator that finds the right element may still be evaluated before the expected state is ready. Use your framework’s condition-based waits or retrying assertions for the outcome as well.
4. Isolate browser state and test data
Tests that depend on earlier tests can fail when run alone, reordered, or in parallel. Give each test a clean browser state and make backend records independent where your application architecture requires it.
- Selenium: create a new WebDriver instance per test, as Selenium recommends for isolation and parallelization.
- Playwright: its test model provides a fresh browser context per test by default, including isolated browser storage.
- Cypress: test isolation behavior resets browser context and test state according to its configuration and documented behavior; check that configuration if tests share state.
- Backend data: a fresh browser context does not isolate accounts, database rows, queues, or shared services. Set up unique or resettable test data appropriate to the application.
Run a suspect test by itself and in the full suite. If it only fails in one setting, inspect both browser storage and the backend data or services it touches. Avoid cleanup that silently depends on another test having run first.
5. Diagnose CI and browser differences
A local-versus-CI difference is useful evidence. Compare the same test while changing one axis at a time: browser and version, operating system, environment configuration, service dependencies, and runner load. Preserve screenshots, traces or replay, console output, request evidence where available, and environment details from the first failure.
- Reproduce the failure with the same test and browser version if practical.
- Reduce it to the smallest test that still reproduces the behavior.
- Compare local and CI configuration, browser versions, and required service readiness.
- Inspect screenshots, trace or replay, console and network evidence, and test logs around the failing step.
- If failures worsen during a run or browsers crash, inspect CPU and memory contention. CI shares resources among the test runner, browser, application server, and background services.
- If browser-to-runner connections reset unexpectedly, check network and operating-system layers, including proxy or security software where relevant.
Resource symptoms such as increasing run duration or browser crashes can point to an overloaded CI machine. Check your provider’s current guidance before changing runner size; requirements vary by provider, browser, application, and test workload.
6. Use retries carefully
A retry can keep a pipeline moving through occasional nondeterminism, but it does not repair the underlying cause. Keep retry counts low, preserve the first failure’s diagnostics, and track whether failures share a browser, test, environment, or resource signature.
A 2023 case study of Chromium CI examined methods for predicting flaky failures. In that study, prediction methods with 99.2% precision were associated with approximately 76.2% of regression faults being missed when failures were classified as flaky. Those figures describe the studied Chromium CI setting and should not be treated as a rate for all test suites. They are a reason to avoid automatically discarding a failure because it passed on rerun.
7. Troubleshooting checklist
| Symptom | Cause to investigate | Fix or next step |
|---|---|---|
| Element is missing intermittently | Test acts before rendering or async work finishes | Wait for the element or state needed by the next step; verify the application actually produces it |
| Click times out or is intercepted | Element is not actionable, covered, disabled, or still moving | Check the rendered page and wait for the intended control to be actionable; fix overlays or app behavior if needed |
| Fixed sleep works only sometimes | Work duration varies; delay is unrelated to readiness | Replace it with a condition tied to the expected state |
| Assertion fails only after another test | Browser state or backend data is shared | Run alone and in suite; create isolated contexts and independent data |
| Test breaks after a CSS or layout change | Selector depends on styling or internal markup | Use a user-facing locator where suitable or a deliberate stable test identifier |
| CI is slower and flaky, then crashes | Runner resource pressure or service contention | Inspect resource use and services; compare runs before changing timeouts |
| Browser-to-runner connection resets | Network, proxy, or operating-system layer issue | Inspect runner logs and network/security configuration for the specific environment |
| Retry passes, original run fails | Nondeterminism remains or failure has been masked | Keep first-run evidence and diagnose the recurring signature; do not label it fixed based only on rerun |
8. Capture failure evidence without adding another flaky step
When a visual state matters, a screenshot can help establish what the browser actually rendered at the failure point. Capture it on failure where the test runner supports that workflow, and retain it alongside logs and traces. A screenshot is one piece of evidence: it does not show all network, backend, or timing causes by itself.
For manual diagnosis, a browser screenshot can record a page state. For automated capture from a URL, ScreenshotNeo is a website screenshot API and MCP server for developers. It can capture PNG, JPEG, WebP, or PDF; its cookie/banner cleanup and verdict headers can help produce cleaner page captures and distinguish clean captures from failures. See the ScreenshotNeo product page and API documentation.
9. Or skip the browser setup
For a screenshot of a page as a separate diagnostic artifact, ScreenshotNeo takes a URL in one GET request. This does not replace the browser test’s assertions or trace; it can provide a page capture without setting up a browser automation runner.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners are accepted and removed, along with known newsletter popups and chat widgets, before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
10. Performance, reliability, and cost
Prefer waiting for the exact state over sleeping for a large fixed interval: arbitrary delays waste time when the app is fast and still fail when it is slower. Isolated tests can run in parallel more safely, but only if backend data and services are independent and the runner has enough resources. A retry adds execution time and can obscure first-run evidence, so keep it low and retain diagnostics.
For hosted screenshot capture, ScreenshotNeo’s stated billing rule is that only clean shots are billed; bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Its plans range from 1,000 free shots per month to paid tiers, and every feature is available on every plan. Use the usage API and response billing headers to inspect consumption and the result classification. See current options in the ScreenshotNeo documentation.
11. A short workflow for the next failure
- Save the first failure’s screenshot, trace or replay, logs, browser version, and environment details.
- Classify it: synchronization, locator/assertion, app behavior, shared state, or environment/resource issue.
- Run it alone and in the suite; check browser state and backend data separately.
- Compare local and CI or browser variants, changing one factor at a time.
- Make one targeted change and see whether the failure signature changes.
- Keep retries low and treat a retry pass as evidence of nondeterminism, not proof of repair.
Frequently asked questions
Why do tests pass locally but fail in CI?
CI may use a different browser or configuration, depend on services with different readiness, or have less available CPU and memory. Compare the same test and inspect its first failure evidence before increasing timeouts.
Should I use CSS selectors or test IDs?
Use a user-facing locator when it clearly represents how someone finds the control. Use a stable team-owned test identifier when the interface does not provide an unambiguous durable locator. Choose based on the behavior and change patterns of your application.
Does a passing retry mean the test is fixed?
No. It shows that the result was nondeterministic or that the original failure did not recur on that attempt. Keep the initial diagnostics and investigate the cause.
Which framework is best for preventing flaky tests?
The reviewed documentation does not establish a universal winner. Compare synchronization, assertion behavior, browser coverage, isolation, debugging evidence, language fit, and your application’s architecture.
Sources
- Selenium: Waiting Strategies
- Selenium: Avoid Sharing State
- Playwright: Best Practices
- Playwright: Writing Tests
- Cypress: Best Practices
- Cypress: Test Performance
- Cypress: Troubleshooting
- Guillaume Haben, Sarra Habchi, Mike Papadakis, Maxime Cordy, and Yves Le Traon, 2023 study of flaky tests and fault-triggering failures in Chromium CI.


