ScreenshotNeo

BlogGuides

Common Causes of Automation Testing Failures and How to Fix Them

Learn how to diagnose failed automation runs, separate product regressions from flaky tests and CI problems, and fix the cause with evidence.

By the ScreenshotNeo team4 October 202612 min read

A failed automation run is a signal to investigate, not proof that the application regressed. The cause may be a product defect, a flaky test, uncontrolled data or state, an external dependency, or an unstable runner. Preserve the evidence, reproduce the failure under controlled conditions, classify the cause, and fix that cause instead of merely making the run green.

This guide answers Why are my automated tests failing? and Why do my tests pass locally but fail in CI? with a practical triage process, fixes for common failure classes, and guidance on keeping suites reliable.

1. Start by preserving evidence

Before rerunning a failure, save the artifacts that may disappear or change on the next run. A rerun can help reproduce a problem, but it cannot replace the original evidence.

  • Test name, run ID, retry number, timestamp, and worker or shard.
  • Application commit or build version, test code revision, browser and operating system versions.
  • Console output, application and dependency logs, network request and response details, and screenshots or traces.
  • Test data identifiers, relevant database state, feature flags, configuration, and environment variables with secrets removed.
  • Runner resource and scheduling information, including concurrent jobs where available.

Compare a failing execution with a successful one. Chromium’s guidance recommends targeted logging and comparing the success and failure paths; Google’s end-to-end testing guidance recommends retaining useful state for debugging.

2. Use a repeatable triage sequence

  1. Reproduce the failure alone. Run the exact case independently with the same application version and, as far as possible, the same environment. Record whether it fails consistently, intermittently, or not at all.
  2. Restore suite conditions. If it passes alone, run it in its original group and order, then under the original concurrency. A change in outcome points toward ordering, shared state, cleanup, or resource contention.
  3. Check the expected state and timing. Use logs and traces to determine whether the application reached the state the test assumed before the next action ran.
  4. Inspect what the user interface actually rendered. At the failure point, check whether the target is absent, obscured, disabled, renamed, or selected through a changed DOM structure.
  5. Compare local and CI conditions. Check operating system and browser versions, resource capacity, scheduling, network access, configuration, and service availability.
  6. Classify before changing code. Record whether the evidence supports a product defect, test-code defect, data/state coupling, external dependency failure, or infrastructure problem.
  7. Apply the matching fix and keep a regression check. Do not weaken an assertion if the evidence points to a genuine behavior regression.

For browser failures, a screenshot or trace can reveal whether the page was blank, still loading, showing a consent overlay, or displaying an unexpected error. Treat it as one artifact in the investigation: it cannot by itself establish whether the application, test, or runner caused the failure.

3. Match the fix to the failure

Timing and synchronization errors

What happens: The test acts before the application is ready, waits for a signal unrelated to the needed state, or assumes asynchronous work completes in a particular order. Browser and WebDriver can also race.

How to diagnose: Compare timestamps for the action, request, response, and resulting UI state. Inspect the trace or logs to see whether the expected condition occurred before the assertion or next action.

How to fix: Wait for the observable condition the test needs, then assert it with a real timeout. Use framework actionability checks where available. A short controlled delay can be a diagnostic experiment, but an arbitrary sleep is not a reliable repair: it may be too short under load and wastes time when the application is already ready. Google advises against arbitrary delays, and Playwright provides condition-based waiting and actionability behavior.

// Playwright example: wait for the user-visible result
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByText('Changes saved')).toBeVisible({ timeout: 10_000 });

Adapt the locator and expected result to the application contract. Do not replace an outcome assertion with a wait that merely proves some unrelated request finished.

Shared state, dirty data, and cleanup gaps

What happens: A test relies on a record left by another case, leaks global state, collides with another parallel worker, or assumes suite order. These problems often explain intermittent failures and failures that appear only under parallel execution.

How to diagnose: Run the case alone, in its original order, and concurrently. Compare a clean environment with a reused one. Inspect setup and teardown, shared database records, browser storage, and global configuration.

How to fix: Create prerequisites explicitly, use unique data per test or worker, restore modified global state, and make cleanup dependable. If isolation is not immediately possible, temporarily serialize the affected cases while removing the coupling. pytest, Playwright, and Chromium all document isolation and uncontrolled state as relevant sources of flakiness.

Brittle selectors and implementation-coupled assertions

What happens: A selector depends on a CSS class, DOM position, or markup structure that changes during a harmless refactor. An assertion may also check an internal detail rather than the behavior a user depends on.

How to diagnose: Inspect the DOM or trace at failure time. Check whether the control exists and whether it is visible, enabled, and reachable. Ask whether the asserted detail is part of the intended behavior contract.

How to fix: Prefer accessible roles and labels when they represent the intended user-facing contract. Where visible wording or structure changes independently of the behavior under test, use an explicit stable test contract. Assert the user outcome. No selector type is universally stable; choose one that expresses what the test is meant to protect.

// Prefer a user-facing role and label when they are the contract
await page.getByRole('button', { name: 'Continue' }).click();
await expect(page.getByRole('heading', { name: 'Review your order' })).toBeVisible();

Uncontrolled third-party services

What happens: A test depends on content, latency, banners, or availability from a system the team does not control. An external change can break a test unrelated to the feature being tested.

How to diagnose: Identify requests to external services, inspect their responses and timing, and compare the failure with a run using a controlled response.

How to fix: Stub or intercept the dependency when testing behavior owned by your application. Keep separate integration coverage for cases where the real service interaction matters, and keep test doubles aligned with the real contract so they do not hide incompatibilities. Playwright documents routing requests to controlled responses; Google’s end-to-end guidance also notes that external components can change and that test doubles can drift.

CI runner, resource, network, or machine instability

What happens: A runner lacks capacity, unrelated processes consume resources, concurrent jobs collide, or network and machine faults interrupt execution. This is a common avenue to investigate when tests pass locally but fail in CI.

How to diagnose: Check whether the application started and inspect runner, system, and network logs. Compare resource use, concurrency, OS and browser versions, and configuration with a successful run. Reproduce at comparable concurrency when possible. If failure appears only at high concurrency, investigate both resource pressure and shared-state collisions.

How to fix: Provide sufficient runner capacity, reduce unrelated load, correct scheduling collisions, and isolate resources where necessary. Keep browser and OS versions consistent when comparing visual output. Do not assume every CI-only failure is a runner problem; the runner may be exposing a real race or data dependency.

Real application or dependency regression

What happens: The system under test is slow, unresponsive, racy, resource-starved, or has changed behavior without a corresponding test update.

How to diagnose: Reproduce against the same version in a controlled environment, compare successful and failing traces, and inspect application and dependency logs. Check whether the observed behavior violates the expected product contract.

How to fix: Repair the application or dependency defect. Update the test only when behavior changed intentionally. Do not weaken the check just to restore a green build when evidence indicates a regression.

Overly broad end-to-end coverage or weak diagnostics

What happens: A browser test exercises many components and inherits their failure modes, while providing too little evidence to identify the cause. End-to-end tests are valuable for critical cross-component behavior, but they are slower and more expensive to maintain than unit or integration tests.

How to improve: Keep browser tests for important user workflows that lower-level checks cannot reliably cover. Keep cases focused, use controlled data, and preserve logs, screenshots or traces, and relevant state. Use unit or integration checks where they can verify the behavior more directly. Selenium’s test automation guidance recommends keeping tests short and using a browser only when needed; Google’s end-to-end guidance makes a similar distinction in scope and maintenance cost.

4. Why tests pass locally but fail in CI

Local and CI environments may differ in runner capacity, scheduling, OS or browser version, network access, service configuration, concurrency, or existing data. A test that depends on a fast machine, a particular run order, or an external service can pass locally by chance.

Observed pattern Investigate first
Fails only under parallel execution Shared records, global state, cleanup, test ordering, worker collisions, and resource contention.
Fails only on a particular CI runner or shard Runner logs, capacity, machine or network errors, shard data, and scheduling.
Fails after a browser or OS update Version-specific behavior, changed rendering, and assumptions in selectors or timing.
Fails intermittently while external calls are slow Dependency availability, response timing, and whether the test should use a controlled response.
Fails consistently on the same build everywhere Application behavior or a deterministic test and configuration defect.

Change one relevant condition at a time where practical. If reducing concurrency makes the failure disappear, that is evidence to investigate scheduling, resource pressure, and shared state; it is not proof that the test is fixed.

5. A decision guide: where should this behavior be tested?

Choose the smallest test scope that can reliably verify the behavior. Compare the behavior’s scope, control over dependencies and data, reproducibility, execution cost, and the diagnostic evidence available when it fails.

Test level Best fit Trade-off to consider
Unit Logic that can be checked in isolation. Does not establish that separate components work together.
Integration Interactions between components or with a controlled dependency. Requires suitable setup and may not prove the complete user workflow.
End-to-end/browser Critical, user-visible workflows and cross-component behavior not reliably covered below. Slower, more exposed to environment and dependency changes, and more costly to maintain.

Avoid moving every failure to a browser test for extra reassurance. Keep browser coverage for the behavior that needs it, and make its dependencies and data as controlled as possible.

6. Retries: useful evidence, poor diagnosis

A retry can show that a failure is intermittent, and a limited retry policy can temporarily reduce disruption while an investigation is underway. A passing retry does not identify the cause or prove the application is healthy. Record the first-attempt failure and retry outcome, track recurring cases, and avoid permanent quarantine without an owner and a plan to resolve the underlying problem. pytest documents reruns as a mitigation while warning that quarantine can be dangerous.

7. Browser screenshots as debugging evidence

A screenshot can help answer what the browser displayed at the point a UI check failed. It can expose an unexpected overlay, a blank page, a missing element, or an error state. Pair it with the test trace, console and network logs, timestamps, and application version; a screenshot alone cannot show why the page reached that state.

For a local Playwright capture, this runnable Node.js example opens a page, waits for a meaningful condition, and saves a full-page PNG:

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();

try {
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  await page.getByRole('heading', { name: 'Example Domain' }).waitFor({ state: 'visible', timeout: 10_000 });
  await page.screenshot({ path: 'failure-context.png', fullPage: true });
} finally {
  await browser.close();
}

Install Playwright and its browser using the commands in the official Playwright getting-started guide. Replace the example URL and heading with a page and condition relevant to your investigation. For a failure artifact, capture the application state and URL actually used in the test, and retain the trace or logs too.

8. Or skip the browser setup

If you need a page capture while investigating a UI or dependency failure, ScreenshotNeo is a website screenshot API and MCP server for developers. Send one GET request with the target URL and receive an image or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and whether the request was billed.
  • An MCP server lets AI agents, including Claude and Cursor, take screenshots with tools for screenshots, page information, and PDF capture.
  • The Free plan includes 1,000 shots a month with no card. Paid plans start at $5 for 3,000 shots. Every feature is available on every plan.

Sign up free for 1,000 screenshots a month, with no card.

9. Reliability, performance, and cost

  • Reliability: Control data and dependencies, wait for application conditions, isolate cases, and preserve enough evidence to diagnose failures. Retries can help expose intermittency but do not replace a root-cause fix.
  • Performance: Arbitrary waits add runtime without guaranteeing readiness. Keep browser tests focused and use lighter test levels for behavior they can verify reliably.
  • Maintenance cost: Broad end-to-end tests tend to involve more components and require more upkeep. Reserve them for important workflows and keep external dependencies controlled or separately covered.
  • Capture cost: If using ScreenshotNeo for debugging captures, only clean shots are billed. Its response includes X-Page-Verdict and X-Billed headers so the outcome and billing status are visible.

10. Troubleshooting checklist

Symptom Likely cause First fix to try
Passes alone, fails in suite Ordering dependency, shared state, or incomplete cleanup. Compare order and state; initialize independent data and restore state.
Fails only in parallel Data collision, shared global state, scheduling, or resource contention. Use per-worker data and isolate shared resources; serialize temporarily while investigating.
Intermittent element timeout Wrong readiness condition, slow application state, overlay, or brittle locator. Inspect trace and rendered page; wait for the intended user-visible condition.
Local passes, CI fails Environment mismatch, runner capacity, network, concurrency, or hidden state. Compare versions and logs, reproduce CI concurrency, and inspect runner resources.
Failure follows an external request Dependency latency, outage, or changed response. Inspect request and response; stub owned behavior and retain integration coverage where needed.
Screenshot shows a blank page Navigation failure, timeout, bot check, or capture before content rendered. Check navigation result, network and console logs, and the condition used before capture.
Retry passes Intermittency; cause still unknown. Keep the first failure artifacts and compare both executions before relying on retries.

11. Frequently asked questions

Why does a test fail only once in a while?

Intermittent failures often involve uncontrolled timing, shared state, concurrency, external dependencies, or runner conditions. Compare the failed run with a successful one and vary one relevant condition at a time.

Should I delete a flaky test?

First determine whether it protects an important behavior and identify why it flakes. Repair or replace an unhelpful test deliberately; deleting a test without checking its coverage can remove a useful regression signal.

Does a green retry mean the bug is gone?

No. It shows that a later attempt passed. Keep the first-attempt result visible and investigate why the outcome changed.

Can a screenshot prove the test failure was a product bug?

No. It records visible browser state. Logs, traces, environment details, and reproducibility help establish whether the cause was product behavior, test code, a dependency, or infrastructure.

Sources