ScreenshotNeo

BlogHow-to

How to Debug Flaky Visual Regression Tests

Trace intermittent screenshot failures to unstable data, timing, resources, capture settings, or rendering environments—and fix the cause without hiding real regressions.

By the ScreenshotNeo team4 October 202610 min read

A flaky visual regression test produces different screenshots across repeated runs even though the code has not changed. To debug one, preserve the failing and passing captures, compare their diffs and capture conditions, inspect trace, network, console, DOM, viewport, and clip details, then stabilize the input or rendering condition that changed. A retry that passes is evidence of variability, not evidence that the original failure was harmless.

A screenshot mismatch can point to a real product regression, a defect in the test or capture setup, or noise from data, timing, resources, or the rendering environment. Diagnose those possibilities before changing a baseline, adding a broad mask, or increasing a wait.

1. Confirm whether the failure is flaky

Run the same test against the same commit and fixture more than once. Keep the existing baseline unchanged while collecting results.

  • Output varies between runs: investigate nondeterministic inputs, unsettled UI, resource loading, and environment differences.
  • The same wrong screenshot appears every run: suspect a real UI or fixture defect, a consistently incorrect baseline, or a stable capture configuration problem. It is related to flakiness but is not itself intermittent.
  • Only one browser project or CI job fails: compare that project’s browser version, OS image, viewport, headless setting, and configuration with the baseline environment.

Chromatic describes an unstable test as one that renders differently across repeated runs without a code change. That is a useful operational definition; it does not explain the cause by itself. Chromatic’s unstable test guide covers common sources and mitigations.

2. Preserve the failure context

For each failing run, save the actual screenshot, expected screenshot, diff, test output, commit or build identifier, browser project, viewport, and any trace. Keep at least one passing capture from the same code and setup for comparison. Avoid updating the expected image until you understand why the actual output changed.

Write down what the test was supposed to capture: the route or story, relevant data state, viewport, scroll position, and whether the assertion covers the full page or a particular element. Without that context, a pixel diff alone can be hard to interpret.

3. Check the rendering environment first

Make the comparison environment match the one used to create the baseline. Check the browser and version, operating system or container image, headless mode, viewport and device scale factor, browser settings, and any project-specific configuration. Rendering can vary across host OS, browser version, settings, hardware, power source, and headless mode; Playwright recommends using the same environment that generated the baseline. See Playwright’s visual comparison guidance.

If a mismatch happens only in CI or one browser, first reproduce it in that same browser project and environment. Do not regenerate the baseline on a different machine just to make the diff disappear.

4. Read the diff alongside the page evidence

Look at the changed pixels, then use the trace and browser artifacts to establish what the page was doing at capture time. Chromatic’s trace viewer can expose network activity, console messages, DOM snapshots, and capture metadata. Its trace viewer documentation describes these artifacts.

  • Network: Did the stylesheet, script, image, or font request fail, return different content, or finish after capture?
  • Console: Did an exception prevent the intended state from rendering? Are there resource or hydration errors?
  • DOM and state: Was the expected component present? Did loading, empty, error, or authenticated state differ? Did content disappear?
  • Viewport and clipping: Do snapshot dimensions, clip rectangle, scroll position, and iframe position match? Could a breakpoint or crop have changed?
  • Timing: Was capture triggered before the target state, font, image, or layout had settled? Does the trace show animation or asynchronous updates?

Visual mismatches are not limited to colors and spacing. A 2026 study of 307 visual-regression pull requests across 103 repositories categorized issues that included layout, appearance, color, text, state, test, and image problems. In that study’s sample, 35 of 189 analyzed issues involved non-stylistic origins such as undefined component state or disappearing content. These figures describe the authors’ dataset and method; they are not industry-wide rates. Read the study.

5. Stabilize the cause you found

Generated, live, or random data

Use fixed fixtures for visual scenarios. Seed random generators when randomness is part of the test, and mock unstable API responses. Keep values such as usernames, counts, avatars, and chart series constant so a data change cannot masquerade as a layout regression.

Current time and date

Freeze the clock when the page displays dates, relative-time labels, countdowns, or time-dependent content. Use a deliberate fixed time in the test and ensure the same value applies to both the application and any mocked API response.

Animation and transient UI

Pause or configure animations when motion is not under test. Ensure the screenshot represents the intended stable state, such as a completed transition or loaded page. Chromatic attempts to pause animations but documents cases where configuration may be needed. A generic sleep can make a race less visible without fixing it; prefer a condition tied to the UI state the test needs.

Fonts, images, and other resources

Make assets available deterministically. Serve stable local or controlled assets where practical, check response status and content, and preload web fonts when appropriate. A fallback font can change line wrapping and move the rest of the page; a missing image can alter layout even if the surrounding CSS is unchanged. Avoid depending on a remote host or CDN response that can vary between runs.

Capture readiness and application state

Wait for a meaningful signal: a target selector, a known loading indicator disappearing, a route-specific state, or network activity settling when that is appropriate for the application. Confirm that the signal corresponds to the visual state being tested. Waiting for all network activity can be unsuitable for pages with long polling or persistent connections, so use a state-specific condition when possible.

Intentional dynamic regions

Decide whether a changing region belongs in a visual assertion. If it does, make its inputs deterministic. If it does not, isolate a stable component or scenario, or mask only the smallest region whose variability is intentional. Broad masks can hide genuine changes.

6. Debug a Playwright failure interactively

Playwright’s Inspector can run a single test, select a browser project, and step through actions. From the project root, replace the path, line, and project below with the failing test’s actual values:

npx playwright test tests/example.spec.ts:10 --project=chromium --debug

Use the Inspector to pause around the capture and inspect the page state. For a test that fails only in CI, first run the same configured project and browser version locally or in a matching CI job. Playwright’s debugging guide documents the Inspector workflow; the exact documentation path and UI can change.

For any visual testing system, retain the same core evidence: expected and actual images, the diff, browser and viewport metadata, resource activity, and the DOM or application state at capture time.

7. Change one thing, then classify the result

  1. Form a specific hypothesis from the artifacts, such as “the fallback font loaded on the failing run.”
  2. Make one targeted change, such as serving the font locally or waiting for its readiness.
  3. Repeat the same test in the same environment and compare multiple captures.
  4. Record the cause and the evidence that the relevant input is now stable.
  5. If the difference remains, return to the traces and compare additional runs. If the change is real and intended, review it and update the baseline deliberately.

Retries are useful for collecting evidence or reducing the impact of a known transient while it is being investigated. They do not repair an unexplained failure. Likewise, quarantine or ignore rules are containment and should have an owner and follow-up plan, not become the final diagnosis.

Symptom-to-cause diagnostic map

Symptom Check first Evidence Likely fix
Text wraps or shifts between runs Font readiness; browser and OS consistency Font and stylesheet requests, DOM, viewport Serve stable fonts, preload when appropriate, and pin the rendering environment.
A timestamp, avatar, number, or chart changes Live data, random generation, current time, API response Fixtures, request log, repeated captures Fix fixtures or random seed, freeze time, and mock unstable responses.
Animation or loading state appears Capture timing and animation policy Trace timeline, DOM, repeated screenshots Configure animation and wait for an explicit stable state.
Image, stylesheet, or font is missing Failed or variable resource host Network panel, console, response status Provide deterministic assets and make sure they are available before capture.
Element is clipped or at the wrong breakpoint Viewport, clip, scroll, iframe position Snapshot metadata, DOM, capture configuration Correct capture dimensions or test where the component is rendered.
Only CI or one browser fails OS image, browser version, headless mode, project config Run metadata and browser-specific trace Reproduce with the baseline environment and pin the relevant configuration.
The same failure occurs every run Application state, fixture, baseline, capture definition Diff, DOM, styles, request status Treat it as a likely real UI or capture defect and investigate directly.

Common errors and fixes

Updating the baseline as soon as a diff appears

Cause: The diff has not been classified, so a real regression or test defect may be accepted as expected output.
Fix: Preserve the old baseline, inspect actual output and trace, then update only after confirming the change is intentional.

Adding a long fixed delay

Cause: The delay hides a timing race on some runs but does not establish readiness. It can also waste time and still fail on a slower run.
Fix: Wait for a relevant selector, application state, or resource readiness signal; use traces to find what was late.

Increasing a pixel threshold or masking a large area

Cause: The assertion is being loosened without finding whether the changed pixels represent noise or a meaningful state or layout defect.
Fix: Identify the changed region and its cause first. Keep any tolerance or mask narrow and justified by the assertion’s purpose.

Assuming a retry proves the failure was harmless

Cause: A later run passed, but the unstable input remains and can still hide or produce a regression.
Fix: Use the passing run as a comparison artifact and stabilize the changing condition.

Reproducing locally with a different browser setup

Cause: The OS, browser build, viewport, headless mode, or settings differ from CI or baseline generation.
Fix: Match the baseline environment before interpreting local output.

Performance, reliability, and cost considerations

Repeated captures and traces consume CI time and, depending on the service, may consume service usage. Keep reruns targeted to the failing test and browser project while diagnosing; once the cause is known, retain enough repeat coverage to confirm the fix. Waiting on a precise readiness signal is usually more efficient and informative than adding a large fixed delay, though the appropriate signal depends on the application.

Reliability improves when both the test inputs and rendering environment are controlled. Pin the browser and CI image used to generate and compare baselines, version fixtures with the test, keep assets reachable, and retain traces for failures. Document intentional masks, tolerance settings, and quarantines so later changes do not silently weaken coverage.

There is no universal cost figure for visual test systems in the evidence used here. Compare the tools and workflows you already use by what evidence they retain, how well they let you reproduce the baseline environment, whether you can debug interactions, how resources are controlled, and whether full-page or element capture metadata is available. These are selection criteria, not a product ranking.

Or skip the browser setup

If you need a clean page capture while diagnosing a visual state, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, along with newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

For a quick capture of Stripe as WebP, use the API examples in the ScreenshotNeo documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo also supports full-page or CSS-selector capture, dark mode, device presets and custom viewports, retina scale, PDF options, custom CSS and JavaScript, clicks, hidden selectors, selector or network-idle waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed public image links, asynchronous jobs with signed webhooks, bulk capture up to 100 URLs per call, a usage API, and an OpenAPI specification. It offers 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. This is a capture option; it does not replace inspecting your test runner’s trace, DOM, or environment when diagnosing flaky assertions.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Should I delete a flaky visual test?

Keep it while you gather evidence if it protects important behavior. If it must be quarantined to unblock work, track the reason and owner and continue the root-cause investigation.

Can visual diffs find bugs outside styling?

Yes. A changed capture can reveal missing content or a wrong application state as well as layout and appearance changes; inspect the DOM and application state alongside pixels.

When should I approve a new baseline?

After confirming the rendered change is intentional and the capture conditions match the baseline workflow. A passing retry alone is not sufficient.