ScreenshotNeo

BlogEngineering

How to Detect and Fix Flaky Visual Regression Tests

Learn how to tell flaky screenshot diffs from real regressions, find the source of unstable captures, and make visual tests deterministic without hiding defects.

By the ScreenshotNeo team4 October 20269 min read

A flaky visual regression test produces different results on repeated runs even though the code and intended state have not changed. Before updating a baseline, repeat the capture, compare the baseline, actual screenshot, and diff, then inspect the rendered state and trace evidence. Fix the source of variation—such as changing data, animation timing, late assets, or environment drift—before adjusting thresholds or accepting new pixels.

1. Confirm that the test is flaky

Keep the application code, test inputs, browser, and intended state fixed. Re-run the same test and compare its baseline, actual image, and diff. If the changed pixels recur consistently, the test may be exposing a real regression. If they vary across otherwise equivalent runs, investigate capture instability. Chromium defines flaky tests as tests that pass sometimes and fail other times (Chromium: Fixing Flaky Tests).

  1. Save the failing run’s screenshot, diff, test report, and trace.
  2. Repeat the same test without changing code or approving a baseline.
  3. Compare the changed regions across runs. Note whether they move, disappear, or consistently show the same UI change.
  4. Record the browser, operating system, test data, and relevant configuration for each run.

A single mismatch is not enough to conclude either that the UI is broken or that the diff is harmless. Do not accept a new baseline until you understand the difference.

2. Inspect what the browser rendered

Use the runner’s report and trace to reconstruct what happened before the screenshot. Playwright recommends traces for debugging failed tests and collecting them on the first CI retry; its trace viewer can show actions and page state around the failure (Playwright best practices, Trace Viewer).

  • Late fonts or images: Did the capture happen before a font swap, image decode, or lazy-loaded asset completed?
  • Changing content: Did a timestamp, randomized value, live counter, rotating message, or server response differ?
  • Animation: Was a transition, carousel, spinner, video, or blinking cursor captured at a different frame?
  • Network variation: Did a third-party request stall, fail, or return different content?
  • Unsettled layout: Did content shift after the screenshot because data, fonts, or responsive measurements arrived late?
  • Test coupling: Did another test modify shared data, storage, or a snapshot with a colliding name?

Chromatic’s instability guidance describes traces that include network requests, console logs, DOM snapshots, and snapshot metadata; these can help identify the source instead of treating the final image as the only evidence (Chromatic: Unstable tests).

3. Make the capture inputs deterministic

A screenshot comparison is meaningful only when the state being compared is controlled. Stabilize inputs before relaxing comparison thresholds.

Control time, randomness, and data

  • Freeze the clock for screens that render dates, relative times, countdowns, or time-based greetings.
  • Seed random values or replace them with fixed fixtures.
  • Use controlled application and database data. Keep a staging fixture unchanged during comparisons where practical.
  • Isolate tests so one test cannot leave storage, server data, or other shared state for another.

Control external resources

Mock or otherwise control third-party responses that affect the captured state. Prefer stable assets and predictable font loading when the visual result depends on them. A live advertisement, embedded widget, or remote image can change independently of your code. Playwright recommends controlling third-party dependencies in tests; Chromatic also recommends stable assets and predictable font loading (Playwright best practices, Chromatic: Unstable tests).

Keep rendering environments consistent

Pin the browser and operating-system environment used to generate and compare snapshots. Browser engines, OS text rendering, installed fonts, and device settings can change pixels. Playwright specifically notes that screenshots can differ across browsers and platforms, and recommends generating and comparing them in consistent environments (Playwright: Visual comparisons). If cross-platform coverage matters, maintain separate baselines for the environments you intend to support instead of comparing unlike renderers as though they were identical.

Control animation only when it is outside the test’s purpose

For a static visual check, disable or pause transitions and animations using the runner’s supported option or a test-only stylesheet. If animation behavior itself is what the test covers, do not remove it; control the animation time or assert a defined frame instead. A blanket animation suppression can hide a defect in motion or transition behavior.

4. Wait for the intended UI state

Synchronize on a meaningful condition: expected content is visible, a loading indicator is gone, a known selector appears, or the application reports that the relevant state is ready. An arbitrary sleep may allow a page to settle in one run while leaving the underlying race intact. Use a delay only when the UI has no reliable readiness signal, and treat it as a fallback to investigate rather than proof of determinism.

Check that “ready” covers the assets that matter. A heading becoming visible does not guarantee that its custom font, hero image, or asynchronously loaded chart has finished rendering. Wait for those specific resources or a stable application signal when they affect the pixels under test.

5. Avoid collisions and shared-state leakage

Give each snapshot a unique, descriptive name that includes its test and the UI state. Cypress documented a case in which two tests used the same Percy snapshot title, coupling their results; its debugging guidance recommends unique names based on the full test and state (Cypress: Debugging flaky visual regression tests).

// Example naming pattern
checkout-form--empty
checkout-form--validation-error
checkout-form--submitted

Also check for shared accounts, reused records, persistent browser storage, and tests running concurrently against the same mutable fixture. Reset or isolate those inputs between tests.

6. Use retries, thresholds, and filters with care

Retries help reveal intermittent outcomes and can preserve artifacts from another attempt. They do not identify or fix the cause. A test that passes on retry may still be unreliable, and a real regression can also behave inconsistently when the affected state is timing-sensitive.

Pixel tolerances can accommodate known rendering noise, but a wider threshold may also accept a real visual change. First stabilize inputs and inspect the remaining diff. If a small difference is an understood rendering artifact, document why the tolerance is appropriate and keep it as narrow as practical.

Chromatic’s flake filter can repeatedly render a test, ignore a detected unstable test for that build, and reevaluate it on later builds. Its documentation says ignored tests still count toward billed snapshot use (Chromatic: Unstable tests). If you quarantine or filter a test, track it, make its status visible, and assign follow-up work. An ignored diff is not evidence that the UI is correct.

7. Choose a workflow that fits the test suite

There is no universally best visual testing tool established by the available documentation. Compare options against your existing runner and review process rather than treating vendor feature descriptions as an independent benchmark.

Consideration Questions to answer
Framework fit Does it work with your browser runner, component stories, and CI workflow?
Rendering coverage Which browser, device, and operating-system environments generate the captures?
Debugging evidence Can reviewers inspect diffs, traces, DOM state, network requests, console output, and snapshot metadata?
Review and policy How are baselines approved? How are unstable tests surfaced, quarantined, and revisited?
Hosting and data Can screenshots and page content be sent to a hosted service under your privacy and security requirements?
Cost How are renders, snapshots, storage, retries, and filtered tests counted? Check current service terms directly.

Playwright provides visual snapshot comparisons with repository-managed baselines that should be reviewed and committed with the tests (Playwright: Visual comparisons). Chromatic documents integrations with Playwright, Cypress, and Vitest (Chromatic documentation). Cypress lists hosted visual testing integrations including Argos, Happo, Percy, Sauce Labs Visual, SmartBear VisualTest, and Chromatic (Cypress: Visual testing). These sources describe their respective workflows; they do not establish a universal ranking. Verify current features, prices, retention, and terms with each provider.

8. Troubleshooting common visual test failures

Symptom Likely cause What to do
Text wraps differently between runs Font not loaded before capture, font fallback, or different OS/browser environment Wait for the intended font, verify it loaded, and pin the rendering environment or use environment-specific baselines.
Image or chart is missing in some captures Capture precedes asset load, lazy loading, or a remote request is unreliable Wait for the specific asset or app-ready state; control the response or use a stable fixture.
Only animated regions differ Capture happened at different animation frames Pause irrelevant animation or set a defined frame; keep animation enabled when it is the behavior under test.
Dates, counters, or labels change Clock, random input, or live data varies Freeze time, seed randomness, and supply deterministic data.
Test passes on retry Race condition, network variation, shared state, or layout timing Compare artifacts and traces from both attempts, then fix the varying input or readiness condition. Keep retry data visible.
A baseline update affects another test Snapshot names collide or state is shared Use unique test-and-state names and isolate mutable data and browser storage.
Diff appears only on CI CI uses another browser, OS, font set, viewport, device scale, or resource timing Compare environment metadata and pin the capture image and browser configuration used for baselines.
A threshold hides a visible change Tolerance is too broad or applied to the wrong region Review the diff, narrow the threshold, and stabilize the source of noise instead of widening tolerance further.
Hosted flake filtering stops a build failure The service classified a test as unstable for that build Keep the issue tracked and revisit the test; filtering contains the build impact but does not establish correctness.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single GET request captures a URL as PNG, JPEG, WebP, or PDF. For visual checks, it can give you a repeatable capture endpoint without setting up browser automation in your own job. It does not replace the need to control test data, state, and rendering conditions when diagnosing a flaky application test.

Cookie and consent banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether it was billed. Its MCP server gives AI agents the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is available on every plan. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Sign up free for 1,000 screenshots a month, with no card required.

Performance, reliability, and cost

  • Keep the suite efficient: Capture only the states that protect meaningful UI behavior. Repeated retries and broad page captures increase work and can make noisy tests slower to diagnose.
  • Make reliability observable: Retain the first failure’s screenshot, diff, trace, and environment metadata. Track flaky tests separately from stable failures so retries do not erase evidence.
  • Account for rendering cost: Hosted tools may count retries, snapshots, storage, or filtered tests differently. The cited sources do not establish a cross-provider price comparison; check current terms before choosing.
  • Protect sensitive content: Review whether captured pages contain private data before uploading images or traces to a hosted service.
  • Use thresholds sparingly: A broad tolerance may reduce review noise while allowing genuine defects through. Fix unstable inputs first.

Frequently asked questions

Should I update the baseline when a screenshot test fails?

Only after reviewing the UI change and confirming it is intended. Reproduce the capture and understand the changed pixels first; accepting a baseline can otherwise encode a transient failure or a real defect.

Is a test that passes on retry safe to ignore?

No. The retry is evidence of intermittent behavior. Keep the trace and screenshots, identify the varying condition, and track the fix even if retries temporarily reduce CI disruption.

Should all visual tests disable animations?

No. Disable or pause motion only when it is irrelevant to the state being checked. Tests for animation behavior need a controlled timeline or defined frame.

Can I compare screenshots from different operating systems?

You can, but treat each rendering environment deliberately. Platform differences in fonts and rendering can require separate baselines; a mismatch across unlike environments does not by itself show an application regression.

Do visual flake filters fix unstable tests?

No. They can contain the effect on a build and identify instability, but the underlying source still needs investigation and follow-up.