ScreenshotNeo

BlogHow-to

How to Fix Flaky Visual Tests

Find why visual tests change between runs, then stabilize data, assets, fonts, and animation before updating a screenshot baseline.

By the ScreenshotNeo team4 October 20268 min read

A flaky visual test produces different screenshots across runs even when you did not intend to change the UI. Start by diagnosing the capture before updating the baseline: inspect the changed region, trace, network requests, console output, DOM state, and capture metadata. Then stabilize the inputs that caused the difference.

Most useful fixes are concrete: make test data deterministic, serve reliable assets, wait for intended fonts and page state, handle motion deliberately, and mask only genuinely volatile regions. Retries can reveal intermittency, but they do not repair its cause.

1. Diagnose the screenshot before changing the baseline

  1. Reproduce the failure with the same browser, viewport, fixtures, and environment as the failing run.
  2. Compare the expected and actual images. Locate the changed pixels and classify the pattern: a font swap, missing image, moving element, timestamp, animation frame, or a real layout change.
  3. Open the test trace or provider diagnostics. Inspect network requests, console errors, DOM snapshots, and snapshot metadata around capture time.
  4. Fix the source of variation, then rerun under controlled conditions. Update the baseline only after confirming the visual change is intended.

This trace-first workflow follows the diagnostic guidance in Chromatic’s unstable-test guide. Do not begin with a longer timeout or a wider mask unless the evidence shows that is the appropriate fix.

2. Remove common sources of nondeterminism

Use deterministic data and application state

Random values, current timestamps, rotating content, and data fetched from changing services can alter a screenshot without a code change. Use fixed fixtures or a seeded generator, freeze time where appropriate, and control the state the test renders. If the test depends on an API, return a known response rather than relying on an unpredictable live service.

Make fonts and assets available consistently

A late web font can change text width, line wrapping, and layout after the first paint. Make sure the intended font files load reliably and preload them when appropriate. Prefer stable local or static image and font resources over unreliable remote hosts, and keep image optimization behavior consistent. Check failed requests and missing stylesheets as well as the visible diff; an asset problem often looks like a layout regression.

Chromatic’s guidance on resource loading discusses retries for assets that fail to load in time and diagnosing missing images, fonts, stylesheets, and domains. Provider-specific retries and capture behavior should not be assumed to apply to another runner.

Handle animation according to what the test asserts

For a static visual comparison, pause or disable CSS transitions, CSS and SVG animations, videos, and blinking cursors so each run captures a comparable state. If motion is the behavior under test, keep it enabled and assert its intended behavior deliberately instead of comparing arbitrary frames.

Capture tools differ in how they pause motion. For example, Chromatic documents pausing CSS transitions, CSS and SVG animations, and videos, with configurable behavior for where an animation is paused. Do not assume another tool has the same defaults.

Wait for a meaningful condition

Wait for the state the assertion needs: a selector to appear, a loading indicator to disappear, a known data response to complete, or the intended content to render. Arbitrary sleeps may hide timing symptoms in one environment while leaving the underlying race in place. Use trace evidence to identify which request or state is late, then wait for that condition or make it deterministic.

Mask only deliberately variable regions

Hide, mask, or normalize content such as a live clock only when it is intentionally variable and outside the behavior the test covers. Keep the region as small as possible: masking a whole panel can hide a real visual regression. Playwright’s screenshot assertions support options for hiding or modifying dynamic regions; consult the API for the syntax supported by the version installed in your project.

3. Example: stabilize a Playwright screenshot assertion

Playwright Test provides screenshot assertions and retry options. This example waits for a meaningful page state, disables motion for the static comparison, and masks a timestamp that is intentionally outside the test’s scope. Adjust selectors and option details for your installed Playwright version.

import { test, expect } from '@playwright/test';

test('dashboard visual state is stable', async ({ page }) => {
  await page.goto('http://127.0.0.1:3000/dashboard');
  await page.getByRole('heading', { name: 'Dashboard' }).waitFor();

  await expect(page).toHaveScreenshot('dashboard.png', {
    animations: 'disabled',
    mask: [page.locator('[data-testid="current-time"]')],
    timeout: 10_000,
  });
});

The masked element should be genuinely volatile and irrelevant to the behavior under test. If the page uses remote data or fonts, control those inputs too; a mask does not make an unstable page healthy. See the Playwright visual comparison guide and the PageAssertions API for current assertion options.

4. Treat retries as a signal, not a repair

Playwright retries are off by default unless configured. Its documentation calls a test that fails initially and passes on retry “flaky.” That label is useful because it surfaces intermittency in results; it does not identify or fix the underlying difference. A retry-only green run is not evidence that the cause is gone. Keep retry results visible, investigate the first failure, and verify the fix across controlled reruns. See Playwright’s retry documentation.

5. Troubleshooting common visual-test failures

Symptom Likely cause What to check and fix
Text wraps differently or elements shift Late or missing web font Inspect font requests and console errors; serve the intended font reliably and wait for the page’s real ready state.
Image, icon, or background is missing Failed, late, or inconsistent asset request Check the trace and resource-loading diagnostics; stabilize the host or serve a fixed asset in the test.
A small region changes on every run Clock, random value, rotating data, or animation Freeze or seed the input, control the fixture, or narrowly mask a region that is outside the assertion.
Failure appears only in CI Different browser, viewport, fonts, network, or timing Reproduce with CI’s browser and viewport; inspect CI trace, requests, and console output before changing timeouts.
A longer timeout makes it pass sometimes Uncontrolled wait condition or late resource Find what is still loading in the trace and wait for a specific state or stabilize the resource source.
Test passes only after retry Intermittent capture or application state Inspect the initial failure and keep retry status visible; retrying contains detection but does not fix the cause.
Large masked area makes the diff pass Mask is hiding meaningful UI Reduce the mask to the truly volatile element and restore assertions for surrounding layout.
Diff consistently shows the same change Intentional or accidental real UI change Review the affected component and product intent; update the baseline only if the change is expected.

6. Choose a visual testing approach

Compare tools by how well they fit your browser runner and component framework, what controls they provide for motion and dynamic content, how they handle fonts and external assets, and what evidence they show when a capture is unstable.

  1. ScreenshotNeo: useful when you need website screenshots through an API or MCP server; cookie banners, popups, and chat widgets are removed before capture, and bot checks, blank pages, failed loads, and cache hits are not billed.
  2. Playwright native screenshot assertions: fit teams already using Playwright who want visual comparisons in their browser tests. The official guide documents screenshot assertions and dynamic-content handling.
  3. Chromatic: its documentation covers hosted visual testing, traces, animation behavior, and resource loading, which can help teams investigating unstable captures.
  4. Percy: Percy’s vendor article describes integrations with Jest, Cypress, Playwright, and Selenium, plus snapshot stabilization behavior. Verify current compatibility and features with the provider.

This is a focused comparison, not a complete current feature or pricing survey. Check each provider’s current documentation for details before choosing.

7. Or skip the browser setup

For a website screenshot without configuring a browser runner, ScreenshotNeo accepts one GET request with a URL and returns an image or PDF. See the ScreenshotNeo API documentation for parameters and formats.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. Response headers report the page verdict and billing status.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.

8. Performance, reliability, and cost considerations

  • Reduce waiting without guessing: fixed fixtures, stable assets, and explicit readiness conditions reduce unnecessary waits while making the capture more predictable.
  • Keep external dependencies controlled: third-party resources can fail or arrive late; local or stable test resources make failures easier to reproduce.
  • Use masks sparingly: a narrowly masked clock is cheaper to reason about than a broad region that can conceal a regression.
  • Budget retries deliberately: retries add executions and time while potentially obscuring intermittent failures. Use them to classify and report flakes, not as proof of a fix.
  • Check provider billing semantics: ScreenshotNeo bills only clean shots; responses identify page verdict and billing status. Its plans include Free (1,000 monthly), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000). Yearly billing gives two months free, and every feature is on every plan.

9. Short FAQ

Should I update the baseline whenever a screenshot fails?

No. First establish whether the difference is intended. A baseline update can otherwise bless missing assets, unstable data, or an actual regression.

Are visual diffs always pixel-perfect?

Comparison and tolerance behavior depends on the runner and configuration. Check the assertion API and provider settings in use rather than assuming all tools compare images identically.

Can I use screenshots from an API as my visual test baseline?

Yes, if the API capture settings and target page state are controlled consistently. Keep the same viewport, format, and relevant capture options between baseline and later captures.

Where should I start if local runs pass but CI fails?

Start with the CI trace and reproduce its browser, viewport, and resource conditions. The difference often points to a missing font or asset, a timing race, or an environment mismatch.