ScreenshotNeo

BlogGuides

How to Stabilize Images in Visual Tests

Make visual tests reliable by controlling page state, resources, animation, and rendering conditions before adjusting screenshot diff tolerance.

By the ScreenshotNeo team4 October 202610 min read

To stabilize images in visual tests, make the page state and capture environment repeatable before loosening pixel-difference thresholds. Fix the route, viewport, locale, user state, and test data; wait for the actual UI and its resources to be ready; disable or pause motion; and keep the browser environment consistent. Use masks only for specific content that is intentionally volatile, then tune tolerances only for small rendering differences that remain.

This guide uses Playwright for a runnable example. Playwright’s toHaveScreenshot assertion waits for two consecutive page screenshots to match before comparing one with the baseline. That helps with transient rendering, but it cannot make changing application data deterministic for you. See the official PageAssertions documentation.

1. Identify what makes a visual test flaky

An unstable snapshot is a visual test whose output changes between runs even though the intended product state has not changed. The cause is usually an uncontrolled input or a page captured before it reaches the intended state.

  • Changing data: timestamps, randomized content, rotating banners, user-specific values, or data from a live API.
  • Late resources: fonts, images, or other assets arrive after the screenshot. Image and font sources that are unreliable can fail to load within a capture window.
  • Motion: a capture lands on a different frame of a CSS transition, GIF, canvas, or JavaScript-driven animation.
  • Unclear readiness: the test waits a fixed number of milliseconds, but the relevant UI is still rendering or loading.
  • Different rendering conditions: viewport, browser, operating system, locale, or device scale changes between baseline and comparison.
  • Overly broad comparison settings: large masks or generous pixel tolerances hide genuine regressions.

Start by inspecting the expected image, actual image, and diff image. Determine whether the mismatch is an intended change, a rendering condition, or genuinely variable page content before updating a baseline.

2. Make the test inputs repeatable

Define the exact state the screenshot represents. Keep these inputs fixed for both baseline creation and comparison:

  • Route, query parameters, viewport, and device scale.
  • Locale, timezone, and any user or permission state that affects rendering.
  • Test records and API responses. Prefer fixtures or deterministic mocks over live changing data.
  • Current date and time, when they appear in the UI. Freeze them or inject a fixed value into the application.
  • Random values. If randomness is needed, use a seed so the same inputs produce the same output.

Chromatic also recommends fixed input data or seeded randomness when debugging unstable tests; its guide is useful even if your capture runs in another system: Unstable tests debugging.

A deterministic route and fixture are more useful than repeatedly increasing a screenshot timeout. If external data is part of what you intend to test, control the test environment and make that data stable for the run.

3. Wait for the UI condition that matters

Wait for a user-visible or application-specific condition that proves the relevant state is ready. Avoid treating an arbitrary delay as proof. Playwright’s retrying assertions and actionability checks wait for expected conditions; its stability check is tied to whether an element is moving or has completed animation. See Playwright auto-waiting.

For example, wait for the heading that identifies the loaded view, then wait for a key image to finish loading before capturing. The exact readiness signal depends on the application. A selector existing in the DOM may not be sufficient if it is still hidden, empty, or awaiting data.

4. Runnable Playwright example

The example below uses Playwright Test with TypeScript. It fixes the viewport, waits for a meaningful page condition, waits for a hero image to load, disables CSS animation for the screenshot, masks one intentionally changing timestamp, and relies on toHaveScreenshot to wait for consecutive stable screenshots.

import { test, expect } from '@playwright/test';

test('product page is visually stable', async ({ page }) => {
  await page.setViewportSize({ width: 1280, height: 800 });
  await page.goto('http://127.0.0.1:3000/products/widget');

  // Wait for the intended page state, not just for elapsed time.
  await expect(page.getByRole('heading', { name: 'Widget' })).toBeVisible();

  // If this image is important to the screenshot, wait for it to be decoded.
  await page.locator('[data-testid="hero-image"]').evaluate(async (element) => {
    const image = element as HTMLImageElement;
    if (!image.complete) {
      await new Promise<void>((resolve, reject) => {
        image.addEventListener('load', () => resolve(), { once: true });
        image.addEventListener('error', () => reject(new Error('Hero image failed to load')), { once: true });
      });
    }
    if (image.decode) await image.decode();
  });

  await expect(page).toHaveScreenshot('product-widget.png', {
    fullPage: true,
    animations: 'disabled',
    mask: [page.locator('[data-testid="updated-at"]')],
    maskColor: '#808080',
  });
});

Replace the example URL and selectors with the route and elements in your app. The image wait is appropriate when that asset is part of the expected visual result; if the image is optional, define the expected fallback state rather than letting the test wait indefinitely.

Playwright documents that the screenshot assertion waits for two consecutive page screenshots to yield the same result before comparing against its expectation. It also supports animation control and masks. See the assertion API and visual comparisons.

Configure a stable browser project

Keep the project that generates the baseline and the project that checks it on the same browser and viewport. For example, in playwright.config.ts:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  projects: [
    {
      name: 'visual-chromium',
      use: {
        browserName: 'chromium',
        viewport: { width: 1280, height: 800 },
        deviceScaleFactor: 1,
        locale: 'en-US',
        timezoneId: 'UTC',
      },
    },
  ],
});

This is a concrete baseline configuration, not a universal requirement to use Chromium or these dimensions. Choose the browser and viewport your users need, then keep them consistent. If you intentionally test multiple viewports or browsers, maintain the corresponding baselines separately.

5. Handle animation and other volatile regions

CSS animation and transitions

Playwright screenshot assertions can disable CSS animations with animations: 'disabled'. Use this when motion is irrelevant to the visual assertion. If the design itself depends on an animated state, pause at a known frame or set the application to a deterministic state instead of removing the motion.

JavaScript-driven animation

CSS animation controls do not guarantee that a JavaScript animation, canvas, video, or application timer is frozen. Add an application-level test hook to pause it, set a fixed frame, or wait for an explicit stable state. Chromatic’s hosted capture can standardize the browser environment and uses network quiescence as a capture heuristic, but its documentation says JavaScript-driven animations still need author intervention: Chromatic snapshots.

Mask only content that cannot be controlled

A mask covers a region so its changing pixels do not fail the comparison. Use it for a narrowly selected value that is meant to vary and is not practical to freeze, such as a live third-party count. Prefer freezing a timestamp or using fixed fixture data if that content is part of the behavior you need to check.

A mask removes coverage: a visual bug inside the masked region will not be caught by that comparison. Keep masks small and specific, and use a visible mask color so the excluded area is apparent in review.

6. Tune diff tolerance only after stabilization

Masks and thresholds address different problems. A mask excludes a region from meaningful comparison; a threshold allows some pixel differences across the compared image. Neither should be widened casually because either can let a real visual regression pass.

Playwright supports maxDiffPixels for setting an allowed number of different pixels. Set it only after stabilizing data, resources, motion, and the browser environment. Prefer a small, justified allowance. Avoid using a broad percentage or pixel budget to make a flaky test green without understanding the diff. Details and configuration examples are in Playwright visual comparisons.

7. Hosted capture and local capture

A hosted visual-testing service can make browser and viewport conditions more consistent by capturing in a standardized environment. Chromatic documents a fleet of standardized browsers and mobile emulators at specified viewport sizes. That consistency does not automatically make application state deterministic, and network quiescence is a heuristic rather than proof that every asset or app-specific condition is ready. JavaScript-driven motion still needs to be paused by the test author.

For any capture approach, assess the same practical dimensions: browser and operating-system consistency, viewport control, control over test data and resources, animation behavior, masks and thresholds, and how reviewers approve or update baselines. The available documentation supports the Playwright controls and Chromatic characteristics above; it does not establish a current, complete vendor-by-vendor comparison or pricing.

8. Troubleshooting flaky visual tests

Symptom Likely cause Fix
The diff changes on every run Random, time-based, user-specific, or live data Use fixed fixtures, freeze displayed time, or seed randomness.
Text moves or wraps differently A font or stylesheet loaded late, or the viewport differs Wait for the intended page state and font readiness; pin viewport and browser settings; check that the same font assets are available.
An image is blank in some captures The asset did not load before capture, or its source is unreliable Wait for the specific image to load and decode; use a stable test asset or deterministic fallback when the source is outside the test’s control.
Only animated areas fail The screenshot catches different animation frames Disable CSS motion for the assertion or pause JavaScript animation at a defined state.
The page is captured before content appears The test uses navigation completion or a fixed delay as its only readiness condition Wait for a locator or application signal that represents the content being ready.
Many unrelated pixels differ Browser, viewport, device scale, locale, or rendering environment changed Compare using a consistent project and capture environment, then regenerate baselines only for intentional changes.
The test passes but a visible bug remains A mask is too large or the threshold is too permissive Narrow the mask or reduce tolerance, then inspect expected, actual, and diff images.
Hosted capture still shows motion Capture standardization does not pause application-specific JavaScript animation Add a test hook that pauses the animation or waits for the intended state.

9. Performance, reliability, and cost considerations

Stability controls add work only where they establish a necessary condition. Waiting for a specific image, font, or application state can increase capture time if that resource is slow; use the smallest set of readiness conditions that represents the screenshot you need. An arbitrary long delay is usually slower and less reliable because it can still miss a late or failed resource.

Retries can help distinguish transient rendering from a consistent mismatch, but repeated retries do not repair nondeterministic inputs. Keep failure artifacts so a reviewer can inspect the expected, actual, and diff images. Update the baseline only after confirming the change is intended.

For service costs, check the current provider’s own plan details before adopting a hosted workflow; the cited documentation here does not establish current prices for the visual-testing services discussed. ScreenshotNeo’s screenshot API pricing is listed below for one-off or automated website captures.

Or skip the browser setup

For a website screenshot outside your local visual-test runner, ScreenshotNeo takes a screenshot with one GET request. See the ScreenshotNeo API documentation for the available parameters. For example, this cURL request saves a WebP capture of Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. For controlled visual regression suites, keep your test inputs and rendering conditions deterministic as described above; a one-call screenshot does not replace those controls.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

Why are my Playwright visual tests flaky?

Common causes are changing data, late fonts or images, motion, and differences in the browser or viewport. Inspect the diff and stabilize the relevant input before changing the threshold.

Does Playwright wait for the screenshot to become stable?

Yes. toHaveScreenshot waits for two consecutive page screenshots to match before it compares the result to the expected image. That does not freeze every source of changing application data.

Should I mask a timestamp?

Freeze it if the timestamp is part of the behavior under test. Mask it if the exact value is irrelevant and cannot be controlled, keeping the mask limited to the timestamp.

Will a hosted capture service stop all animations?

No general assumption is safe. Chromatic documents standardized browser capture, but says authors must pause JavaScript-driven animations themselves.

Sources