ScreenshotNeo

BlogHow-to

How to compare Playwright screenshots in CI without flaky failures

Use Playwright screenshot assertions with a stable CI environment, deterministic page state, and carefully reviewed diff settings to catch real visual regressions.

By the ScreenshotNeo team4 October 20269 min read

Use Playwright Test’s toHaveScreenshot() assertion, and run both baseline creation and CI comparisons in the same pinned browser and operating-system environment. Make the page state deterministic before capture, mask or hide only content that is irrelevant to the behavior under test, and review every changed baseline. A generous pixel allowance cannot make an unstable screenshot meaningful.

Playwright’s screenshot assertion creates an expected image on the first run and compares later captures against it. It waits for two consecutive captures to match; by default, it disables animations, hides the caret, and uses CSS-pixel scale. Those defaults help, but they do not control your app’s test data, network responses, browser version, or host environment. See the official visual comparisons guide.

1. Install and configure Playwright Test

The examples below use JavaScript with Playwright Test. Keep the Playwright package and its browser installation aligned in local development and CI. The snapshot files are expected test artifacts: commit them, review their diffs, and update them only when the UI change is intentional.

npm init playwright@latest

A focused configuration can pin the browser project and centralize comparison policy. Use the same configuration when generating and checking snapshots.

// playwright.config.js
import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  projects: [
    {
      name: 'chromium',
      use: { ...devices['Desktop Chrome'] },
    },
  ],
  expect: {
    toHaveScreenshot: {
      animations: 'disabled',
      caret: 'hide',
      scale: 'css',
      // Choose a limit only after reviewing your UI's risk.
      maxDiffPixelRatio: 0.001,
    },
  },
});

The ratio above is an example policy value, not a universal recommendation. Start with no diff allowance if your page renders consistently, inspect actual diffs, then choose a narrow limit justified by the interface. You can instead use maxDiffPixels to cap an absolute count. Avoid setting both without a clear reason. The API documents these as separate extent limits; threshold is a different per-pixel color-difference tolerance.

2. Write a screenshot assertion

Use a whole-page assertion when the page composition is part of the test. Use a locator assertion for a component or region when that is the behavior you need to protect. A focused assertion is often easier to diagnose and less exposed to unrelated page content.

// tests/dashboard.spec.js
import { test, expect } from '@playwright/test';

test('dashboard layout is stable', async ({ page }) => {
  await page.goto('/dashboard');
  await page.getByRole('heading', { name: 'Dashboard' }).waitFor();

  // Control app-owned data through a test fixture or deterministic API response.
  await expect(page).toHaveScreenshot('dashboard.png', {
    fullPage: true,
  });
});

test('navigation component is stable', async ({ page }) => {
  await page.goto('/dashboard');
  const navigation = page.getByRole('navigation', { name: 'Primary' });
  await expect(navigation).toHaveScreenshot('primary-navigation.png');
});

For pages whose meaningful state loads asynchronously, wait for the state your test actually cares about—such as a heading, completed-results indicator, or known fixture—before asserting. A generic fixed sleep may waste CI time and still fail to synchronize correctly. Control dates, randomized values, user-specific data, and network responses in the application’s test setup where those values are not the subject of the visual check.

3. Keep capture conditions consistent

Rendering can differ across operating systems, browser versions, settings, hardware, power source, and headless mode. Generate baselines and compare them in the same environment wherever practical; the Playwright guide calls out these environment differences explicitly.

  • Pin the Playwright version and install the matching browser version in CI.
  • Use the same operating-system family and container image for baseline generation and CI checks.
  • Keep project settings, viewport, device scale factor, color scheme, locale, and fonts consistent.
  • Do not casually regenerate baselines on a developer machine with a different browser or OS than CI.
  • When upgrading Playwright, its browser, the OS image, or fonts, expect that the rendering environment change may require reviewed baseline updates.

There is no universally correct CI image or tolerance: choose and document an environment your team can reproduce. A platform-specific snapshot name may include project or platform context, so inspect the generated paths and keep project naming stable.

4. Handle animation and volatile content narrowly

Playwright’s screenshot assertion disables animations and hides the caret by default. Keep those defaults unless the animation or caret itself is what you need to test. The assertion also waits for two consecutive screenshots to be identical before comparison, but this stabilization does not replace making the application state deterministic.

For intentionally irrelevant variation—such as a timestamp or rotating ad—mask only the relevant element, or use a screenshot stylesheet. Do not mask areas whose visual behavior is part of the feature under test.

test('account page ignores its changing timestamp', async ({ page }) => {
  await page.goto('/account');
  await expect(page).toHaveScreenshot('account.png', {
    mask: [page.getByTestId('last-updated-time')],
  });
});

A stylesheet can hide or normalize targeted elements for the screenshot. stylePath was added in Playwright v1.41, so confirm the installed version before using it. Keep the rule scoped and review the resulting image to ensure it does not conceal a defect.

/* tests/visual-snapshot.css */
[data-testid="last-updated-time"] {
  visibility: hidden !important;
}
test('page snapshot with targeted style', async ({ page }) => {
  await page.goto('/account');
  await expect(page).toHaveScreenshot('account.png', {
    stylePath: './tests/visual-snapshot.css',
  });
});

5. Set diff tolerances deliberately

Playwright’s pixelmatch comparator uses a YIQ color difference. Its documented default threshold is 0.2, on a scale from strict 0 to lax 1. This threshold controls how different an individual pixel may be before it counts as different. maxDiffPixels and maxDiffPixelRatio instead constrain how many pixels, or what share of the image, may differ. See the SnapshotAssertions API.

Option What it limits Use it for
threshold Per-pixel color difference Small rendering or color variation that is acceptable for this UI
maxDiffPixels Absolute count of different pixels A fixed-size component or screenshot where a count has clear meaning
maxDiffPixelRatio Share of pixels allowed to differ Images whose dimensions may vary and where a proportional limit is appropriate

These tolerances are not substitutes for a stable capture. Prefer a strict threshold that produces useful diffs, and set one extent limit after inspecting representative failures. A permissive allowance can hide real regressions, especially on small components where a handful of changed pixels matter.

6. Generate, inspect, and update baselines

  1. Run the screenshot test in the pinned environment with --update-snapshots to create the initial expected image.
  2. Inspect and commit the snapshot files alongside the test that explains what they protect.
  3. On later runs, inspect the expected, actual, and diff artifacts when an assertion fails.
  4. Decide whether the difference is a bug, capture noise, or an intentional UI change.
  5. For an intentional change, regenerate with --update-snapshots, review the image diff, and commit the updated expectation.
npx playwright test --update-snapshots
npx playwright test

Snapshots are stored under a directory based on the test file, typically with a -snapshots suffix. Snapshot names can also incorporate browser or project context. Keep the baseline changes reviewable in version control; do not automatically accept every changed image in CI.

7. Common failures and fixes

Symptom Likely cause Fix
Large diffs on every CI run Different OS, browser build, fonts, viewport, or device scale factor Pin the environment and project settings used to create the baseline and compare it.
Only a timestamp, avatar, or rotating region differs Volatile content is part of the captured page Stabilize test data, or narrowly mask/style that specific element if it is outside the test’s scope.
Screenshot changes between repeated captures App state is still changing, or asynchronous content has not reached the target state Wait for the meaningful state, control app-owned data and responses, and inspect animations or timers. The built-in two-capture check cannot stabilize an indefinitely changing page.
Assertion times out Navigation or the target locator never reaches the expected state, or the page remains visually unstable Check the navigation outcome and locator; wait on a specific state rather than increasing a timeout blindly.
Baseline is missing No expected snapshot has been created for that test, or the snapshot path/project naming changed Generate it in the intended environment with --update-snapshots, then review and commit it.
Failures start after a Playwright upgrade The browser or screenshot behavior changed with the installed version Confirm package/browser version alignment, inspect diffs, and update baselines only for understood changes.
Many genuine changes pass unexpectedly Diff limits are too permissive, or too much of the page is masked/hidden Reduce the allowance and remove masks that cover behavior the test should catch.
stylePath is rejected or ignored The installed Playwright version predates its support Check the version; stylePath was added in v1.41.

8. CI workflow and operating costs

Visual assertions add browser rendering and image comparison work to the test job. Keep the number and scope of captures aligned with the risks they cover: a locator screenshot may be cheaper to diagnose and less sensitive to unrelated page content than a full-page image. Whole-page coverage remains useful when layout across the page is itself important. Stabilizing the state and environment usually saves more CI time than repeatedly rerunning flaky captures.

For reliability, make visual failures produce artifacts developers can inspect, keep baselines in the same repository workflow as the test, and require review for baseline updates. Avoid treating retries as a fix: retries can reveal intermittent behavior, but a passing retry does not explain or remove the source of capture variance.

The research sources do not establish a universal runtime, cost, or flake-rate reduction. Measure your own CI job before increasing coverage, and use a small set of representative, high-value screenshots as a starting point.

9. Or skip the browser setup

If your goal is to capture a page for a report, fixture, or downstream review rather than assert against a committed Playwright baseline, ScreenshotNeo provides a screenshot API and MCP server. A single GET request can return PNG, JPEG, WebP, or PDF. It is not a replacement for a deterministic visual regression assertion, but it can handle the browser capture setup for you.

See the ScreenshotNeo API documentation for the API options. This cURL example saves a WebP capture:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

The equivalent Python call is:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the request was billed. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Can I compare screenshots without Playwright Test?

The documented toHaveScreenshot() assertion is part of Playwright Test’s runner workflow. For a reliable visual baseline check, use that runner or choose another comparison workflow with explicit baseline and review behavior.

Should I use full-page or locator screenshots?

Choose full-page capture when the page’s overall composition is under test. Choose a locator when a specific component is the risk you need to monitor and unrelated content would make the assertion noisy.

Does a passing screenshot assertion prove the page is correct?

No. It means the captured image is within the configured comparison limits of the expected image. The test still needs appropriate coverage, meaningful state setup, and review of its baseline.

Which tolerance is right for every project?

There is no universal value. Select a threshold and pixel limit based on the UI, rendering environment, and defects the test must catch, then revisit them when those conditions change.