ScreenshotNeo

BlogGuides

Playwright Image Comparison: Visual Regression That Stays Reliable

Learn how Playwright compares screenshots, updates baselines, controls flaky diffs, and scales visual regression testing.

By the ScreenshotNeo team1 October 202610 min read

Use Playwright Test’s toHaveScreenshot() assertion. The first run creates a reference image; later runs capture the page again and compare it with that reference. Playwright waits for two consecutive screenshots to match before it compares the result, which removes many timing-related false positives.

For a stable workflow:

  1. Stabilize data, time, network responses and interaction state.
  2. Capture only the page or locator region that matters.
  3. Keep browser, operating system and rendering settings consistent.
  4. Review actual, expected and diff artifacts when a check fails.
  5. Update snapshots only after deliberately approving a visual change.

1. Create a screenshot comparison test

Install Playwright Test and create a test file such as tests/home.spec.ts:

import { test, expect } from '@playwright/test';

test('home page matches its visual baseline', async ({ page }) => {
  await page.goto('http://localhost:3000/', { waitUntil: 'networkidle' });
  await expect(page).toHaveScreenshot('home.png', {
    fullPage: true,
  });
});

Run it once to create the reference:

npx playwright test tests/home.spec.ts

Commit the generated snapshot directory to version control. On later runs, Playwright writes the new capture, compares it with the committed reference and reports a failure when the difference exceeds your configured limits. The official snapshot guide covers the baseline directory and naming rules: Playwright screenshot comparisons.

Choose a deterministic snapshot name

A descriptive name prevents collisions when several tests use the same page:

test('checkout is stable', async ({ page }) => {
  await page.goto('http://localhost:3000/checkout');
  await expect(page).toHaveScreenshot('checkout-empty-cart.png');
});

Compare a focused component or region

Use a locator assertion when the rest of the page contains unrelated or volatile content:

test('pricing table matches', async ({ page }) => {
  await page.goto('http://localhost:3000/pricing');
  const pricing = page.locator('[data-testid="pricing-table"]');
  await expect(pricing).toHaveScreenshot('pricing-table.png');
});

Scoped screenshots usually produce smaller, more understandable diffs and reduce maintenance when navigation or surrounding content changes.

2. Establish and update baselines deliberately

Generate a baseline with:

npx playwright test --update-snapshots

Use this command only for an intentional interface change. Review the resulting image files in the same code review as the product change. A snapshot update should explain what moved, changed color, changed typography or was added or removed.

To update one test while investigating:

npx playwright test tests/home.spec.ts --update-snapshots

After updating, run the test again without --update-snapshots. This confirms that the new reference is accepted by the normal comparison path.

3. Configure the comparison

Screenshot options can be set per assertion or in the expect.toHaveScreenshot section of playwright.config.ts:

import { defineConfig, expect } from '@playwright/test';

export default defineConfig({
  expect: {
    toHaveScreenshot: {
      animations: 'disabled',
      caret: 'hide',
      scale: 'css',
      threshold: 0.2,
      maxDiffPixels: 100,
    },
  },
});
Option What it controls Use it when
fullPage Captures the full scrollable page instead of the viewport. The page layout below the fold is part of the contract.
animations Disables or allows CSS and web animations during capture. Disable animations unless animation itself is under test.
caret Controls whether a text caret is visible. Inputs appear in the region being compared.
mask Masks one or more locators with a solid color. Content is intentionally volatile but its layout matters.
stylePath Applies a stylesheet during capture. A project needs a shared way to hide or freeze dynamic regions.
threshold Per-pixel perceived color difference accepted by the comparator. Small rendering variation is expected. The documented pixelmatch default is 0.2; zero is strict and one is lax.
maxDiffPixels Maximum number of differing pixels. You can define an absolute visual-change budget.
maxDiffPixelRatio Maximum differing-pixel ratio relative to image area. The same percentage is meaningful across different image sizes.
scale Uses CSS pixels or device pixels. Choose one value consistently across baseline and comparison jobs.

The total-difference limits are unset unless you configure them. Set the smallest tolerance that accommodates a known source of rendering variation. A larger tolerance can hide a real regression, so do not use it as a substitute for investigating unstable content. See the SnapshotAssertions API for the current option set.

Mask volatile elements

test('dashboard layout is stable', async ({ page }) => {
  await page.goto('http://localhost:3000/dashboard');
  await expect(page).toHaveScreenshot('dashboard.png', {
    fullPage: true,
    mask: [
      page.locator('[data-testid="current-time"]'),
      page.locator('[data-testid="avatar"]'),
    ],
    maskColor: '#FF00FF',
  });
});

Masking hides changing pixels while still checking the surrounding layout. If the changing value affects layout, freeze the data instead of masking it.

Apply a capture-only stylesheet

/* tests/visual-freeze.css */
[data-testid="live-clock"],
[data-testid="rotating-promo"] {
  visibility: hidden !important;
}

*, *::before, *::after {
  animation-duration: 0s !important;
  animation-delay: 0s !important;
  transition: none !important;
}
test('account page is stable', async ({ page }) => {
  await page.goto('http://localhost:3000/account');
  await expect(page).toHaveScreenshot('account.png', {
    stylePath: 'tests/visual-freeze.css',
  });
});

4. Remove the causes of flaky screenshots

Screenshot comparison is only as reliable as the state being captured. Before the assertion:

  • Seed the database or mock API responses so the same records appear.
  • Freeze dates, clocks and randomized identifiers.
  • Wait for the specific content needed by the assertion, not an arbitrary sleep.
  • Move the pointer away from controls that change on hover.
  • Close menus, dialogs and focus rings unless they are part of the scenario.
  • Wait for fonts and images that affect layout.
  • Use one browser version, operating system image and headless mode for baselines and comparisons where possible.
test('search results are deterministic', async ({ page }) => {
  await page.route('**/api/search**', async route => {
    await route.fulfill({
      status: 200,
      contentType: 'application/json',
      body: JSON.stringify({
        results: [
          { id: 1, title: 'Alpha' },
          { id: 2, title: 'Beta' },
        ],
      }),
    });
  });

  await page.goto('http://localhost:3000/search?q=example');
  await page.locator('[data-testid="result-list"]').waitFor();
  await page.mouse.move(0, 0);
  await expect(page).toHaveScreenshot('search-results.png');
});

Operating system fonts, browser versions, hardware, power settings and headless mode can change rendered pixels. Keep those inputs consistent in CI. If developers generate references on laptops but CI compares them on a different image, expect avoidable diffs.

5. Diagnose a failed comparison

When an assertion fails, inspect all three artifacts:

  • Actual: what the test captured.
  • Expected: the committed baseline.
  • Diff: where the pixels differ.

Classify the failure before changing code or tolerance:

Symptom Likely cause Fix
The entire page shifts Different viewport, device scale factor, font or browser. Pin the project and CI environment; verify viewport and scale.
A clock, avatar or ad changes Volatile data or third-party content. Mock or freeze it, mask it, or hide it with stylePath.
Only hover styles differ Pointer remains over an interactive element. Move the mouse away before the assertion.
Text wraps differently Font not loaded, different font fallback or changed width. Wait for the page’s fonts and use the same font files and viewport.
Images are missing Capture starts before images load or the network response is unstable. Wait for the relevant locator or image state and control the response.
Only a moving animation differs Animation or transition is still active. Use the default animation disabling, or apply a capture stylesheet.
Failures occur intermittently Race between application state and capture. Wait for a meaningful state, remove arbitrary sleeps and inspect traces.
Many tiny anti-aliased differences Rendering environment variation. Align environments first; use a narrow threshold only when the variation is understood.

6. Tune tolerances without hiding regressions

threshold applies to each pixel’s perceived color difference in YIQ color space. maxDiffPixels limits the absolute number of changed pixels, while maxDiffPixelRatio limits the proportion of the image that may differ.

test('marketing hero allows a known one-pixel edge variation', async ({ page }) => {
  await page.goto('http://localhost:3000/');
  await expect(page).toHaveScreenshot('hero.png', {
    threshold: 0.1,
    maxDiffPixels: 50,
  });
});

Start strict. If a diff is legitimate, identify why it occurs, make the input deterministic, and then add the smallest documented allowance. A broad ratio on a full-page image can permit a large visible change, so prefer a focused locator or a small absolute budget when practical.

7. Organize snapshots in CI

  1. Run visual tests in a pinned browser and operating-system image.
  2. Upload actual, expected and diff artifacts when a job fails.
  3. Make visual changes in the same pull request as the code that causes them.
  4. Require a reviewer to inspect changed baselines.
  5. Run the normal comparison command after any approved update.

Keep visual tests separate from unit tests when that makes artifact review and retry policy clearer. Retries can help diagnose infrastructure noise, but they should not automatically approve a changed baseline.

8. Local snapshots or hosted review

Playwright’s built-in snapshots are a good fit when your team can keep rendering environments consistent and wants image files reviewed in the repository. A hosted workflow can be useful when you want a service to maintain base builds and present visual changes in a review interface. BrowserStack documents a Percy integration that routes existing Playwright screenshot assertions to Percy; check the vendor guide for current package and runtime requirements before pinning versions: Percy with Playwright.

Choose explicitly whether a visual difference should fail the job immediately or enter an approval queue. That decision affects CI behavior, access controls and baseline administration.

9. Or skip the browser setup

If you need rendered images for documentation, monitoring, previews or a separate visual-review system, ScreenshotNeo provides a website screenshot API. It accepts one GET request and returns PNG, JPEG, WebP or PDF. The API can load lazy images, capture a CSS-selected element, set a viewport or device preset, use dark mode and retina scale, apply custom CSS or JavaScript, click before capture, wait for a selector, delay or network idle, block ads or resource types, send headers and cookies, set timezone or geolocation, resize images, cache with a chosen TTL, create signed image links, run asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call and expose usage data. See the ScreenshotNeo documentation.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const fs = require('node:fs/promises');

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo removes cookie and consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, failed loads and cache hits are not billed. Each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

10. Performance, reliability and cost

Performance

  • Prefer locator screenshots over full-page captures when the test only needs one component.
  • Keep test data and network mocks local where possible.
  • Run independent pages in parallel, while avoiding shared mutable state.
  • Use full-page captures selectively because taller images take longer to render, compare and review.

Reliability

  • Pin browser binaries and CI images.
  • Control fonts, locale, timezone and viewport.
  • Keep third-party content out of visual contracts or replace it with deterministic fixtures.
  • Inspect artifacts before changing tolerances.

Cost

Playwright’s local comparison uses your existing test infrastructure; the main costs are CI time, artifact storage and review effort. Hosted visual review adds a service and its administration, so evaluate how many builds, browsers and reviewers your team needs. For on-demand rendered screenshots, ScreenshotNeo bills only clean shots; failed loads, bot checks, blank pages and cache hits cost nothing. Its plans are Free: 1,000/month, Starter: $5 for 3,000, Growth: $15 for 15,000, Pro: $39 for 60,000, Scale: $99 for 250,000 and Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan.

11. Checklist for a maintainable Playwright image comparison suite

  • Use toHaveScreenshot() through Playwright Test.
  • Commit snapshots and review updates with code changes.
  • Scope assertions to locators where full-page coverage is unnecessary.
  • Freeze data, time, fonts, network responses and interaction state.
  • Disable animations and handle hover, focus and caret state.
  • Mask or hide only known volatile regions.
  • Pin browser and operating-system environments.
  • Set narrow, explained tolerance limits.
  • Preserve actual, expected and diff artifacts in CI.
  • Re-run without snapshot updates after approving a change.

FAQ

Does Playwright compare screenshots automatically?

Yes, when you call expect(page).toHaveScreenshot() or the locator equivalent in Playwright Test. The first run creates a baseline and later runs compare against it.

Why does Playwright capture twice?

The assertion waits for two consecutive screenshots to be identical before comparing the final capture. This helps settle layout and animation-related changes.

Should I use a full-page screenshot for every test?

No. Use a locator screenshot for a component or focused region, and reserve full-page captures for layouts where below-the-fold content matters.

When should I update snapshots?

Update them after reviewing an intentional UI change and its diff. Run npx playwright test --update-snapshots, commit the new files, then run the ordinary test command again.

What is the safest way to handle timestamps?

Freeze or mock the time when possible. Masking is appropriate only when the changing value does not affect layout or behavior under test.

Can ScreenshotNeo replace Playwright visual assertions?

It serves a different workflow: it renders URLs through an API and can clean pages before capture. Use Playwright for assertions inside your browser tests; use ScreenshotNeo when you need a screenshot service, PDFs, bulk captures or MCP tools.