ScreenshotNeo

BlogHow-to

How to use Playwright to compare website screenshots in CI

Build reliable visual regression checks with Playwright Test: create reviewed baselines, reduce screenshot noise, and keep CI comparisons reproducible.

By the ScreenshotNeo team4 October 20269 min read

Use Playwright Test’s expect(page).toHaveScreenshot() assertion to compare a rendered page with a reviewed reference image. The first run creates the reference; later runs compare against it. Reliable results depend on generating and comparing snapshots in matching environments: keep the operating system, Playwright and browser versions, viewport, scale, fonts, and page state consistent between baseline creation and CI. A difference can indicate an application change, but first check for environment and content variation.

This guide shows how to add the assertion, generate and commit a baseline, run it in CI, control known volatility, diagnose failures, and update references safely. Examples use TypeScript and Playwright Test.

1. Add a visual regression assertion

Start with an ordinary browser test that reaches the state you want to protect. Assert that meaningful content is ready before capturing the page. Playwright’s screenshot assertion waits until two consecutive screenshots are identical before comparing with the reference image.

import { test, expect } from '@playwright/test';

test('home page visual baseline', async ({ page }) => {
  await page.goto('/');
  await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
  await expect(page).toHaveScreenshot('home.png');
});

Save this as a Playwright test file, for example tests/home.visual.spec.ts. The URL / is resolved using the test configuration’s baseURL, if one is set. Otherwise use a complete URL in page.goto().

toHaveScreenshot() is a Playwright Test runner assertion. Use it for page screenshot comparisons instead of passing a screenshot buffer to toMatchSnapshot(); the snapshot assertion documentation cautions against that approach for screenshots.

2. Create and review the reference image

Run the test once in the environment you intend to use for baseline generation:

npx playwright test tests/home.visual.spec.ts

When no expected screenshot exists, Playwright writes a reference image. Inspect the generated file at the reported snapshot path before accepting it. The test file name contributes to the generated snapshot directory. Once reviewed, add the reference image to version control so later local and CI runs have a stable comparison target.

Do not treat a generated image as automatically correct just because the test created it. Check that it shows the intended page, viewport, content, and state. If your repository needs a different snapshot layout, configure the snapshot path template using the Playwright Test configuration supported by your installed version.

3. Keep baseline creation and CI reproducible

Use the same material screenshot environment when creating references and running comparisons. Playwright documents that output can vary with host operating system, browser version, settings, hardware, power source, and headless mode. Its visual comparison guide advises: “For consistent screenshots, run tests in the same environment where the baseline screenshots were generated.”

  • Pin @playwright/test through your package manifest and lockfile.
  • Install the browser binaries that match that Playwright version in CI.
  • Keep the operating system, browser project, viewport, device scale factor, color scheme, locale, timezone, and relevant browser settings aligned.
  • Make the page state deterministic: use stable test data, wait for meaningful content, and avoid time-dependent or randomly generated content.
  • Review CI-generated diffs before changing a baseline.

A minimal configuration might look like this:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  use: {
    baseURL: 'http://127.0.0.1:3000',
  },
  webServer: {
    command: 'npm run start',
    url: 'http://127.0.0.1:3000',
    reuseExistingServer: !process.env.CI,
  },
});

Adjust the server command and URL for your application. In CI, install dependencies and the matching browser before running the suite. For example, a Linux Chromium job can use:

npm ci
npx playwright install --with-deps chromium
npx playwright test

Playwright documents npx playwright install for browser binaries and npx playwright install --with-deps chromium for Chromium plus required system dependencies on supported Linux environments. If your project tests multiple browsers, install the browsers required by those projects. Linux headed browser tests need Xvfb; Playwright’s default browser execution is headless.

Browser binary caching is generally not recommended in Playwright’s CI guidance: restoring the cache can take about as long as downloading browsers, and Linux system dependencies cannot be cached this way. If you do cache browser binaries, key the cache to the Playwright version so a package update does not reuse mismatched binaries.

4. Make multi-project comparisons intentional

If you deliberately run projects on different browsers or operating systems, their rendered output may differ. Give each materially distinct project its own appropriate expectation set rather than comparing its output with an unrelated environment’s baseline. For example, Chromium and WebKit runs may need separate references. Playwright’s WebKit is derived from WebKit main branch sources and is not branded Safari; for the closest Safari experience, Playwright advises running WebKit on macOS. Platform-sensitive behavior, including codecs, can differ.

Do not assume that Linux WebKit screenshots are identical to screenshots from Safari on macOS. Decide which platform and browser combinations are part of the behavior you want to guard, then create and review references in those same projects.

5. Reduce noise without hiding meaningful changes

Screenshot assertions disable animations by default. For other known volatile regions, you can provide a stylesheet with stylePath to neutralize or hide them. Playwright documents that the stylesheet can affect content inside frames and Shadow DOM as well.

Example stylesheet, saved as tests/visual-stability.css:

/* Mask only the timestamp documented as volatile for this test. */
.test-timestamp {
  visibility: hidden !important;
}

Then configure the assertion:

import { test, expect } from '@playwright/test';

test('home page visual baseline', async ({ page }) => {
  await page.goto('/');
  await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
  await expect(page).toHaveScreenshot('home.png', {
    stylePath: 'tests/visual-stability.css',
  });
});

Keep exclusions narrow and explain why each excluded element is unstable. Hiding a whole header, navigation area, or other user-visible region can make the test pass while real visual regressions go unnoticed. If dynamic data can be controlled through test fixtures or application state, stabilizing that state is often preferable to masking its rendered result.

6. Choose screenshot format and difference tolerances

Screenshot assertions use PNG by default. You can use a .webp extension for lossless WebP. PNG is the straightforward default; WebP can be an alternative for reference storage if your team is comfortable reviewing that format. Keep the format consistent within a reference set.

Playwright’s threshold sets the acceptable perceived per-pixel color difference in YIQ space; the documented default is 0.2. maxDiffPixels and maxDiffPixelRatio can cap the number or ratio of changed pixels. Those budgets are unset unless configured. Start with the defaults and change them only after reviewing real diffs and agreeing on a policy.

import { defineConfig } from '@playwright/test';

export default defineConfig({
  expect: {
    toHaveScreenshot: {
      threshold: 0.2,
      maxDiffPixels: 10,
    },
  },
});

The value 10 above is an illustrative team policy, not a Playwright recommendation. You can also configure screenshot comparison options per project or per assertion. Prefer a narrowly justified tolerance to a large global budget: a loose threshold may hide a visible change. When a test fails, inspect the actual image and diff before deciding that a tolerance should change.

7. Update a baseline deliberately

When a UI change is intentional, update the reference explicitly:

npx playwright test --update-snapshots

Inspect the updated image and its diff, then commit the accepted reference change with the application change. Do not run snapshot updates automatically after ordinary CI failures. Automatically replacing the reference would erase the comparison that exposes unexpected changes.

8. Troubleshoot failed comparisons

Symptom Likely cause What to check or change
A screenshot is written instead of compared No reference image exists for this test and project. Review the newly generated reference, then add it to version control. Confirm CI checks out the committed snapshot directory.
CI reports a large diff immediately after a dependency update The Playwright version or its browser binary changed. Check the lockfile and installed browser version. Install browser binaries using the updated Playwright package, then regenerate references only if the reviewed rendering change is intentional.
The same commit passes locally but fails in CI The OS, browser, fonts, rendering settings, headless mode, hardware, or power conditions differ. Compare environments and run baseline creation in the CI-matching environment. Check the actual screenshot and diff before updating.
Only dates, rotating content, or personalized areas differ The test page contains unstable data or state. Use stable test data and assert the intended state. If a region is genuinely volatile and irrelevant to the assertion, use a narrowly scoped stylePath rule and document it.
A screenshot differs around a font or text wrap Font availability, OS, browser build, viewport, or scale may differ. Align the environment and screenshot settings. Verify font loading and the rendered page state before changing tolerances.
A browser fails to launch in Linux CI Required browser binaries or system dependencies are missing. Install the matching browser and dependencies, for example with npx playwright install --with-deps chromium. For headed Linux runs, provide Xvfb; headless is the default.
WebKit output does not match Safari Playwright WebKit is not branded Safari and platform behavior can differ. Use WebKit on macOS when a closer Safari experience is needed, and maintain platform-appropriate references.
A tolerance change makes failures disappear, but a visual issue remains The accepted difference budget is too broad or masks a meaningful area. Review the diff, narrow the budget, or remove it. Prefer stabilizing the source of noise to allowing more pixels to differ.

9. Or skip the browser setup

If you need a website image without maintaining a Playwright browser job and committed visual baselines, ScreenshotNeo is a website screenshot API and MCP server. Its API accepts a URL in one GET request and can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
  writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))
);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, with no card required.

10. Performance, reliability, and cost considerations

A Playwright visual test runs a browser, loads your application, and captures a page. Its runtime therefore includes application startup and page loading as well as screenshot comparison. Keep the test focused on the state and pages whose visual behavior matters, and avoid repeatedly capturing unchanged screens without a reason. The research sources do not provide a general runtime benchmark; measure your own suite in its CI environment.

For reliability, use a pinned dependency lockfile, install matching browser binaries, keep the page state deterministic, and make baseline updates reviewed changes. Browser caching may not reduce CI time once restore cost is considered, and it does not replace installing Linux system dependencies. For cost planning, account for CI minutes, browser installation, and snapshot storage in your own setup; exact costs depend on your runner and repository. Playwright’s cited documentation does not establish universal cost or performance figures.

Playwright visual assertions and a screenshot API address related but different jobs. A committed baseline assertion detects whether a page changed relative to a reviewed image. A URL-to-image API is useful when you need captures without owning that baseline workflow. ScreenshotNeo’s per-plan pricing and billing behavior are described above; use the Playwright workflow when the comparison itself is the requirement.

FAQ

Does toHaveScreenshot() compare the very first run against anything?

No reference exists on the first run, so Playwright creates one. Review it and commit it; later runs compare against that reference.

Should visual baselines be committed?

Yes. They are the reviewed reference used by future comparisons, so keep them in version control alongside the test.

Can I use this approach with multiple browsers?

Yes. Define the browser projects you need and use appropriate references for materially different browser or operating system environments.

Does disabling animation eliminate every screenshot difference?

No. It addresses animation, while environment differences and dynamic page content can still change output.

Is the documented default threshold a ten-pixel budget?

No. The documented default threshold is 0.2 in YIQ color-difference space. Pixel-count and ratio budgets are separate options and are unset unless configured.

Official Playwright references