Playwright Screenshot Testing: Visual Regression That Stays Reliable
Learn how to build reliable Playwright screenshot tests, manage baselines, tune diffs, stabilize CI, and debug visual failures.
Playwright screenshot testing compares a page or component against a reviewed reference image. Use expect(page).toHaveScreenshot() for a page and expect(locator).toHaveScreenshot() for a specific region. The first run creates the baseline; later runs fail when the rendered pixels differ beyond your configured tolerance.
For reliable results, pin the browser and operating-system environment, keep dynamic content deterministic, leave animations disabled, review snapshot changes in version control, and use narrow tolerances. The complete workflow is below.
1. Install Playwright Test
npm init playwright@latest
If Playwright is already installed, verify the test runner and browsers:
npm install -D @playwright/test
npx playwright install
The examples use TypeScript, but the same test APIs work in JavaScript.
2. Write a full-page screenshot assertion
import { test, expect } from '@playwright/test';
test('landing page visual baseline', async ({ page }) => {
await page.goto('https://example.com');
await expect(page).toHaveScreenshot('landing-page.png');
});
Run the test once:
npx playwright test
When no reference exists, Playwright reports the missing snapshot and writes the captured image as the baseline. Run the test again to compare against that file. Snapshot paths are created under the test’s snapshot directory and should be committed with the test code.
3. Test one component with a locator
Use a locator assertion when the visual contract is a component. This reduces unrelated differences elsewhere on the page.
import { test, expect } from '@playwright/test';
test('header visual baseline', async ({ page }) => {
await page.goto('https://example.com');
await expect(page.getByRole('banner')).toHaveScreenshot('header.png');
});
Other useful locators include page.locator('[data-testid="checkout-summary"]'), page.getByRole('dialog'), and a CSS selector for a stable component boundary.
4. Create and update baselines safely
Review the first baseline before committing it. If a UI change is intentional, update snapshots explicitly:
npx playwright test --update-snapshots
Review every changed image, then commit the snapshot directory with the corresponding code change. Do not update snapshots automatically in a normal CI run: an accidental layout change would become the new expected result without review.
5. Configure screenshot behavior
Screenshot assertions wait for two consecutive screenshots to be identical before comparing them. This removes many capture-time races. By default, CSS animations and Web Animations are disabled for the assertion; finite animations are fast-forwarded and infinite animations are canceled.
Full-page capture
await expect(page).toHaveScreenshot('full-page.png', {
fullPage: true
});
Mask dynamic regions
Mask timestamps, rotating promotions, avatars, or user-specific values that are outside the purpose of the assertion:
await expect(page).toHaveScreenshot('account.png', {
mask: [
page.locator('[data-testid="current-time"]'),
page.locator('[data-testid="recommendations"]')
]
});
Masking is a test decision, not a substitute for deterministic fixtures. Keep the data stable whenever possible.
Control animation and hover state
await page.mouse.move(0, 0);
await expect(page.locator('.menu')).toHaveScreenshot('menu.png');
Moving the pointer away prevents accidental hover styles from becoming part of the baseline. Keep the default animation disabling unless the test specifically verifies an animation frame.
Set a custom screenshot timeout
await expect(page).toHaveScreenshot('slow-page.png', {
timeout: 10000
});
The documented project configuration default for the expect timeout is 5,000 ms. Increase it only when the page genuinely needs more time to settle.
6. Choose diff strictness deliberately
Playwright exposes three controls:
| Option | What it limits | Use it when |
|---|---|---|
threshold |
Per-pixel color difference (pixelmatch tolerance) | Minor anti-aliasing or color variation is expected |
maxDiffPixels |
Absolute number of different pixels | A fixed amount of noise is acceptable |
maxDiffPixelRatio |
Different pixels as a ratio of the image | The image size varies and a proportional limit is more meaningful |
await expect(page).toHaveScreenshot('dashboard.png', {
threshold: 0.2,
maxDiffPixels: 100,
maxDiffPixelRatio: 0.001
});
Pixelmatch’s documented default threshold is 0.2. Set only the controls you need: combining generous values can hide a real regression. Treat any tolerance increase as a reviewed policy change.
Set project-wide defaults
import { defineConfig } from '@playwright/test';
export default defineConfig({
expect: {
toHaveScreenshot: {
animations: 'disabled',
threshold: 0.2,
maxDiffPixelRatio: 0.001
}
}
});
7. Keep rendering deterministic in CI
Visual baselines are tied to rendering conditions. Use the same operating-system and browser versions for baseline creation and comparison. Host settings, browser version, hardware, power source, and headless mode can all influence output. A practical CI setup pins the Playwright version, installs the same browser revision, and runs screenshots in one known image.
Also control:
- Fonts installed in the runner.
- Timezone and locale.
- Viewport and device scale factor.
- Seed data and authenticated account state.
- Network responses for timestamps, ads, rotating content, and experiments.
- Hover, focus, caret, and scroll position.
Do not compare a macOS-created baseline with a Linux CI capture unless you have accepted and reviewed the rendering differences.
8. A complete example with setup and stable data
import { test, expect } from '@playwright/test';
test.beforeEach(async ({ page }) => {
await page.addInitScript(() => {
Date.now = () => 1700000000000;
});
await page.route('**/api/recommendations', async route => {
await route.fulfill({
status: 200,
contentType: 'application/json',
body: JSON.stringify({ items: ['Plan A', 'Plan B'] })
});
});
});
test('checkout summary is stable', async ({ page }) => {
await page.goto('https://example.com/checkout');
await page.mouse.move(0, 0);
await expect(page.locator('[data-testid="checkout-summary"]'))
.toHaveScreenshot('checkout-summary.png', {
mask: [page.locator('[data-testid="user-avatar"]')],
maxDiffPixels: 50
});
});
9. Diagnose a visual failure
- Open the expected, actual, and diff images produced by the failed test.
- Decide whether the change is intentional, environmental, or flaky.
- For an intentional UI change, update the baseline after review.
- For a rendering issue, fix the environment or test setup before changing tolerances.
- For a timing issue, inspect the page state and wait for a meaningful condition rather than adding arbitrary delays.
Use Playwright Trace Viewer for CI diagnosis. A trace provides a test timeline and DOM snapshots, which helps identify whether the page captured before data loaded, after an unexpected navigation, or with a different state. Tracing every test can be expensive; configure it for retries or targeted runs.
10. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Snapshot does not exist | First run or wrong snapshot path | Run once to create it, verify the test name and snapshot directory, then commit the file. |
| Large diff after a browser update | Different browser revision or rendering engine | Pin the Playwright version and browser installation; regenerate baselines only after review. |
| Text differs between runs | Fonts, locale, timezone, timestamps, or data changed | Install the same fonts, set locale/timezone, freeze or mock time, and use deterministic fixtures. |
| Only hover styles differ | Mouse remained over an interactive element | Move the pointer away before the assertion or assert the intended hover state explicitly. |
| Animated region produces intermittent diffs | Animation or transition still changes during capture | Keep animations: 'disabled'; mask the region if it is not under test. |
| Screenshot times out | Page or locator never reaches a stable state | Check navigation and selector waits, remove a hanging resource, or increase timeout for a known slow page. |
| Full-page image is unexpectedly tall | Lazy content loads while scrolling or the page has expanding regions | Wait for required content, stabilize layout, and use a locator assertion when the page height is not the contract. |
| CI fails but local passes | Different OS, browser, fonts, hardware, or headless mode | Run both in the same pinned environment and inspect the trace and diff. |
11. Performance and reliability guidance
- Prefer locator screenshots for component tests; they produce smaller images and isolate failures.
- Use full-page assertions for page composition, route-level smoke coverage, and important responsive layouts.
- Keep snapshot suites focused. A screenshot test still pays for navigation, page setup, and image comparison.
- Reuse authenticated storage state where appropriate instead of logging in through the UI for every test.
- Run tests in parallel only when fixtures and test data are isolated; shared mutable data can create visual races.
- Capture traces on retry or targeted failures rather than every passing test.
- Review image diffs in pull requests so baseline updates remain auditable.
12. When to use toMatchSnapshot()
Playwright also documents comparing await page.screenshot() with toMatchSnapshot(). For screenshot assertions, the snapshot reference recommends toHaveScreenshot(), which includes screenshot-specific waiting and configuration. Use toMatchSnapshot() for non-image values or a deliberate lower-level workflow.
13. Or skip the browser setup
If you need screenshots for documentation, monitoring, previews, or batch jobs rather than an in-process browser test, ScreenshotNeo provides a single HTTP request. The API accepts the URL and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', body));
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account and get 1,000 screenshots a month without a card.
FAQ
Should I test the whole page or a component?
Use a page assertion when layout and composition are the contract. Use a locator assertion when the component is the contract and unrelated page changes should not fail the test.
How often should baselines be regenerated?
Only after reviewing an intentional visual change or a deliberate environment migration. Do not regenerate snapshots as an automatic response to CI failures.
What tolerance should I start with?
Start with Playwright’s defaults and add the smallest necessary threshold, maxDiffPixels, or maxDiffPixelRatio value after understanding the source of the variation.
Why are screenshots still flaky when animations are disabled?
Other causes include changing data, fonts, browser or OS versions, hover state, lazy content, and network responses. Use deterministic fixtures and a pinned rendering environment before relaxing comparison limits.


