ScreenshotNeo

BlogHow-to

How to Compare Chromium Screenshots for Visual Regression Testing

Build stable Chromium screenshot tests with Playwright, reviewed baselines, useful diff tolerances, and practical ways to reduce rendering noise.

By the ScreenshotNeo team4 October 202610 min read

For visual regression tests of a web application in Chromium, use Playwright Test’s toHaveScreenshot() assertion. It creates a reference screenshot on the first run and compares later captures against that reviewed baseline. Keep baseline generation and comparison in the same operating system, browser version, fonts, viewport, and rendering mode; otherwise environmental differences can look like regressions. For Chrome’s own desktop interface, use Chromium Pixel Tests with Skia Gold instead.

A visual test answers “did this output change from the approved image?” It does not prove that the page is universally correct. Review the diff, decide whether the change is intended, and update the baseline only after that review.

1. Choose the right Chromium screenshot workflow

What you are testing Recommended approach Baseline and review
Your web application or component in Chromium Playwright Test with toHaveScreenshot() Snapshot files live with the tests; review diffs in your development or CI workflow.
Chrome browser UI maintained in the Chromium project Chromium Pixel Tests, using Kombucha Screenshot or ScreenshotSurface verbs for new tests Approved images are compared through Skia Gold; inspect failures before accepting or rejecting changes.
Hosted review, centralized approvals, or rendering beyond one controlled Chromium environment Consider a hosted visual testing workflow after deciding which browsers, widths, and approval process you need Percy documents Playwright integration; Chromatic documents a Playwright upload and cloud review workflow.

For an application already tested with Playwright, start with its built-in assertion: it needs no separate visual testing service. Chromium Pixel Tests are intended for Chrome’s own browser UI, not as a drop-in alternative for ordinary app tests. Chromium’s guide prefers Kombucha screenshot verbs for new pixel tests and describes TestBrowserUi as legacy support. See the Chromium Pixel Tests guide and the Playwright visual comparisons guide.

2. Add a Playwright screenshot assertion

Install Playwright Test and its Chromium browser if they are not already part of the project:

npm init playwright@latest
npx playwright install chromium

Create a test such as tests/visual.spec.ts:

import { test, expect } from '@playwright/test';

test('pricing page appearance', async ({ page }) => {
  await page.setViewportSize({ width: 1280, height: 800 });
  await page.goto('https://example.test/pricing');
  await expect(page).toHaveScreenshot('pricing-page.png');
});

Run it with npx playwright test. On the initial run, Playwright creates the expected screenshot. Inspect and commit that file as the reference. Later runs capture the page again and compare the new image to the reference. When the comparison fails, Playwright reports a diff for review.

Playwright’s screenshot assertion waits until two consecutive screenshots match, then compares the last capture with the expected image. Screenshot assertions disable animations and hide the caret by default. These defaults reduce some noise, but they do not make a page deterministic by themselves. Consult the PageAssertions API for current assertion options.

Make the first baseline intentional

  1. Use a stable test account and fixed data. Avoid depending on live records that change between runs.
  2. Choose a fixed viewport and use the same browser project in baseline creation and CI.
  3. Wait for the state under test: for example, a loaded route and visible component, rather than an arbitrary pause.
  4. Run the test once to create the snapshot, then inspect the image at normal size before committing it.
  5. Run it again in the intended CI environment and confirm it compares successfully there.

The first-run image is only a reference; it is not automatically an approved design. A bad or partially loaded page can become a bad baseline.

3. Stabilize content and rendering before comparing

Browser rendering can vary with host operating system, browser version, settings, hardware, power source, headless mode, fonts, and other conditions. Playwright specifically advises creating and comparing snapshots in the same environment. Pin the browser version through the project’s Playwright setup, keep the CI image stable, and avoid generating baselines on one platform and comparing them on another unless platform variation is an intentional part of the test.

  • Data: Seed fixtures or use fixed responses for changing API data, user names, prices, and counts.
  • Fonts and assets: Ensure local fonts and images have loaded before capture. A fallback font can shift line wrapping and move large parts of the page.
  • Time and randomness: Freeze dates or provide deterministic values when timestamps, rotating content, or random identifiers are visible.
  • Animation: Screenshot assertions disable animations by default. For application-specific transitions or videos, set the page to a stable state and verify the resulting capture.
  • Network: Wait for the relevant page state or element, not merely for navigation to begin. Avoid making a test depend on an unrelated third-party request.
  • Environment: Keep OS/container image, fonts, viewport, browser settings, and headless mode consistent. If testing multiple OS environments, maintain and review the corresponding baselines separately.

Prefer fixing the source of instability over raising the allowed difference. Playwright’s visual comparison documentation explains why host differences matter: Visual comparisons.

4. Capture only the region that answers the test

A full-page image is useful when the page’s overall composition is the subject. For a focused behavior, capture the component or region whose appearance matters. Smaller, relevant captures are less likely to fail because unrelated content changed.

// Capture a component instead of the entire page.
await expect(page.locator('[data-testid="checkout-summary"]'))
  .toHaveScreenshot('checkout-summary.png');

For Chromium browser UI, the Pixel Tests guide distinguishes Screenshot, which captures the named UI element, from ScreenshotSurface, which captures the containing dialog or window. Avoid including neighboring elements likely to change independently. Chromium’s Pixel Tests documentation describes capture and Skia Gold review.

5. Set diff tolerance with reviewed evidence

Playwright supports screenshot comparison options such as maxDiffPixels. Use a tolerance only after you have a stable capture and have inspected the differing pixels. There is no universally correct threshold: an acceptable number depends on the image size, what the test covers, and which visual changes matter to the product.

await expect(page).toHaveScreenshot('pricing-page.png', {
  maxDiffPixels: 25,
});

Start with strict comparison. If small, irreducible rendering variation remains in a controlled environment, choose the smallest threshold that accommodates it and document why. A broad tolerance can conceal a genuine layout or color regression. See the PageAssertions API for comparison options, including pixel and ratio-based difference settings.

Filter genuinely volatile areas

If a timestamp or rotating element is not part of the behavior being tested, remove or stabilize it for the capture. Playwright’s stylePath option applies a stylesheet during screenshot comparison:

await expect(page).toHaveScreenshot('dashboard.png', {
  stylePath: './tests/visual-snapshot.css',
});
/* tests/visual-snapshot.css */
.test-only-clock,
[data-visual-volatile] {
  visibility: hidden !important;
}

Use this narrowly. Hiding a region that is part of the feature under test can make the screenshot pass while the user-visible output is broken. The assertion API documents stylePath and other options: PageAssertions.

6. Review failures and update snapshots deliberately

  1. Open the expected image, actual image, and diff. Identify the changed region.
  2. Check whether the page state, test data, fonts, browser build, or environment changed unexpectedly.
  3. Classify the difference as an intended product change, a real regression, or capture noise.
  4. Fix the page or test when the change is a regression or noise. Do not approve noise by increasing tolerance without understanding it.
  5. For an intended visual change, update snapshots with npx playwright test --update-snapshots, inspect the new files, and include them in the same review as the code change.

Chromium Pixel Tests similarly expect failed images to be reviewed through Skia Gold before accepting or rejecting an image change. Chromium web tests also use PNG baselines and can compare reference output. See Web Test Expectations and Baselines.

7. Chromium Pixel Tests for Chrome UI

If the target is Chrome’s own desktop UI, follow the Chromium project’s Pixel Tests infrastructure rather than adapting an app-level Playwright test. The current guide describes approved screenshot images and Skia Gold review. For new tests, it prefers Kombucha’s Screenshot and ScreenshotSurface verbs. Use the element capture when a specific control is under test, and a surface capture when the dialog or window as a whole is the intended boundary.

Keep the captured surface focused: Chromium cautions against including other elements likely to change on their own. On failure, inspect the actual output and expected image in the review workflow before accepting a new baseline. Follow the project’s current build and test instructions in the Chromium Pixel Tests documentation; the exact test setup depends on the Chrome UI code being changed.

8. Hosted visual testing when local snapshots are not enough

A local Playwright baseline is a strong fit when the team wants Chromium comparisons in its existing test suite and wants image changes reviewed alongside code. A hosted service may fit when the team needs centralized review, broader browser or responsive rendering coverage, or a shared approval workflow.

BrowserStack documents Percy’s visual comparison and Playwright integration. Chromatic documents a Playwright integration that captures page archives during test runs, uploads them to its cloud, creates snapshots, and performs pixel diffing; its visual testing guide describes cloud snapshots and review. These are optional workflows, not prerequisites for Playwright screenshot assertions: Chromatic Playwright integration and Chromatic visual tests.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It is useful when you need a rendered page image without maintaining browser capture code. The one-call request returns an image or PDF; see the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.test/pricing \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.test/pricing"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.test/pricing',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month, with no card required.

10. Troubleshooting

Symptom Likely cause What to do
Many pixels differ on every run Unstable data, animation, late assets, or inconsistent rendering environment. Fix test data, wait for the intended state and loaded assets, then keep OS, browser, fonts, viewport, and headless mode consistent.
Text wraps or shifts between machines Different fonts, font loading, OS rendering, or viewport dimensions. Use the same CI image and fonts, wait for the page state that includes the font, and use a fixed viewport. Maintain separate baselines if testing platforms intentionally.
Snapshot differs after a browser or dependency update The rendering environment changed, or the update changed page output. Inspect the diff. Decide whether this is expected; update baselines only after review.
Snapshot includes a loading state or incomplete image The test captured before the relevant content was ready. Wait for a meaningful locator or application state before the assertion; avoid relying on a fixed sleep when a state check is possible.
One irrelevant timestamp causes repeated failures Volatile content is included in the comparison. Freeze the value or narrowly filter it with stylePath if it is outside the test’s purpose.
Small threshold makes real changes pass The allowed difference is too broad for the capture boundary. Reduce or remove the tolerance, narrow the screenshot to the behavior under test, and review the actual diff.
Snapshot update changes many files The environment or page output changed broadly, or an update was run too widely. Review the set of changed images, confirm the cause, and update only the intended tests where possible.

11. Performance, reliability, and maintenance

  • Runtime: Screenshot capture and image comparison add work to each test. Keep visual tests focused on representative pages and states, and avoid capturing large surfaces when a component is the subject.
  • Reliability: Deterministic fixtures and a fixed rendering environment usually matter more than permissive thresholds. Playwright’s consecutive-capture stability check helps with transient changes but cannot make external data or environment drift stable.
  • Repository size: Baseline images become part of the test artifacts or repository workflow. Keep their scope purposeful and review image changes with the related code.
  • Cost: Playwright’s built-in screenshot assertion does not require a separate visual testing service. Hosted tools may be useful for workflow or rendering coverage, but compare their current plans and capabilities directly before adopting them; this guide makes no price or performance comparison.
  • Coverage: One Chromium viewport checks only that configuration. Add viewports or platforms when those are part of the product’s supported behavior, and manage the resulting baselines explicitly.

12. Frequently asked questions

Does a passing screenshot test prove the design is correct?

No. It shows that the capture matches its approved reference within the configured comparison rules. People still need to review the baseline and intended design.

Should every page use a full-page screenshot?

No. Use full-page capture when page composition is the requirement. For a specific component or dialog, capture that boundary to avoid unrelated changes causing failures.

Can I compare snapshots created on macOS with Linux CI?

Rendering differences can result from the host environment. Generate and compare in the same configured environment, or maintain platform-specific references when cross-platform output is intentionally tested.

Is Percy or Chromatic required for Playwright visual tests?

No. Playwright Test includes screenshot assertions and local baselines. Hosted services are optional for teams that need their review or rendering workflows.

References