ScreenshotNeo

BlogHow-to

How to compare screenshots of a website before and after a CSS change

Build a reliable visual regression workflow: capture a baseline, control browser and page state, compare changes, and review diffs without hiding real regressions.

By the ScreenshotNeo team4 October 20268 min read

To compare screenshots before and after a CSS change, capture the same page in the same browser, operating system, viewport, and page state, then compare the new image with a reviewed baseline. Playwright Test provides a built-in baseline workflow with expect(page).toHaveScreenshot(): the first run creates a reference screenshot, and later runs compare against it. A pixel difference is a signal to inspect, not proof that the CSS change is wrong.

This guide shows a practical Playwright workflow, ways to reduce noisy diffs, how to review intentional changes, and when a hosted comparison service or screenshot API fits better.

1. Choose what to capture

Start with the page states the CSS change could affect. A screenshot comparison only covers the state that was captured, so include the relevant routes, responsive sizes, and interaction states.

  • Route: capture the page or component where the CSS is used, plus related routes if the stylesheet is shared.
  • Viewport: include desktop and mobile widths when responsive rules may change.
  • Interaction state: capture menus, dialogs, selected tabs, validation messages, or other states whose styles are affected.
  • Capture area: use a full-page screenshot to catch page-wide layout shifts; focus on a component when diagnosing a localized change.

Keep each rendering environment’s baseline separate. Different browsers and operating systems can render fonts, controls, and scrollbars differently, so a cross-browser diff can reflect the platform as well as your CSS change. See Percy’s cross-browser documentation.

2. Set up a Playwright visual baseline

Playwright Test’s screenshot assertion is the most direct option if your project already uses its test runner. Keep the browser version, operating system, viewport, device scale, and headless settings consistent between baseline generation and later comparisons. Playwright documents those environment factors as sources of screenshot variation in its visual comparison guide.

Install the test runner

npm init -y
npm install --save-dev @playwright/test
npx playwright install chromium

Create a screenshot test

Save this as tests/homepage.visual.spec.ts. The local server must serve your page at the URL in the test. The shown pixel ratio is an example of configuration syntax, not a universal recommendation; begin with a strict comparison and tune only after reviewing repeatability and diffs.

import { test, expect } from '@playwright/test';

test('homepage visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1280, height: 800 });
  await page.goto('http://localhost:3000');

  await expect(page).toHaveScreenshot('homepage.png', {
    maxDiffPixelRatio: 0.001,
  });
});

Create and review the baseline

  1. Run npx playwright test. On the first run, the screenshot assertion creates the expected reference image. Review that image and add it to version control as part of the test workflow.
  2. Make the CSS change, then run the same test in the same environment.
  3. Inspect the actual, expected, and diff images. Decide whether each visible change is intended and acceptable.
  4. If the appearance change is intended, update the baseline deliberately after reviewing it, then commit the updated reference with the CSS change.

Do not update snapshots automatically as part of an ordinary test run: that can replace evidence of an unintended regression with a new reference before anyone reviews it. Playwright’s snapshot guide describes creating and comparing reference screenshots.

3. Make captures repeatable

CSS is not the only source of a screenshot diff. Rendering environment, page timing, animation, live content, and incomplete loading can all change pixels. Stabilize those inputs before relaxing comparison sensitivity.

Keep the environment fixed

  • Use the same browser engine and version for baseline and comparison runs.
  • Use the same operating system, viewport dimensions, device scale factor, and headless or headed mode.
  • Keep fonts and other local assets available and loaded before capture.
  • Use the same color scheme, locale, timezone, and test data if they affect the page.
  • Use one baseline per browser and platform when you need cross-browser coverage.

Playwright’s screenshot assertions wait for consecutive screenshots to match, disable CSS animations and Web Animations by default, and support masking or a stylesheet for volatile content. See the PageAssertions API for the exact behavior and options.

Control dynamic content

Replace changing values with deterministic test data where possible. If a timestamp, rotating banner, avatar, or live counter is irrelevant to the CSS review, mask that region or hide it with a screenshot stylesheet. Keep the mask narrow: masking a large region can conceal a real layout break.

await expect(page).toHaveScreenshot('homepage.png', {
  mask: [page.locator('[data-testid="live-clock"]')],
  stylePath: 'tests/visual-stability.css',
});

Example tests/visual-stability.css:

[data-testid="rotating-promo"] {
  visibility: hidden !important;
}

Only use a stylesheet to suppress content that is genuinely irrelevant to the comparison. The purpose is to stabilize the capture, not to hide the effect under review.

Handle animation and page readiness

Playwright handles CSS and Web Animations during screenshot assertions, but JavaScript-driven animation may continue changing the page. Pause or disable it in test mode when it makes captures unstable. Wait for a meaningful page condition, such as a visible heading or loaded component, rather than relying only on a fixed sleep.

Hosted capture systems may use readiness heuristics as well. Chromatic documents network quiescence as one readiness signal and notes that JavaScript-driven animations can require explicit handling. See Chromatic’s snapshot documentation.

4. Read and tune the diff

Playwright screenshot assertions support per-pixel color sensitivity and limits on the total differing pixels or ratio. The PageAssertions API documents options including threshold, maxDiffPixels, maxDiffPixelRatio, screenshot scale, masking, and styling.

  • threshold: controls how different a pixel’s color can be before it counts as a mismatch.
  • maxDiffPixels: allows a specified maximum count of differing pixels.
  • maxDiffPixelRatio: allows a specified maximum proportion of differing pixels.
  • scale: chooses CSS-pixel or device-pixel screenshot output; use the same choice for reference and actual captures.
  • mask and stylePath: conceal or stabilize selected volatile regions so they do not dominate the comparison.

There is no universal correct threshold. Begin strict enough to reveal changes, repeat the test to check stability, inspect highlighted pixels, then adjust only for known rendering variation. A permissive threshold can hide small but important shifts; a strict threshold can produce noise when the environment is not controlled.

5. Choose a workflow for the team

The right comparison workflow depends on where you want baselines to live, how you review changes, and which rendering environments you need.

Approach Fits Trade-offs
Playwright Test Teams using Playwright who want screenshot assertions in the test suite. Baselines follow the project’s test workflow; stable results depend on controlling the render environment. Screenshot assertions use the Playwright test runner.
Chromatic with Playwright Teams that want hosted snapshots and a dedicated diff-review workflow. Requires hosted-service setup, and capture and review follow that service’s workflow. See Chromatic for Playwright.
Percy Teams that need browser-specific visual comparisons. Browser and operating-system rendering differences are part of the comparison, so maintain the relevant environment-specific references. See Percy’s cross-browser documentation.

Compare tools by baseline ownership, local versus hosted review, integration with your test suite, browser matrix, dynamic-content handling, and how reviewers inspect diffs. None of these workflows removes the need to decide whether a visual change is intended.

6. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. Its API returns a screenshot or PDF from one GET request. Use the same page URL and relevant viewport settings for before and after captures, then compare the saved files or keep them as review artifacts. The ScreenshotNeo API documentation lists the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed; response headers report the page verdict and billing status.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

7. Troubleshooting visual comparisons

Symptom Likely cause What to do
A test fails on every run with widespread pixel changes The browser, operating system, fonts, viewport, device scale, or headless setting differs from the baseline environment. Restore the original environment or create and review a separate baseline for the new environment.
Only a few regions change between repeated runs Live data, rotating content, timing, or JavaScript animation is changing. Use deterministic fixtures, wait for a stable page condition, pause the animation, or narrowly mask the volatile region.
Text wraps differently A font did not load, the viewport or scale differs, or platform font rendering changed. Verify font assets and capture dimensions; compare within the same browser and operating system.
The screenshot is blank or incomplete The page was captured before navigation or important content finished loading. Wait for a page-specific selector or state that indicates the content is ready; check failed resource requests.
The diff is huge after a small CSS edit A parent layout, font, global stylesheet, or shared token may have changed; a baseline may also have been captured in another environment. Inspect actual, expected, and diff images from the top down, then isolate the affected component or route.
Small visible differences are ignored The color threshold or allowed diff count/ratio is too permissive. Reduce tolerance and inspect repeatability before accepting the change.
A harmless animation still changes pixels The animation is driven by JavaScript or media behavior outside automatic CSS animation handling. Pause it explicitly in test mode or stabilize the component before capture.

8. Performance, reliability, and cost

  • Keep the suite focused: capture the routes, states, and viewports affected by the change. Extra captures add runtime and snapshots to review.
  • Make failures diagnosable: retain the expected, actual, and diff images with the test result so reviewers can see what moved.
  • Separate environments: run cross-browser and cross-platform comparisons as distinct capture configurations with their own reviewed references.
  • Avoid excessive retries: retries can help identify intermittent infrastructure failures, but they do not make unstable screenshots deterministic. Fix the source of variation.
  • Account for hosted-service costs: hosted comparison products require service setup and may have plan-specific usage terms. Check the service’s current pricing and capture limits before adopting it; this guide makes no unsupported cost comparison.
  • Use API capture for artifact workflows: a screenshot API can simplify generating images outside a local browser test. It does not replace browser-baseline diffing unless you also store, align, and review reference captures under stable conditions.

FAQ

Does a pixel diff tell me whether a CSS change is a bug?

No. It identifies visual differences. A reviewer must decide whether they are expected and whether the resulting interface is correct.

Should I use one baseline for every browser?

No. Browser and operating-system rendering can vary. Keep references specific to the rendering environment you intend to verify.

Should every visual test capture the whole page?

Only when page-wide effects matter. A focused component capture can make local styling changes easier to inspect, while full-page capture can expose shifts elsewhere.

When should I accept an updated baseline?

After reviewing the diff and confirming the new appearance is intentional. Update the reference alongside the code change so the review history explains why it changed.