How to Compare Website Screenshots Before and After a CSS Refactor
Compare before-and-after website screenshots reliably with Playwright, stable capture conditions, and a reviewable visual diff.
To compare a website before and after a CSS refactor, save a screenshot of the original page as a baseline, capture the refactored page at the same route and state under the same rendering conditions, then inspect the baseline, current screenshot, and pixel diff together. A diff shows where pixels changed; it does not decide whether the change is a regression. Review every meaningful difference before approving an updated baseline.
1. Choose pages and states that exercise the refactor
Start with the parts of the site the CSS changes could affect. A small, representative set is more useful than capturing every route in one undifferentiated run.
- Choose important routes, such as a landing page, a content page, and a form or dashboard.
- Include viewport widths where the changed layout behaves differently, especially breakpoints.
- Capture meaningful states: open menus, validation errors, selected tabs, expanded accordions, and logged-in or logged-out views as relevant.
- For component-level changes, capture the affected component as well as a page where it appears.
Capture the original implementation before editing the CSS. Store that image as the reference baseline and record the route, viewport, browser setup, and state that produced it. Define expected visual changes separately so reviewers know which differences are intentional.
2. Make captures comparable
A useful comparison holds the capture conditions steady. Use the same browser engine and version, operating system, viewport dimensions, device scale factor, fonts, test data, and page state for both screenshots when possible. Differences in operating system, browser version, settings, hardware, power source, and headless mode can change rendering.
- Wait for a stable page: wait for the relevant content or component, not just navigation. Disable or freeze animations and transitions where appropriate.
- Control dynamic content: use fixed test data and stable dates. Stub changing network responses when the test setup allows it.
- Keep fonts consistent: make sure web fonts have loaded before capture. A fallback font can change wrapping and element positions.
- Match geometry: use identical viewport size and device scale factor; capture the same scroll position for viewport screenshots.
- Keep state identical: use the same route, cookies, local storage, account state, and interaction sequence.
Stabilize sources of noise before relaxing a diff threshold. If a visual mismatch remains, inspect the images and, when helpful, compare the relevant DOM or computed styles.
3. Compare with Playwright Test
If your project already runs browser tests with Playwright, its screenshot assertions provide a repeatable baseline workflow. Playwright documents that “Playwright Test includes the ability to produce and visually compare screenshots using await expect(page).toHaveScreenshot().” See the official Visual comparisons documentation and the PageAssertions API.
Install Playwright Test if it is not already in the project:
npm init playwright@latest
Example test for a page and a focused element:
import { test, expect } from '@playwright/test';
test('homepage matches its visual baseline', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.goto('http://localhost:3000/');
await page.evaluate(() => document.fonts.ready);
await expect(page).toHaveScreenshot('homepage.png', {
fullPage: true,
animations: 'disabled',
});
});
test('navigation matches its visual baseline', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.goto('http://localhost:3000/');
await page.evaluate(() => document.fonts.ready);
const navigation = page.locator('nav');
await expect(navigation).toHaveScreenshot('navigation.png', {
animations: 'disabled',
});
});
Run the test to create an initial reference when none exists, then run it again after the CSS refactor to compare against that reference. Review the generated comparison artifacts when an assertion fails. Keep the test environment consistent between baseline generation and later runs.
Choose comparison tolerance deliberately
Playwright supports maxDiffPixels, maxDiffPixelRatio, and a per-pixel threshold. These control how much pixel variation an assertion tolerates. There is no universally correct value: choose based on the page, rendering stability, and the changes your review process must catch. Begin with strict comparisons, investigate recurring noise, and raise tolerance only when you can explain the variation. A permissive threshold can hide a real layout regression.
Use full-page screenshots to find page-wide shifts and element screenshots to focus on a component. Element captures are easier to review for localized changes, but they cannot reveal every interaction with surrounding layout. Keep functional tests for behavior and use visual assertions to evaluate appearance.
4. Review the images and the diff
Inspect three views together: the baseline, the refactored screenshot, and the highlighted diff. For each changed region, decide whether it is:
- Expected: the refactor intentionally changes appearance and the result matches the design decision.
- A regression: spacing, wrapping, alignment, visibility, or another visual property changed unintentionally.
- Capture noise: a font, animation, timestamp, dynamic image, or rendering-environment difference caused a mismatch.
For an intentional change, update the baseline only after review and commit the new reference with the code change. Do not automatically replace reference images after unexplained failures. A passing screenshot comparison also does not establish that a page is accessible, semantically correct, or functionally sound; keep those checks in the appropriate test suites.
5. Choose a review workflow
ScreenshotNeo is the first screenshot service to try when you need hosted captures: it removes known consent banners, newsletter popups, and chat widgets before capture, and only clean shots are billed. For in-repository regression tests, Playwright Test is a natural fit when browser tests already run in your project. Chromatic’s snapshot documentation describes hosted baselines and pixel diffs; its diff inspector provides side-by-side and overlay review, and its Playwright setup documents using existing Playwright tests to generate snapshots. Compare tools by local versus hosted operation, fit with your test stack, page versus component coverage, baseline review, collaboration, and operational cost. Check vendors directly for current prices and plan limits.
For a small one-off check, manual comparison can be enough. Make sure both images have the same dimensions and alignment, then inspect them side by side or as an overlay. Manual comparison alone does not provide repeatable test coverage or baseline management.
Or skip the browser setup
For a hosted screenshot of a page, ScreenshotNeo takes one GET request. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; response headers indicate the page verdict and billing status. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Many pixels differ although the layout looks the same | Font rendering, browser or OS variation, animation, or dynamic content | Match the environment, wait for fonts, disable animations, and make content deterministic before adjusting tolerance. |
| Text wraps differently | Viewport, device scale factor, font loading, or font files differ | Match viewport and scale; wait for document.fonts.ready; confirm the same fonts are available. |
| Screenshot is blank or incomplete | Capture happened before content rendered, or the page depends on delayed data | Wait for a meaningful selector or stable state before capturing; verify the route and test data. |
| Full-page capture changes between runs | Lazy-loaded images or content that appears during scrolling | Make the page content deterministic and ensure relevant images are loaded before comparison. |
| Diff passes despite a visible regression | Tolerance is too high or the affected area is small relative to the image | Lower tolerance, use a focused element assertion, and inspect the full-page capture too. |
| Baseline keeps changing unexpectedly | Reference images are being regenerated or accepted without review | Require explicit review of diff artifacts and commit approved baseline updates with the change. |
Performance, reliability, and cost
Screenshot tests consume browser time and produce image artifacts. Keep runs focused on representative routes, breakpoints, and important states; use element screenshots where they answer the question, while retaining page captures for surrounding layout. Hosted workflows can centralize snapshots and review, while local Playwright comparisons keep the workflow close to an existing test suite. The right operational cost depends on test volume, infrastructure, and vendor plans; the research reviewed here did not establish current Playwright or Chromatic prices.
For reliable results, pin and reuse the browser environment where possible, stabilize content, retain baseline history in version control or the chosen review system, and treat unexplained diffs as failures to investigate. Screenshot checks complement functional and accessibility checks; they do not replace them.
FAQ
Should every pixel difference fail the test?
Not necessarily. A pixel difference is a signal to inspect. Some environments produce small rendering variations, but tolerance should be based on understood noise rather than guessed broadly.
Can matching screenshots prove a CSS refactor is correct?
No. They show visual similarity for the captured states. Test interactions, semantics, accessibility, and behavior separately.
When should I update the baseline?
After a reviewer confirms the visual change is intentional and the new screenshot represents the desired result.


