ScreenshotNeo

BlogHow-to

How to Create Visual Regression Tests from Website Screenshots

Build reliable visual regression tests by comparing repeatable website screenshots with reviewed baselines using Playwright, Storybook, or Chromatic.

By the ScreenshotNeo team4 October 20268 min read

Visual regression tests render a page or component, capture a screenshot, and compare it with an approved reference image. A difference is a review signal: investigate whether it reveals an unintended change or an intentional design update. For a code-first workflow, Playwright Test can generate and compare screenshot baselines with toHaveScreenshot(). Keep the browser, operating system, viewport, data, and other rendering conditions consistent between baseline and later runs.

1. Choose what to test

Start with the UI states where a visual change could matter. A broad test suite that captures every page, size, and interaction can be slow and noisy; a narrow suite may miss regressions. Select a representative set and expand it when a bug or important state calls for more coverage.

  • Component states: use Storybook stories for individually addressable variants such as empty, loading, error, and populated states.
  • Pages and journeys: use Playwright to navigate to routes, authenticate when needed, and establish the state that matters before capture.
  • Viewports: include the desktop and mobile sizes your users rely on. Treat each viewport as a distinct reference.
  • Interactions: capture menus, dialogs, selected tabs, validation errors, or other states after driving the UI into that state.

Make each test deterministic. Use stable test data, avoid relying on production content that changes during the test, and ensure animations, clocks, and network-dependent state do not make the captured page vary unexpectedly. A test should reproduce the same visible state when run again.

2. Create a Playwright screenshot test

Install Playwright Test in a JavaScript project and install its browser. The following is a minimal runnable setup for a page you can reach locally.

npm init -y
npm install --save-dev @playwright/test
npx playwright install

Create tests/visual.spec.js:

const { test, expect } = require('@playwright/test');

test('home page matches its approved visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 900 });
  await page.goto('http://127.0.0.1:3000', { waitUntil: 'networkidle' });
  await expect(page).toHaveScreenshot('home.png', {
    fullPage: true,
  });
});

Start the application in another terminal, then run the test:

npx playwright test tests/visual.spec.js

On the first run, Playwright creates a reference image. Review that image carefully, then commit the accepted snapshot with the test. On subsequent runs, Playwright captures the page and compares it with that reference. A failed assertion produces comparison output for inspection. Do not treat an automatically generated baseline as approved without reviewing it.

Capture a focused element

For a component or region, assert against a locator. This reduces unrelated page content in the comparison and helps identify which part changed.

const { test, expect } = require('@playwright/test');

test('pricing card matches its baseline', async ({ page }) => {
  await page.goto('http://127.0.0.1:3000/pricing');
  const card = page.locator('[data-testid="pricing-card"]');
  await expect(card).toHaveScreenshot('pricing-card.png');
});

Use a stable selector such as a test ID when possible. If the target is absent or matches more than one element, make the locator unambiguous and ensure the page has reached the state where the element exists.

Run in a pinned, repeatable environment

Playwright cautions that the operating system, browser version, settings, hardware, power source, and headless mode can affect rendering. Its guidance is to run tests in the same environment used to generate the baseline. Pin the Playwright dependency, use the same installed browser version and execution image in local and CI workflows, and keep viewport and test data fixed. A baseline generated on one operating system may differ from a run on another even when the application code has not changed. Playwright visual comparisons documentation explains screenshot comparison and environment consistency.

3. Review and update baselines deliberately

  1. Run the visual test on the change.
  2. Inspect the actual screenshot, approved baseline, and diff output.
  3. Decide whether the difference is an unintended regression or an intended UI change.
  4. For an intended change, regenerate snapshots with npx playwright test --update-snapshots.
  5. Review the regenerated files and commit them with the code change.

Updating a baseline changes what future runs consider approved. Avoid accepting every changed image in bulk without review: doing so can turn a real regression into the new reference. For intentional design work, include baseline changes in the same review as the corresponding UI change.

4. Choose local snapshots or hosted visual review

Workflow Good fit What to consider
Playwright Test snapshots Page-level routes, flows, and focused element checks in an existing test suite Reference images live with the project; keep their generation and comparison environment consistent.
Storybook with Chromatic Component libraries whose stories represent the states to check Storybook documents visual testing that captures stories and compares them with earlier versions. Its versioned setup instructions vary; check the current docs for the Storybook version in use.
Chromatic with Playwright Teams that want hosted review for snapshots produced during Playwright tests Chromatic documents extending Playwright’s test and expect utilities and uploading an archive of test states for comparison.

Storybook documents its visual testing approach for Storybook 9 and Storybook 8. The cited Storybook 8 instructions for the @chromatic-com/storybook addon require Storybook 7.6 or higher; follow current version-specific instructions before setup. Chromatic also documents Playwright integration and its snapshot configuration, including variation across browser, viewport, theme, and configuration.

When choosing a workflow, compare the states and viewports covered, ownership of reference images, capture environment, review process, and the team’s appetite for maintaining capture infrastructure. Hosted capture can provide a standardized environment and review interface; repository snapshots keep references alongside code. Select based on how the team works and what it needs to review.

5. Handle rendering noise and diff thresholds

Screenshot comparisons can flag small rendering differences caused by fonts, antialiasing, dynamic content, or environment changes. First remove avoidable variation: stabilize test data and state, pin the browser and runtime environment, and use the same viewport and device pixel ratio. Check whether an image, font, or third-party resource loaded differently before relaxing comparison settings.

There is no universal pixel threshold that fits every website. A permissive threshold can hide a real visual regression; a strict comparison can fail on harmless rendering noise. Choose comparison settings based on the UI and the failures the team needs to catch, then review diffs as evidence rather than treating a score alone as a verdict. Mask only genuinely dynamic regions where their content is irrelevant to the visual contract, and keep the rest of the page under comparison.

Device pixel ratio deserves special attention when using hosted capture. Chromatic’s snapshot documentation notes that its Capture 9 uses a device pixel ratio of 2.0; a mismatch can change every pixel. This is version-sensitive, so consult the current Chromatic snapshot documentation when diagnosing broad diffs.

6. Troubleshooting

Symptom Likely cause Fix
Many pixels differ although the UI looks unchanged Different operating system, browser build, device pixel ratio, rendering settings, or headless mode Compare in the baseline environment. Pin the browser and CI image, then regenerate a reviewed baseline only if the environment change is intentional.
The test fails intermittently Unstable data, animation, delayed resources, or a page captured before it reaches the intended state Use fixed test data, wait for a meaningful element or state, and remove or control sources of motion and changing content.
The first test run fails because a snapshot is missing No reference image has been created yet Run the test to generate the snapshot, inspect it, and commit it only after approval.
The comparison shows a genuine design change The rendered UI intentionally changed, but the stored baseline is old Review the diff, update snapshots with --update-snapshots, inspect the new reference, and commit it with the UI change.
An element screenshot times out or cannot resolve The locator is wrong, ambiguous, or the element is not yet visible Use a unique stable locator and wait for the expected page state before asserting.
Storybook or Chromatic setup instructions do not match the installed version The documentation path or addon setup is version-specific Use the current documentation for the installed Storybook and integration versions; do not copy version 8 instructions blindly into a different version.
Hosted snapshots differ everywhere after a capture change Browser, viewport, theme, or device pixel ratio changed Check snapshot configuration and current provider documentation, then determine whether the capture change or application change is intended.

7. Performance, reliability, and cost

Test runtime grows with the number of states, routes, and viewport combinations. Prioritize high-value states, capture focused elements when full-page coverage is unnecessary, and parallelize only when the test data and environment remain isolated. Keep reference images in review so comparison failures are actionable rather than just red CI status.

Reliability depends on repeatable rendering and deliberate baseline ownership. A visual test is useful when its failure is trustworthy and someone can review it. Control the capture environment, avoid volatile third-party content where possible, and record intentional reference changes with the code that caused them.

Costs depend on the workflow: local Playwright snapshots use your existing test infrastructure and repository storage; hosted review services such as Chromatic are commercial services, so check their current plans and limits directly. The sources cited here do not establish a universal runtime, cost, or defect-detection benchmark. Measure your own suite and select coverage that fits your CI budget and review capacity.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. For a visual regression pipeline, you can use its image response as a capture input, while keeping baseline approval and image comparison in your test workflow. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = require('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing. Responses identify page verdict and billing status in headers.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots each month without a card.

FAQ

Should every visual difference fail CI?

A difference should trigger investigation. Whether it blocks a change depends on whether it is an unintended regression and on the team’s review policy.

Can a screenshot test prove a page works?

No. It checks rendered appearance for captured states. Keep functional assertions for behavior, navigation, and accessibility alongside visual checks.

How often should references be refreshed?

Refresh them when an intentional visual change has been reviewed and accepted, then commit the new references with that change.