ScreenshotNeo

BlogHow-to

Visual Review for Websites: Compare Screenshots and Approve Changes

Build a repeatable visual review workflow: capture consistent screenshots, compare them with accepted baselines, and approve intentional changes.

By the ScreenshotNeo team4 October 20269 min read

Visual review compares a new screenshot of a website with an accepted reference, then asks a reviewer to decide whether the differences are intentional. A difference is a signal to inspect, not proof of a defect or permission to update the reference automatically.

A reliable process is: choose important pages and states, make captures repeatable, compare against an approved baseline, inspect the changes, and accept a new baseline only after review. For local Playwright tests, use its built-in screenshot assertions. For a hosted review flow, consider a service such as Chromatic, Applitools Eyes, or Percy after checking its fit and current terms.

1. Choose what to review

Start with the pages and states that protect important parts of the product. A useful set might include a landing page, a frequently used form, a key account state, and a checkout or confirmation step. Pick checkpoints based on what your team needs to protect; capturing every possible page and state can make maintenance harder without making review more useful.

For each checkpoint, record the conditions that define it:

  • Route and meaningful application state, such as an open menu or validation error.
  • Browser and version, viewport size, device scale factor, and operating system where relevant.
  • Test data, locale, timezone, and any feature flags.
  • How the test waits for the page and handles animations or changing content.

Use the same conditions for the baseline and each new run. Otherwise, a screenshot can differ because the capture environment changed rather than because the interface changed.

2. Make captures repeatable

Visual comparisons are sensitive to what is on the page at capture time. Data that changes on every run, rotating content, timestamps, asynchronous images, animation, font loading, and browser rendering differences can all create changes that deserve investigation but may not represent a product regression.

Stabilize the test before tuning the comparison:

  1. Use predictable test data and a fixed route state.
  2. Wait for the content that matters to appear. Avoid relying only on a fixed delay when the test can wait for a specific selector or application condition.
  3. Disable or finish animations when they make capture timing inconsistent.
  4. Keep browser, viewport, device scale, fonts, and operating system consistent where practical.
  5. Mask or ignore a dynamic region only when its changing pixels are not part of what the test needs to protect.

Noise controls trade coverage for stability. For example, masking a changing advertisement can help focus on the surrounding layout, but masking too much can hide a real overlap or spacing defect. Keep the ignored area as small and explicit as possible.

3. Compare against an accepted baseline

A baseline is the reference image that a reviewer has accepted for a particular checkpoint. A new capture is compared with that reference. When they differ, inspect the result instead of assuming either that the implementation is broken or that the new image is correct.

Playwright Test supports screenshot assertions with await expect(page).toHaveScreenshot(). Its documentation also describes updating reference images with npx playwright test --update-snapshots; commit snapshot files to version control and review their changes like other code changes. See the Playwright visual comparisons documentation.

4. Add screenshot comparison to Playwright

The following is a complete minimal Playwright Test example. Install the test package, create the test file, and run it once to create its initial reference. The first run establishes the baseline; inspect and commit that image before treating later runs as a regression check.

npm init -y
npm install --save-dev @playwright/test
npx playwright install chromium

Create tests/homepage.spec.ts:

import { test, expect } from '@playwright/test';

test('homepage matches the approved visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 900 });
  await page.goto('http://localhost:3000', { waitUntil: 'networkidle' });
  await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
  await expect(page).toHaveScreenshot('homepage.png', {
    fullPage: true,
    animations: 'disabled',
  });
});

Run the test with:

npx playwright test tests/homepage.spec.ts

Replace the local URL and heading with an actual route and stable readiness condition in your application. If the site keeps network connections open or continuously polls, networkidle may never be appropriate; wait for a page-specific selector instead.

Playwright’s screenshot assertion options let you tune capture behavior and comparison sensitivity. Commonly useful choices include a stable screenshot name, full-page capture, animation handling, masking dynamic elements, and a pixel difference threshold. Consult the current Playwright documentation for supported options and environment constraints rather than copying configuration across versions without checking it.

When a change is intentional, first inspect the diff. Then update snapshots explicitly:

npx playwright test --update-snapshots

Review the changed image files and test changes before committing. Updating every baseline in response to a failing run without inspecting it can bless accidental layout or styling regressions.

5. Inspect and approve changes

Use an overlay, side-by-side comparison, or difference view when available. Look for changes in:

  • Position, size, alignment, and spacing.
  • Text wrapping, clipping, missing content, or unexpected font changes.
  • Color, borders, shadows, and contrast.
  • Responsive behavior at the viewport sizes the team supports.
  • Unexpected loading states, blank areas, overlays, or broken images.

For each difference, decide whether it matches an intended code or design change. If it does, approve the new reference and record the reason in the pull request or review system. If it does not, investigate the change or fix the test setup. Make ownership clear: a named reviewer or team should be responsible for accepting baseline changes.

6. Choose a review workflow

Tools differ in where they store baselines and how they present changes. The right fit depends on your existing test stack, review process, noise controls, coverage needs, and service requirements.

Approach What it provides Trade-off to assess
Playwright screenshot assertions Visual assertions integrated with Playwright Test and snapshots managed in the project workflow. Your team manages baseline changes and the review process in its repository and CI workflow.
Chromatic with Playwright Chromatic documents hosted snapshot comparison and a review flow where changes can be approved or rejected. Adds a hosted service and its workflow; check fit with your CI and service requirements.
Applitools Eyes with Playwright Applitools documents a visual checkpoint integration and describes Visual AI features for filtering rendering noise. Treat noise-filtering and product advantage statements as vendor claims; validate with representative pages and states.
Percy by BrowserStack Percy describes a CI/CD workflow for capturing screenshots, comparing baselines, and reviewing changes. Assess its integration and review workflow against your framework, coverage, and operating needs.

Read the vendors’ current documentation for details: Chromatic’s Playwright setup, Applitools’ Playwright integration, and Percy. These are vendor descriptions, not independent comparative benchmarks. Current pricing and plan limits are not established here, so verify them directly before choosing a hosted service.

Compare options on the questions that affect your team:

  • Integration: Does it work with the browser framework and CI you already run?
  • Baseline ownership: Are references stored and reviewed in the repository, or managed by a hosted service?
  • Review clarity: Can reviewers understand the changed region, add context, and record approval?
  • Noise controls: Can you manage dynamic data, font-rendering variation, antialiasing, and sensitivity? Verify filtering claims with your own representative cases.
  • Coverage: Can you maintain checks across the browsers, viewports, routes, and components you care about?
  • Team and cost: Who reviews changes, how are approvals recorded, and what are the current service limits and prices?

7. Capture clean reference screenshots

If you need a screenshot of a live URL as an input to a visual review, use a repeatable capture configuration and keep its URL, viewport, wait condition, and relevant state aligned across runs. A hosted capture can provide the image; your review process still needs an accepted reference and a human decision about whether a difference is acceptable.

ScreenshotNeo is a website screenshot API and MCP server for developers. Its API returns a PNG, JPEG, WebP, or PDF from one GET request. You can set a viewport or device preset, capture a full page or selected element, wait for a selector, delay, or network idle, and configure options such as dark mode, cookies, headers, custom CSS, and JavaScript. See the ScreenshotNeo API documentation for request parameters.

Or skip the browser setup

Use the same target URL for repeat captures and save the returned image for your review workflow.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for free.

8. Troubleshooting visual diffs

Symptom Likely cause What to do
A screenshot assertion fails on every run The rendered page varies, or the baseline was never reviewed and committed. Open the actual and expected images; stabilize data and capture conditions; approve a baseline only after reviewing it.
Text or layout differs only in CI Browser, fonts, operating system, device scale, or viewport differs from the baseline environment. Align the capture environment and viewport. Regenerate references in the intended environment after review.
The screenshot is blank or incomplete The capture happened before the application or important assets were ready, or the URL reached an unexpected state. Wait for a meaningful page selector, verify the route and test data, and inspect the screenshot rather than immediately updating the baseline.
Tests hang waiting for the page A network-idle condition may not occur because the site polls or keeps requests open. Wait for the specific content needed for the assertion instead of requiring all network traffic to stop.
Diffs move between runs Animation, rotating content, timestamps, random data, or asynchronous rendering changes the captured pixels. Control the changing input, disable animation when appropriate, wait for the stable state, or narrowly mask irrelevant regions.
An approved UI change still fails The updated reference was not generated, reviewed, or committed for the correct test environment. Update only the affected snapshot, inspect the file change, and ensure it is included in the branch.
A baseline update hides a regression Snapshots were refreshed without a reviewer inspecting the visual difference. Restore the prior reference, inspect the diff against the intended design, and update only after the change is understood.

9. Performance, reliability, and cost

Screenshot checks add browser work and image comparisons to the test run. Keep the suite focused on high-value routes and states, and avoid duplicate captures that do not protect a distinct behavior. Run the same browser and viewport for a given baseline so the team is not reviewing environment drift as product change.

Reliability comes from deterministic setup and explicit review. A screenshot assertion can tell you that pixels changed; it cannot decide whether a product change is intended. Keep baseline updates visible in code review, assign approval ownership, and fix recurring noise at its source instead of repeatedly accepting it.

For local Playwright, account for CI time, reference image storage, and maintenance of test fixtures. Hosted services add service and plan considerations; compare current pricing and limits directly because they can change. ScreenshotNeo offers 1,000 shots a month free without a card, then plans of $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. These are ScreenshotNeo’s stated prices and may change; check its site when planning usage.

10. Review checklist

  • The tested page and state protect a meaningful user or product flow.
  • Reference and new capture use consistent browser, viewport, data, and readiness conditions.
  • Dynamic regions are controlled or narrowly excluded with a clear reason.
  • Diffs are inspected before a baseline is accepted.
  • The reviewer and approval record are clear to the team.
  • Snapshot changes are committed and visible in the normal review process.
  • Recurring false alarms lead to test or fixture improvements rather than routine blind approvals.

FAQ

Does a screenshot difference mean the test found a bug?

No. It means the captured image differs from its reference. Review the affected area and decide whether the change is intended.

Should I capture every page and viewport?

Usually not. Begin with important routes, states, and supported viewport sizes, then expand when a real coverage gap appears.

Can an automated tool approve a new baseline safely?

Automation can capture and compare images, but baseline approval should follow review of the change against the intended design and product behavior.

Where should baselines live?

That depends on the team’s workflow. Playwright’s documented approach keeps snapshots in the project and version control; hosted services provide their own review and baseline workflows. Choose the model that makes changes visible and approvals accountable.