ScreenshotNeo

BlogHow-to

How to Scan Websites for Visual Bugs

Find visual regressions by comparing repeatable screenshots with approved baselines. Set up Playwright checks, reduce noisy diffs, and choose a workflow that fits your team.

By the ScreenshotNeo team4 October 20268 min read

To scan a website for visual bugs, capture representative pages in a repeatable browser environment, compare each capture with an approved screenshot baseline, and inspect every difference before deciding whether it is a regression. Playwright Test includes screenshot assertions with toHaveScreenshot(), so teams already using Playwright can add visual checks to their existing tests.

A difference is evidence that a rendering changed; it is not proof that users will see a defect. The reliable workflow is to control the capture conditions, review the diff, and update the baseline only when the change is intended.

1. Choose pages, states, and viewports

Begin with pages where visual mistakes would matter: key entry points, high-traffic pages, forms, checkout or account flows, and shared components whose regressions could affect multiple routes. Include important interaction states, such as an open menu, validation message, expanded panel, or selected tab. A screenshot of only the initial page cannot detect a broken state the test never visits.

  1. List the routes and UI states that matter to the release.
  2. Choose viewport sizes that reflect your audience and layout breakpoints. Start with a small representative set rather than every possible width.
  3. Make test data and account state predictable. Use stable fixtures instead of live or changing content where possible.
  4. Decide which browsers to cover based on your audience and release risk. Add browser coverage when it answers a real compatibility question.

Keep the first suite focused. Once it gives useful, reviewable results, expand route, state, viewport, or browser coverage.

2. Set up Playwright screenshot comparisons

Install Playwright Test in your project if it is not already present. Its toHaveScreenshot() assertion creates a reference screenshot the first time it runs; later executions compare the current capture with that reference. Treat the first run as a baseline creation step and review the generated images before relying on them.

npm init playwright@latest
npx playwright install

Create a test such as tests/visual.spec.ts. This example assumes the development server is available at http://127.0.0.1:3000 and that the page has a stable heading.

import { test, expect } from '@playwright/test';

test('home page visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 900 });
  await page.goto('http://127.0.0.1:3000');
  await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
  await expect(page).toHaveScreenshot('home-desktop.png', {
    fullPage: true,
    animations: 'disabled',
  });
});

Run the test once to generate its baseline, inspect the resulting screenshot, then run it again to confirm the comparison is stable.

npx playwright test tests/visual.spec.ts
npx playwright test tests/visual.spec.ts

The assertion supports screenshot comparison options, including full-page capture and animation handling. Consult the Playwright visual comparisons documentation for the current options and baseline behavior. A full-page image is useful for long pages, while a viewport capture is often faster to review and can isolate above-the-fold regressions.

3. Review diffs and manage baselines

When a comparison fails, inspect the actual image, expected baseline, and diff produced by the test runner. Look for shifted elements, missing content, unexpected wrapping, changed colors, clipped controls, and differences in interactive state.

  1. Confirm the page loaded the expected route and data.
  2. Check whether the changed pixels correspond to a user-visible issue or an intended design change.
  3. If it is a defect, fix the code and rerun the comparison.
  4. If the change is intentional, review it and update the baseline as part of the same change. Do not update baselines merely to make a failing build pass.

Keep baseline changes reviewable in code review. A baseline is an approved rendering reference, not a record to refresh automatically after every mismatch.

4. Reduce noisy and flaky visual diffs

Screenshot rendering can vary with the host operating system, browser version, settings, hardware, power source, and headless mode. Playwright recommends generating and comparing screenshots in the same environment for consistency. Pin or standardize the environment used to create and check baselines wherever possible.

  • Standardize browser and operating system: use a consistent CI image and browser version for baseline generation and comparisons.
  • Fix viewport dimensions: explicitly set width and height instead of relying on machine defaults.
  • Stabilize content: use deterministic fixtures, fixed dates, and known account state. Avoid depending on live APIs or changing production data.
  • Control motion: disable or wait out animations and transitions when motion is not what the test is checking.
  • Wait for meaningful readiness: wait for a visible element or a known page state before capture. A navigation event alone may occur before images or application content are ready.
  • Review dynamic areas deliberately: timestamps, rotating promotions, avatars, and personalized content can create diffs unrelated to a code regression. Make them deterministic or exclude them from the comparison where appropriate.
  • Use focused assertions: capture a specific component when the risk concerns that component; use full-page screenshots when page-level layout is the concern.

These are operational practices derived from the documented sources of rendering variation. There is no universal threshold or configuration that makes every screenshot test stable; identify the source of noise in your own pages.

5. Test responsive layouts and browser coverage

Repeat the relevant checks at viewport sizes that exercise meaningful layout changes, such as a mobile navigation state and a desktop navigation state. Avoid adding many nearly identical widths unless they cover a distinct breakpoint or user scenario.

If your users rely on multiple browsers, include those browsers in the workflow where the release risk warrants it. Hosted visual testing products describe ways to expand browser and device coverage and support review workflows. Applitools describes Playwright integration, configurable match levels, dynamic-content handling, and hosted cross-browser or device rendering; Chromatic describes Playwright visual checks, responsive and cross-browser coverage, and captured page context for debugging; Percy by BrowserStack describes baseline-based visual testing in CI and responsive or cross-browser checks. These are vendor descriptions, not independent performance comparisons. Evaluate supported environments, diff review, integration, and cost against your own needs.

6. Choose a workflow for your team

Workflow Fits when Consider
Playwright Test Your team already uses Playwright and wants local control over screenshot assertions. Baseline upkeep, consistent execution environments, and the team’s review process.
Applitools Eyes You want to evaluate its Playwright integration, match settings, dynamic-content handling, or hosted browser and device rendering. How its matching and review workflow fits your pages and team.
Chromatic You want to evaluate its Playwright checks, responsive or cross-browser coverage, and page-context debugging. How it fits your existing test and approval process.
Percy by BrowserStack You want to evaluate baseline-based CI checks and responsive or cross-browser review. Required browsers, baseline review needs, and hosted-service fit.

These options are described by their respective vendors; the available research does not establish independent comparative performance or pricing. Pick based on the browser test stack you already maintain, the environments your audience needs, how reviewers approve changes, and the cost your team accepts.

7. Troubleshooting

Symptom Likely cause Fix
The same test produces different diffs on different machines. OS, browser version, settings, hardware, power state, or headless rendering differs. Generate and compare screenshots in a standardized environment with pinned browser versions and explicit viewport dimensions.
A screenshot is blank or captured before content appears. The application or its data had not reached the expected state when capture ran. Wait for a meaningful visible element or application-ready condition before asserting the screenshot; verify the test data and route.
Only animated elements differ. The capture occurred at a different animation frame. Disable animations for the visual assertion or wait for a stable state, unless animation itself is the behavior under test.
Differences appear in timestamps, promotions, or user-specific content. Dynamic data changed between runs. Use fixed test data or isolate changing content so it does not dominate the comparison.
A baseline update hides a real regression. The reference was refreshed without inspecting the diff. Review actual, expected, and diff images before accepting a new baseline; require code review for baseline changes.
A change appears at one width but not another. The affected breakpoint or responsive state is only exercised at one viewport. Add a capture at the relevant breakpoint and test the corresponding interaction state.

8. Performance, reliability, and cost

Visual checks add browser navigation, rendering, image capture, and comparison to a test run. Keep the suite useful by prioritizing high-risk pages and states, reusing stable setup, and adding viewport or browser combinations only when they answer a real release question. Full-page captures can include more content and take longer to inspect than focused viewport or element captures.

Reliability depends on repeatable inputs and execution conditions. A passing comparison means the rendering matched the approved baseline under that run’s conditions; it does not establish correctness in every browser or prove that the design is usable. Pair visual checks with functional assertions and manual review for important changes.

Cost depends on the workflow and service selected. The research sources do not establish pricing for the hosted visual testing products. For a self-managed Playwright workflow, account for CI browser execution and baseline maintenance in your existing infrastructure. For hosted services, check current plan details directly before choosing.

Or skip the browser setup

If you need screenshots of pages without building and maintaining a browser capture setup, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Use it to gather page captures; keep your approved visual baselines and review process so you can still decide whether a changed screenshot represents a bug.

See the ScreenshotNeo API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
  • Cookie banners are accepted and removed before capture; 60+ known consent platforms, newsletter popups, and chat widgets are removed. Each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing. Responses identify the page verdict and whether the request was billed.
  • An MCP server lets AI agents, including Claude and Cursor, use screenshot, page-info, and PDF-capture tools.
  • The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo and get 1,000 screenshots a month free, with no card.

FAQ

Does a screenshot diff always mean there is a visual bug?

No. It means the current rendering differs from its reference. Review the change to determine whether it is intentional or a regression.

Should I take a screenshot of every page?

Start with representative, high-risk routes and states. Expand coverage as the suite proves useful and the team can review its results.

Can visual checks replace functional tests?

No. Screenshot comparisons check rendered appearance. Keep functional assertions for behavior such as navigation, form submission, and validation.

How often should I update baselines?

Update them when an intended visual change has been reviewed and approved. Each baseline change should remain tied to the code change it represents.

Sources