ScreenshotNeo

BlogGuides

Screenshot Testing: A Practical Guide

Learn how visual regression tests compare approved screenshots, set them up with Playwright or Cypress, reduce flaky results, and choose local or hosted tools.

By the ScreenshotNeo team4 October 20268 min read

Screenshot testing checks whether a web interface still looks as expected. In visual regression testing, you capture a deliberately chosen UI state, compare it with an approved baseline image, then review any difference. A screenshot alone is only a capture: the comparison and review are what make it a visual test.

Playwright Test includes screenshot assertions with toHaveScreenshot(). Cypress can capture screenshots, but image comparison requires a separate plugin or service. Either approach depends on stable application state and rendering conditions; changing data, fonts, time, browser versions, or animations can produce noisy diffs.

1. How screenshot testing works

  1. Capture a known state. Navigate to a page or component, provide predictable data, and wait until the interface is ready.
  2. Compare it with an approved baseline. A diff highlights changed pixels or regions. The baseline represents the appearance the team has accepted.
  3. Review the change. Decide whether the visual change is a defect or intentional. If it is intentional, review and approve the updated baseline.

A diff is a review signal, not an automatic verdict. A redesign may be correct even when many pixels change; a tiny change may still matter if it affects a shared control or important text.

2. How do I do visual regression testing with Playwright?

Playwright Test has built-in screenshot comparison. The first run of a screenshot assertion creates a baseline; later runs compare the page against it. The following TypeScript test is runnable in a project configured with @playwright/test and a page served at the configured base URL:

import { test, expect } from '@playwright/test';

test('landing page visual state', async ({ page }) => {
  await page.goto('/');
  await expect(page).toHaveScreenshot('landing-page.png');
});

For a first baseline, run npx playwright test and inspect the generated snapshot. To intentionally regenerate approved references after a reviewed design change, run npx playwright test --update-snapshots. Review the generated image changes in version control; baseline updates change what future test runs treat as expected.

Set a stable viewport and wait for the intended state

Use a fixed viewport in your Playwright configuration, and wait for a meaningful condition such as a heading or loaded component before taking the screenshot. Avoid capturing during page transitions. A project configuration can pin the viewport and browser project, for example:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  use: {
    baseURL: 'http://127.0.0.1:3000',
    viewport: { width: 1280, height: 800 },
  },
  projects: [{ name: 'chromium', use: { browserName: 'chromium' } }],
});

Start the application before running the test, or configure Playwright’s web server in the project so the expected local URL is available. Keep the browser, operating system, fonts, and headless settings consistent between baseline generation and comparison. Snapshot names can include project or platform context because different rendering environments may need distinct references.

Control diff tolerance deliberately

toHaveScreenshot() uses pixel comparison. Options such as maxDiffPixels can allow a small number of differing pixels when that is an explicit policy. Keep a threshold narrow and explain why it exists. A large tolerance can hide real regressions and does not solve an unstable page.

See the official Playwright visual comparisons documentation for assertion options and snapshot behavior. Playwright advises running comparisons in the same environment used to generate the baselines.

3. Does Cypress support visual testing?

Cypress can take screenshots with cy.screenshot() locally, in cypress run, and in CI. Cypress itself does not compare screenshot images, so add a visual comparison plugin or a hosted service. Local open-source plugins commonly compare pixels against baselines stored with the code. Managed services may provide rendering infrastructure, baseline storage, and a review workflow.

Cypress documents integrations including Applitools, Argos, Chromatic, Happo, LambdaTest SmartUI, Percy, Sauce Labs Visual, SmartBear VisualTest, and Wopee.io. This is a list of documented integrations, not an endorsement.

// Cypress capture only; this does not perform image comparison by itself.
cy.visit('/');
cy.get('[data-testid="product-grid"]').should('be.visible');
cy.screenshot('product-grid');

Follow the chosen plugin or service’s setup instructions to create and compare baselines. Cypress’s visual testing guide explains the capture, comparison, and review cycle. Its screenshots and videos guide covers screenshot capture.

4. How do I reduce flaky screenshot tests?

Flaky visual checks usually mean that the same test does not render the same state every time. Stabilize the inputs and environment before adjusting comparison thresholds.

  • Use controlled data. Seed fixtures or stub network responses so each run sees the same content, ordering, and loading state.
  • Control time. Freeze clocks or use a fixed date when timestamps, relative dates, countdowns, or rotating content appear. Cypress documents cy.clock() for clock control.
  • Wait for readiness. Wait for a meaningful selector or application condition. Avoid arbitrary short sleeps and snapshots taken while fonts, images, or transitions are still changing.
  • Disable or finish animation. Capture after motion has ended, or configure the test environment to suppress animations consistently.
  • Pin the rendering environment. Keep the OS/container, browser version, viewport, font files, and headless settings stable. Generate and compare local pixel baselines in the same CI environment.
  • Mask only uncontrollable regions. If a third-party widget or ad cannot be controlled, mask its small region where the tool allows it. Do not relax the whole page to accommodate one noisy element.
  • Keep checkpoints intentional. Focus on important pages, shared components, and meaningful interactions. Element-level checks are useful when a component has clear ownership; full-page captures are useful when page layout is what you need to protect.

Playwright lists host OS, browser version, settings, hardware, power source, and headless mode among factors that can affect rendering. Its best practices also recommend keeping tests isolated and controlling their inputs.

5. What should you test?

Choose states where a visual change would be useful to catch and someone can own the review. Common checkpoints include:

  • A key landing page at a standard viewport.
  • A shared design-system component in representative states, such as default, error, disabled, or selected.
  • A meaningful interaction, such as opening a menu or submitting a form and seeing validation feedback.
  • A responsive layout at selected widths that represent actual breakpoints.

A smaller, deliberate suite is easier to understand and review than snapshots attached to every incidental state. Use full-page screenshots when page-level layout is the concern; use element-level comparisons when a component’s appearance is the concern. Review changed pixels before approving a baseline update.

6. Should we use a visual testing service or compare screenshots in CI?

Both can work. A local workflow keeps image storage and baseline operations in your repository or infrastructure, while a managed service may provide hosted rendering, baseline approval, dashboards, and broader browser or responsive coverage. These are general trade-offs described in Cypress documentation, not an independent price or performance comparison.

Decision Local or open-source comparison Managed visual testing service
Cost Local plugins may be free; the team owns CI and maintenance. Typically a paid subscription; check current plans.
Baselines Stored and updated in the team’s repository or infrastructure. Service may manage image storage and approvals.
Rendering Team maintains the container, browser, fonts, and environment. Service may supply consistent hosted rendering.
Coverage Usually the configured environment for each run. May offer multiple browsers and responsive widths; verify current coverage.
Review Review diffs locally or as CI artifacts. May include a dashboard and pull-request review workflow.
Good starting point Small, controlled suites when the team can own baseline operations. Teams needing shared approvals, managed baselines, or wider rendering coverage.

Before choosing a paid provider, verify its current framework support, browser coverage, security and data handling, plans, and terms. If you already use Playwright, start with its native assertions. If you use Cypress, decide whether your team wants to operate a local comparator or adopt a managed review workflow.

7. Accessibility and visual testing are different checks

A matching screenshot does not establish that a page is accessible. Pixel comparison cannot determine whether contrast meets a standard, whether keyboard interaction works, or whether screen-reader semantics are correct. Pair visual checks with accessibility scans and manual or application-specific checks. Automated scans have limits; Cypress notes that no automated scan can prove an interface fully accessible and usable for people with disabilities.

Read the Cypress accessibility testing guide for its explanation of scan coverage and complementary manual checks.

8. Troubleshooting common screenshot test failures

Symptom Likely cause Fix
First run reports a missing snapshot No reference image has been generated yet. Run the test to create the baseline, inspect it, and commit it as an approved reference.
Many unrelated pixels differ in CI Baseline and test use different OS, browser, fonts, viewport, or rendering settings. Generate and compare baselines in the same pinned CI environment.
Only text edges differ Font files or font rendering differ, or capture happened before fonts loaded. Load the same fonts in both environments and wait for the page’s intended ready state.
Images or cards vary between runs Uncontrolled data, external requests, random ordering, or rotating content. Use fixtures or stubbed responses and control ordering and time-dependent content.
Screenshot catches a half-open menu or transition Capture ran before the interaction or animation completed. Wait for a stable selector/state and complete or disable the animation consistently.
A harmless widget causes recurring diffs Uncontrolled third-party content changes independently of the application. Stub it, remove it in the test environment, or mask only its small region.
Raising tolerance makes failures disappear The threshold is hiding noise without fixing its source. Stabilize state and environment first, then use a narrow, documented tolerance if needed.
Baseline update hides an unexpected change References were regenerated without reviewing the diff. Inspect image changes in code review and update only after deciding the visual change is intentional.

9. Performance, reliability, and cost

Visual checks add browser work and image comparison to a test run, so keep the suite focused on states that provide useful coverage. Capture fewer, well-owned checkpoints before scaling to many pages and viewports. Reuse the same stable setup and data fixtures, and separate slow visual jobs if they make the main feedback loop too long.

Reliability depends more on repeatable state and environment than on a permissive pixel threshold. Local tools avoid a service subscription, but require the team to maintain browsers, baselines, diff artifacts, and review conventions. A managed service can move some rendering and baseline workflow into a hosted product, typically for a subscription; actual coverage, data handling, and price depend on the provider and should be verified.

Or skip the browser setup

For a standalone website screenshot, ScreenshotNeo provides a one-request screenshot API and MCP server. It does not replace baseline comparison or visual test review, but it can capture a page without installing and managing a browser for that capture.

One cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo API documentation

Cookie banners are accepted and removed before capture, along with supported newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. AI agents can use its MCP server tools to take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.

FAQ

Is screenshot testing the same as visual regression testing?

Screenshot testing can mean capture alone. Visual regression testing adds a comparison with an approved baseline and a review of changes.

Should every component have a screenshot test?

No. Choose representative shared components and important states that someone can review and maintain.

Can visual tests replace accessibility tests?

No. They check rendered appearance, not semantics, keyboard use, screen-reader behavior, or all accessibility requirements.

Can I approve a changed screenshot?

Yes. If the visual change is intentional, review it and update the baseline so future runs use the newly approved appearance.

Sources