ScreenshotNeo

BlogHow-to

How to Update Playwright Screenshot Baselines Without Hiding Regressions

Update Playwright screenshot baselines safely: control the rendering environment, inspect every image diff, and keep intentional changes reviewable.

By the ScreenshotNeo team4 October 20267 min read

To update Playwright screenshot baselines without hiding regressions, first confirm the UI change is intentional, then regenerate only the affected snapshots in the same rendering environment as the existing references. Inspect the expected, actual, and diff images; verify every change matches the intended UI work; and commit the reviewed snapshots alongside the code change. The update command replaces the expected image. It does not approve the new appearance.

What a baseline update changes

A Playwright screenshot baseline is the expected image used for later visual comparisons. When you update it, you replace the reference image with a newly captured one. That can be appropriate for a deliberate design or content change, but it can also erase evidence of an unintended change if you accept snapshots without reviewing them.

Playwright documents npx playwright test --update-snapshots for updating references. Its visual comparison guide also recommends committing snapshot files and reviewing their changes. See Playwright visual comparisons.

Safe update workflow

  1. Identify the intended UI change. Name the implementation change and the tests whose screenshots should change. Do not regenerate references just to clear unexplained failures.
  2. Use the reference rendering environment. Run the update in the same OS, browser version, browser settings, and headless configuration used to create the existing baselines. Playwright notes that rendering can also vary with hardware and power source. If the environment itself is changing, review that as a deliberate baseline migration.
  3. Scope the update. Run only the affected test, file, or project where practical. A narrower run makes unrelated snapshot changes easier to spot.
  4. Use the documented update flag. For example: npx playwright test tests/settings.spec.ts --update-snapshots. Replace the path with the affected test file. Check the CLI documentation matching the Playwright version installed in your project before relying on update modes or defaults.
  5. Inspect each changed image. Compare expected, actual, and diff views. Decide whether each visible difference follows from the intended change. A generated image is evidence to review, not a verdict.
  6. Review code and snapshots together. Check the implementation diff and snapshot diff in the same change review. Ask for an explanation if the image changes extend beyond the expected UI area.
  7. Commit the reviewed references. Keep snapshots in version control so the expected appearance and the code that caused it can be reviewed together.

Updating snapshots from the command line

Run commands from the project directory containing the Playwright configuration. The basic documented command updates snapshots during a test run:

npx playwright test --update-snapshots

For a narrower update, pass a test file or other supported test filter:

npx playwright test tests/settings.spec.ts --update-snapshots

Use the equivalent package-manager invocation if your project does not invoke npx directly. The important part is that the test run uses the same project and rendering setup as the references. Do not assume that a command’s defaults are stable across Playwright versions.

Update modes and version differences

The current Playwright CLI reference lists the update modes changed, missing, all, and none. It documents a run without the flag as defaulting to missing, and the flag without a mode as defaulting to changed. The CLI also lists source update methods patch, 3way, and overwrite. These details are version-sensitive: consult the CLI reference for the version you use before choosing a mode or depending on a default.

As a review policy, choose the narrowest mode that updates references needed for the intended change. Broad regeneration can make review noisy and make unrelated visual changes harder to distinguish. This is workflow guidance, not a guarantee about what any particular mode will update.

Keep screenshot output comparable

Visual comparisons only provide useful evidence when the reference and new capture are comparable. Playwright warns that host OS, browser version, settings, hardware, power source, and headless mode can affect rendering. Fonts and platform rendering can also make screenshots differ between machines. With multiple projects, project names can be part of snapshot names in place of a browser name.

  • Record which browser project and operating-system environment creates the references.
  • Update and compare baselines in that environment, especially when CI is the reference environment.
  • Keep browser versions and relevant settings consistent when investigating unexplained diffs.
  • When intentionally changing the rendering environment, treat the new output as a migration: inspect the resulting differences rather than silently accepting them.

For project-specific snapshot naming and comparison settings, use the visual comparison documentation for the installed version.

Review the expected, actual, and diff images

For each changed reference, ask: does the difference correspond to the intended UI change, and is its extent reasonable? A small change to a button label should not silently explain a large change across the whole page. Check the expected image, the actual capture, and the diff image rather than relying only on a test pass after updating.

Playwright’s Trace Viewer documentation describes comparing the image diff, actual screenshot, and expected screenshot when investigating a visual regression. Those views help locate and understand differences; a human reviewer still decides whether the new appearance is correct.

Tolerances and dynamic content

Playwright supports maxDiffPixels to set a tolerated pixel difference. It also supports stylePath to apply a stylesheet while taking a screenshot, which can be used to filter volatile elements. These controls change what the comparison treats as acceptable or visible, so use them only when the permitted variation or excluded content is understood.

There is no universally safe pixel threshold established by the documentation. If your team sets a tolerance, document why that amount of variation is harmless and check that it will not make meaningful regressions too easy to pass. Likewise, hide or filter dynamic content only when excluding it is an intentional test decision.

Troubleshooting

Symptom Likely cause What to do
Snapshots differ on a developer machine but not in CI, or the reverse. The captures use different operating systems, browser versions, settings, headless modes, fonts, or hardware. Run the comparison in the environment that created the reference. Record and align the browser project and relevant rendering setup before updating.
Many unrelated snapshots change after a small UI edit. The update run was broader than the intended change, or the rendering environment changed. Review the full file list and image diffs. Re-run a scoped update in the reference environment and investigate unrelated changes instead of accepting them.
A test still fails after using the update flag. The failing comparison may not be the snapshot you intended to update, or another test/load issue may remain. CLI behavior can also depend on the installed version and mode. Read the failure details, identify the test and project, and check the matching CLI documentation. Do not treat a remaining failure as a reason to regenerate every reference.
Text or layout shifts slightly between runs. Rendering inputs or dynamic page content may be unstable. Stabilize the rendering environment and test data where possible. If volatile content must be excluded, consider a documented stylePath rule and review what it hides.
A tolerance makes a noisy test pass, but reviewers cannot explain the allowed difference. maxDiffPixels may be set without a clear test policy. Review the threshold and its scope. Keep only a tolerance whose harmless variation is understood; the documentation does not prescribe a universal value.

Performance, reliability, and cost

Snapshot updates run the relevant tests and write image references; keeping the run scoped limits unnecessary test work and keeps the review focused. Stable rendering inputs improve the reliability of comparisons. A broad update can increase review effort because every changed image needs explanation.

Playwright baseline updates are a development workflow using your existing test setup and source control. The supplied Playwright documentation does not specify a per-snapshot service price or universal runtime, so cost and speed depend on the project and environment. Avoid quoting a generic benchmark: measure your own suite if runtime matters.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. For a one-off capture, one GET request returns an image or PDF. Use a disposable test page or URL when generating a supplementary visual artifact; Playwright snapshot baselines still belong to your Playwright test and review workflow.

See the ScreenshotNeo API documentation. Example using cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie and consent banners are accepted like a visitor and removed, along with supported newsletter popups and chat widgets, before the shot.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots. Yearly billing gives two months free, and every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does updating a snapshot mean the visual change is safe?

No. It changes the expected image. Review the image diff and connect it to the intended code change before committing.

Should every Playwright project share one baseline?

Use the project and environment strategy appropriate to your browsers and platforms. Playwright notes that project names may be included in snapshot names when multiple projects are used.

What pixel threshold should I set?

The documentation provides maxDiffPixels as an option but does not give one threshold that is safe for every application. Choose and document a value based on the variation your team considers harmless.

Can I use ScreenshotNeo as the Playwright baseline updater?

The API returns screenshots, while Playwright’s documented baseline update workflow runs its tests with the update flag. Use ScreenshotNeo for website captures; keep Playwright’s test references under the Playwright workflow.