How to Review Only the Changes in Visual Tests
Compare visual test results with approved baselines, separate intended UI updates from regressions, and update snapshots only after review.
To review only the changes in a visual test, compare the fresh screenshot with an approved baseline, inspect the reported differences in context, and update the baseline only after confirming that each change is intentional. A diff shows that pixels changed; it does not decide whether the change is a design update or a regression.
A baseline is the previously approved screenshot used for comparison. Playwright Test can keep reference screenshots with the project and compare new captures with toHaveScreenshot(). Hosted tools such as Chromatic and Percy add web-based review and approval flows. The right workflow depends on whether you want version-controlled snapshots, hosted review, or both.
1. Keep captures stable
Before interpreting a diff, make the baseline and fresh capture as comparable as possible. Playwright notes that screenshots can vary with the host operating system, browser version, browser settings, hardware, power source, and headless mode. Run captures in a consistent environment, especially in CI.
- Use the same browser and version for baseline creation and comparison.
- Keep viewport dimensions, device scale, color scheme, locale, and relevant test data consistent.
- Wait for the page content that matters to appear before capturing.
- Reduce animation, caret, scrollbar, and other transient noise where your capture tool supports it.
Capture stabilization details are tool-specific. For example, Argos documents waiting for fonts and images to settle, hiding carets and scrollbars, pausing GIFs, and stabilizing sticky elements. Treat these as Argos behaviors, not universal settings.
2. Compare the fresh result with its baseline
With Playwright Test, use a screenshot assertion. On a first run, Playwright generates a reference screenshot; on later runs it compares the current capture with that reference. Keep tests focused on a meaningful page or component state so that a failure is straightforward to review.
import { test, expect } from '@playwright/test';
test('home page matches its approved appearance', async ({ page }) => {
await page.goto('https://example.com');
await expect(page).toHaveScreenshot('home.png');
});
Replace the example URL with the page under test. Run the test using your project’s configured Playwright Test command. The first run establishes the reference; subsequent runs compare the screenshot against it. Playwright’s official guide covers screenshot assertions and reference images: Playwright visual comparisons.
When a comparison fails
- Open the actual screenshot, expected baseline, and generated diff provided by the test run.
- Locate the changed area and inspect the surrounding layout, text, controls, and imagery.
- Check the page state and capture conditions before deciding the UI itself changed.
- If the difference is unintended, fix the application or make the capture deterministic, then rerun.
- If the difference is intended, update the reference only after review.
To update Playwright reference screenshots after approving an intended change, run the project’s Playwright command with --update-snapshots, for example:
npx playwright test --update-snapshots
Review the resulting snapshot changes in version control before committing them. Avoid updating snapshots simply to make a failing run pass: doing so can replace evidence of a regression with an unreviewed baseline.
3. Inspect differences and decide what they mean
A useful review asks whether the observed change is expected, whether it appears in the relevant browser and viewport captures, and whether any nearby behavior or content has also shifted. Check the screenshot in context rather than judging only a highlighted pixel region: small changes in wrapping, spacing, or alignment can affect nearby controls.
| What you see | What to check | Decision |
|---|---|---|
| Text wrapping or font changes | Font loading, content changes, viewport, and browser version | Fix unstable capture or application issue; approve only if the new appearance is intended. |
| Image or icon difference | Asset URL, load completion, responsive variant, and page data | Confirm the intended asset and state before accepting. |
| Broad layout shift | Viewport size, page state, CSS change, and environment consistency | Investigate as a likely meaningful UI change before updating. |
| Small scattered pixel differences | Rendering environment and transient content | Stabilize the capture if possible; do not approve noise without understanding it. |
These checks are review guidance: a diff identifies visual change, while the reviewer determines whether it is intended.
4. Approve changes at the right scope
In Chromatic, review the changed snapshots in a build and accept or reject them; accepting changes updates the baselines. In Percy, approvals can apply at build, matching-change group, or snapshot level. Percy documents that a snapshot includes its captured browser and width combinations, so an individual screenshot within that snapshot cannot be approved separately. Check the approval scope before accepting a group of changes.
- Chromatic documentation: hosted snapshot review and approval workflow.
- Percy approval documentation: approval scope for builds, groups, and snapshots.
Whichever system you use, leave a clear review record when the reason for a visual change may not be obvious from the diff. This makes future baseline changes easier to understand.
5. Choose a workflow that fits your team
| Workflow | Useful when | Review point |
|---|---|---|
| Playwright reference screenshots | You want snapshots in the project and use Playwright Test. | Keep the capture environment consistent and review snapshot updates in version control. |
| Chromatic | You want a hosted build review and approval flow. | Review changed snapshots before accepting the build changes. |
| Percy | You want hosted visual review with approval scopes for builds, groups, or snapshots. | Understand that a snapshot approval covers its captured browser and width combinations. |
| Argos | You want its documented capture stabilization behaviors as part of visual comparison. | Its documented handling of fonts, images, carets, scrollbars, GIFs, and sticky elements is specific to Argos. |
Compare local, version-controlled references with hosted review; how diffs are grouped; approval granularity across snapshots, browsers, and widths; noise controls; and integration with your test runner and CI. See the official Playwright, Chromatic, Percy, and Argos documentation for each product’s current workflow.
6. Troubleshooting visual test diffs
| Symptom | Likely cause | What to do |
|---|---|---|
| Many unrelated diffs appear at once | Browser, operating system, settings, hardware, or headless mode changed. | Restore a consistent capture environment and regenerate a baseline only if the environment change is deliberate and reviewed. |
| Text differs or wraps unexpectedly | Fonts or page content were not settled, or viewport and browser conditions differ. | Wait for the intended content and fonts where supported; verify viewport and environment. |
| Images are missing or different | The asset did not finish loading or the page selected different content. | Wait for relevant images and confirm the page state before comparing. |
| A diff appears on every run | The capture includes changing or transient content, or the page is not in a stable state. | Stabilize the page state and use the capture tool’s supported noise controls. |
| Snapshot update hides a real bug | The baseline was updated before the difference was understood. | Restore the prior reference, fix the regression, rerun, and update only for a confirmed intended UI change. |
| Approval changes more captures than expected | The review tool groups browser or width captures under a snapshot or build. | Inspect the tool’s approval scope; in Percy, a snapshot spans its captured browsers and widths. |
7. Performance, reliability, and cost considerations
Visual comparisons add screenshot capture and review work to a test run. Keep the set of screenshots focused on states that matter, and run the same predictable capture setup in CI so environment drift does not create avoidable review work. Hosted review adds a service workflow; local Playwright references stay with the project. The research sources do not establish comparable performance benchmarks or pricing, so check each vendor’s current documentation and plan details when evaluating cost.
Reliability comes from repeatable page state and environment, plus deliberate review before baseline changes. A passing comparison means the capture matched its reference under the run’s conditions; it does not establish that the reference itself is correct or that untested states are unchanged.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It can capture a URL as PNG, JPEG, WebP, or PDF, with options such as full-page capture, CSS selector capture, viewport and device presets, custom CSS and JavaScript, and wait conditions. For visual review, use it to capture a consistent current page image; keep your approved baseline and comparison workflow in your visual test system.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
See the ScreenshotNeo API documentation for request options and response headers. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Does a visual diff prove there is a bug?
No. It reports a difference from the approved baseline. Reviewers decide whether it is expected or a regression.
When should I update a baseline?
After confirming that the rendered change is intentional and reviewing the affected captures. Fix unexpected changes first.
Can I approve one browser capture inside a Percy snapshot?
Percy’s documentation says a snapshot approval covers its captured browsers and widths; an individual screenshot within that snapshot cannot be approved separately.
Why does the same page produce different screenshots on another machine?
Rendering can vary with the operating system, browser version and settings, hardware, power source, and headless mode. Keep those conditions consistent for baseline and comparison runs.


