Visual Regression Testing: Review and Approve UI Changes
Learn how to review visual diffs, approve intentional UI changes, and keep screenshot baselines reliable with Playwright, Chromatic, or Percy.
Visual regression testing captures a rendered page or component, compares it with an accepted screenshot baseline, and surfaces differences for review. A diff is a signal, not a diagnosis: inspect the changed UI, decide whether it is intended, then approve the baseline update or fix the regression. Playwright keeps screenshot snapshots in the repository; Chromatic and Percy provide hosted review workflows.
1. What visual regression review does
A visual test asks whether a new render matches an accepted render under defined capture conditions. When they differ, the reviewer determines why. The difference may be an intended design change, an unintended layout or styling bug, or capture noise caused by unstable content or environment differences.
Approving an intentional change advances the expected baseline for future comparisons. Rejecting an unexpected change means fixing the implementation or capture setup before accepting a new baseline. Do not approve a diff just to clear a build.
Visual tests complement functional tests: they can show that pixels moved or changed, but they do not establish whether a control works or whether the changed design is correct.
2. A review and approval workflow
- Choose meaningful coverage. Select components, pages, and states where a visual defect would matter. Include representative viewports and states such as empty, populated, validation-error, or open-menu views where applicable. Treat coverage as risk-based; there is no universal number of screenshots that fits every product.
- Make capture conditions repeatable. Fix the viewport, browser, data, fonts, animation state, and other inputs that affect rendering. Avoid timestamps, randomized content, rotating banners, and external content in baseline captures when possible.
- Create an approved baseline. Capture the intended UI and review it before treating it as the expected result. Playwright uses screenshot snapshot files; Chromatic establishes a hosted baseline from snapshots.
- Capture after changes. Run visual checks in CI or the pull-request workflow. A native Playwright assertion is the comparison mechanism; your project’s CI configuration determines when it runs.
- Inspect each diff in context. Relate the changed region to the pull request, inspect surrounding layout, and check related states or viewports. Confirm whether the change is intended and whether it introduces a defect elsewhere.
- Approve or fix. Accept an intentional design change so the baseline moves forward. Reject or leave an unexplained change unresolved, then fix the code or test setup. In Chromatic, denying a change marks it as a regression and fails the build.
- Keep the decision tied to the current result. Review the latest build for the branch. Chromatic documents that branch baselines are independent until merge and that comments on old builds are disabled to keep review tied to the latest UI.
3. Playwright: repository-managed screenshot baselines
Playwright Test provides screenshot assertions through toHaveScreenshot(). Add a test that loads a deterministic page state and compares its screenshot to a stored snapshot. The first run can create a baseline; review and commit that file as part of the code change. On later runs, a mismatch fails the assertion and produces comparison artifacts for inspection.
import { test, expect } from '@playwright/test';
test('pricing page visual state', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 900 });
await page.goto('http://127.0.0.1:4173/pricing');
await page.evaluate(() => document.fonts.ready);
await expect(page).toHaveScreenshot('pricing-page.png', {
fullPage: true,
animations: 'disabled',
});
});
Run it using the project’s configured Playwright Test command, commonly npx playwright test. For a new test, inspect the generated snapshot before committing. For a deliberate change, update snapshots using Playwright’s update-snapshots option, review the resulting binary diff and test change, then commit both together.
Playwright snapshot files are maintained in version control, so the team owns baseline changes and review conventions in the repository. Keep updates narrow: a broad baseline refresh can hide unrelated regressions. See the Playwright visual comparisons documentation.
4. Choosing between Playwright, Chromatic, and Percy
| Option | Baseline and review | Useful fit | Tradeoff to consider |
|---|---|---|---|
| Playwright screenshot assertions | Snapshot files live with the project and are reviewed through repository changes. | Teams already using Playwright that want comparisons close to their tests and code. | The team manages snapshot files, rendering consistency, and review conventions. |
| Chromatic | Hosted snapshots and branch baselines; supports CI UI Tests and a separate pull-request UI Review for stakeholder feedback. | Teams that want a hosted review surface for developers, designers, or product managers, or use its component and story workflow. | Check current plan, limits, security terms, and workflow fit before purchasing. |
| BrowserStack Percy | Hosted approval UI with approval at build, matching-group, or individual-snapshot scope. Snapshot approval applies across browser and width combinations represented by that snapshot. | Teams that want approval controls at several granularities in Percy’s hosted workflow. | Check current plan, limits, security terms, and workflow fit before purchasing. |
Chromatic distinguishes automated visual testing from stakeholder review. Its documentation explains: “UI Review is different than UI Tests because it shows you what will change on the base branch when you merge a pull request.” See Chromatic’s pull-request workflow, quickstart and baseline workflow, and Playwright integration. For Percy’s approval scopes, see its approval workflow documentation.
5. Make captures stable and useful
Control the page state
- Use fixed test data and deterministic routes. Stub changing APIs where the test framework and application permit it.
- Wait for a meaningful ready condition, such as a target element or application-specific loaded state, instead of relying only on a short delay.
- Disable or finish animations and transitions. Hide a live clock, rotating content, or other known volatile region only when that region is not what the test intends to validate.
- Ensure fonts and required images have loaded before capturing. A screenshot taken during font swapping can create broad false diffs.
Control the rendering environment
- Keep browser version, operating system, viewport, device scale, and fonts consistent for repository-managed comparisons.
- Use a small set of representative viewport widths, including breakpoints where layout changes. Confirm a fix at neighboring widths when it affects responsive behavior.
- Do not treat antialiasing or tiny rasterization differences as automatically harmless. Decide whether the changed region matters and whether the environment is stable enough to make that decision.
Choose screenshot scope deliberately
- Use a component or element capture when the component is the unit under test and surrounding page variation adds noise.
- Use a full-page capture when vertical layout, overflow, or lower-page content matters.
- Capture meaningful interaction states explicitly. A default page screenshot cannot cover an opened menu, validation message, or other state that was never rendered.
6. Performance, reliability, and cost considerations
Visual checks add browser rendering and image comparison work to a pipeline. Runtime depends on the number of captures, page readiness, browser setup, and whether tests run serially or in parallel; measure in your own CI rather than relying on a universal benchmark. Keep the suite focused on high-value states and avoid waiting on unrelated network activity.
Reliability depends on consistent capture conditions and trustworthy review. A noisy baseline creates repeated false alarms; an overly broad approval can normalize real defects. Separate intentional baseline changes from code changes where possible, and make reviewers inspect the changed images before approval.
Cost depends on the chosen product and current plan terms. This research does not establish current Chromatic or Percy prices or limits, so verify their official plan and security information before committing. Repository-managed Playwright avoids a hosted visual-review service, but still uses CI and engineering time.
7. Troubleshooting common visual diffs
| Symptom | Likely cause | What to do |
|---|---|---|
| Large text or layout diff across the page | Font not loaded, different browser or OS, changed viewport, or late content shifting layout. | Wait for fonts and a stable app condition; compare the browser, viewport, and environment with the baseline. |
| Diff appears only intermittently | Animation, random data, time-dependent content, or a race with network rendering. | Freeze data and time-dependent UI, disable animations, and wait for a specific ready condition. |
| Images are blank or inconsistent | Capture began before images loaded, or remote assets vary or fail. | Wait for required images, use controlled fixtures where possible, and inspect failed network requests. |
| Only one viewport fails | A breakpoint, overflow, or responsive rule changed at that width. | Inspect the failing viewport and nearby widths; fix the layout or update the baseline only if the change is intended. |
| Baseline update produces many unrelated changes | Snapshot refresh covered too much, or capture environment changed. | Revert broad updates, isolate the intended state, and establish why rendering changed before refreshing snapshots. |
| Hosted review has stale discussion | Review is attached to an older build while a newer branch build exists. | Open the latest branch build and continue the approval decision there. |
| CI fails but local capture looks different | Local and CI rendering conditions differ, including browser version, OS, fonts, or device scale. | Align environments where possible and inspect the CI-produced artifacts as the authoritative render for that run. |
8. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return a screenshot or PDF. This is useful for capturing reference pages during visual review; it does not replace a baseline comparison or decide whether a UI change should be approved. See the ScreenshotNeo website and API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.
9. FAQ
Does a visual diff mean the test found a bug?
No. It means the rendered output differs from the accepted baseline. A reviewer determines whether the change is intended or a regression.
Should every visual change update the baseline?
Only after the changed UI has been reviewed and accepted as the intended result. An unresolved or unexplained difference should not be approved just to make CI green.
Can visual tests replace functional tests?
No. They compare appearance. Keep functional checks for behavior and interaction outcomes.
Should I start with native or hosted review?
Start from your existing capture workflow and review needs. Playwright fits repository-managed assertions; hosted tools can add a dedicated review surface and approval workflow. Confirm current product terms before selecting a paid service.


