ScreenshotNeo

BlogGuides

Visual Review for Pull Requests: A Guide to UI Changes

Build a repeatable pull request workflow for reviewing UI changes with human feedback and screenshot comparisons, from local Playwright checks to shared review.

By the ScreenshotNeo team4 October 202610 min read

A visual review asks whether a rendered UI change is intentional and correct. Screenshot comparisons help reveal what changed; they cannot decide whether the change is right. A reliable pull request (PR) process combines the author’s explanation, human inspection of relevant states, automated screenshot checks against accepted baselines, and explicit approval before merge.

This guide uses Playwright Test for a local, repository-based workflow. It also explains when a hosted review workflow may help a team. The core practice applies whichever screenshot tooling you choose: define the expected change, inspect the rendered result, and update reference screenshots only when the change is intentional.

1. Decide what the PR should change

Before reviewing pixels, understand the intended user-visible effect. Read the PR description and identify the affected surfaces, states, and users. If the change is hard to infer from code, ask the author to include a preview link or screenshots in the PR description.

  • Which page, component, or flow changed?
  • What should look different, and what should remain stable?
  • Which states matter: empty, loading, error, success, hover, focused, expanded, or authenticated?
  • Which screen sizes, themes, locales, or content lengths could expose a problem?
  • Does the change alter layout, copy, imagery, interaction, or accessibility cues?

Write down the intended change in concrete terms, such as “the navigation collapses below the tablet breakpoint” or “the error message now appears below the field.” This gives reviewers a standard for interpreting a detected difference.

2. Review the rendered UI by hand

Open the PR preview or run the app at the relevant route. Inspect the intended change first, then scan for unintended effects nearby. Compare the result with the design intent and existing product patterns.

  1. Check layout and alignment, including whether content overlaps, clips, or shifts unexpectedly.
  2. Read visible text at realistic lengths. Look for truncation, wrapping, missing labels, and confusing error copy.
  3. Inspect imagery, icons, typography, spacing, and color consistency.
  4. Exercise the changed interaction and at least one nearby state, such as keyboard focus, validation, loading, or dismissal.
  5. Repeat at the relevant narrow and wide viewport sizes. Check a different locale or theme when the change could affect it.
  6. Confirm that the page still communicates state without relying only on color or a transient animation.

A screenshot diff is evidence that pixels changed, not proof of a defect. A large diff may be the intended redesign; a tiny diff may hide a broken focus indicator or a one-pixel overflow that matters. Use the image as a prompt to inspect the interface and decide whether the result matches the stated intent.

3. Add local screenshot checks with Playwright

Playwright Test provides screenshot assertions with toHaveScreenshot(). The assertion captures the rendered page or element and compares it with a stored reference image. The first run can create a reference; subsequent runs compare against it. Review and commit reference images deliberately so the repository records the expected visual state.

Install and configure

In a JavaScript or TypeScript project, install Playwright Test and its browser. The following commands and files provide a minimal runnable example. Adapt the start command and route to the application.

npm init -y
npm install --save-dev @playwright/test
npx playwright install chromium

Create playwright.config.ts:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  use: {
    baseURL: 'http://127.0.0.1:3000',
    browserName: 'chromium',
    viewport: { width: 1280, height: 800 },
    screenshot: 'only-on-failure',
  },
  webServer: {
    command: 'npm run dev -- --host 127.0.0.1',
    url: 'http://127.0.0.1:3000',
    reuseExistingServer: !process.env.CI,
    timeout: 120_000,
  },
});

Add a script to package.json:

{
  "scripts": {
    "test:visual": "playwright test"
  }
}

Create tests/home.visual.spec.ts:

import { test, expect } from '@playwright/test';

test('home page visual baseline', async ({ page }) => {
  await page.goto('/');
  await expect(page.getByRole('main')).toBeVisible();
  await expect(page).toHaveScreenshot('home.png', {
    fullPage: true,
    animations: 'disabled',
  });
});

Run the test:

npm run test:visual

When the reference does not yet exist, Playwright reports that it created a snapshot. Inspect the generated image before committing it. Keep the reference images with the test artifacts according to the team’s chosen snapshot workflow; Playwright documents snapshot paths and comparison behavior in its visual comparisons documentation.

Capture a specific element or state

For focused coverage, assert a component instead of the whole page. Make the state deterministic first: load known data, dismiss consent or onboarding only when that is the intended state, and wait for a meaningful readiness condition.

test('account panel visual state', async ({ page }) => {
  await page.goto('/account');
  const panel = page.getByTestId('account-panel');
  await expect(panel).toBeVisible();
  await expect(panel).toHaveScreenshot('account-panel.png', {
    animations: 'disabled',
  });
});

Prefer role, label, or stable test-id locators that describe the intended element. Avoid selectors tied to incidental DOM structure if a refactor should not invalidate the visual test.

Make the screenshot repeatable

Visual assertions are sensitive to environment and timing. Reduce avoidable variation before interpreting a diff:

  • Use fixed test data and predictable application state.
  • Choose a fixed viewport and run comparisons in a consistent browser and operating system environment.
  • Disable or finish animations when motion is not what the test is checking.
  • Wait for the content that matters rather than relying on an arbitrary short delay.
  • Keep dynamic content stable or mask it only when the changing area is irrelevant to the review.
  • Use fonts and assets that are available before the screenshot is taken.

Do not mask a region simply because it causes failures. First determine whether the variation reveals a real issue. If a timestamp or rotating advertisement is deliberately outside the test’s purpose, isolate that specific region and document why it is excluded.

4. Inspect the diff and update baselines intentionally

When a screenshot assertion fails, open the actual image and the diff output. Relate the changed pixels to the PR’s stated intent. Check for both the expected change and collateral changes elsewhere on the page.

  • Expected and correct: approve the visual change and update the reference image through the repository’s normal review process.
  • Unexpected: fix the implementation or test setup, then rerun the comparison.
  • Unclear: ask the author or design owner to clarify the intended appearance before accepting a new baseline.

Playwright supports updating snapshots with --update-snapshots. Use it only after deciding that the observed rendering is the new expected result:

npx playwright test --update-snapshots

Review the resulting image changes in the PR. A baseline update is a change to the team’s visual expectation; it should receive the same scrutiny as the code that produced it. See the Playwright snapshot documentation for reference screenshot handling and update behavior.

5. Choose local checks or hosted visual review

Local screenshot comparison and hosted visual review solve related but distinct workflow needs. Playwright can run screenshot assertions in the existing test stack and compare against repository reference images. Hosted services can provide a shared place to inspect diffs and collect review feedback.

Approach What the sourced workflow supports Questions for your team
Playwright Test (local/repository workflow) Screenshot assertions with toHaveScreenshot(), reference image comparison, and intentional baseline updates. Who reviews snapshot file changes? Is the CI browser environment consistent? How will the team keep dynamic regions stable?
Chromatic (hosted review and tests) Chromatic documents UI Tests against accepted baselines and a separate UI Review workflow for comparing PR changes. Its documented test dimensions include browsers, viewports, themes, locales, and CSS media features. Do designers or product stakeholders need a shared review surface? Which coverage dimensions are important? Who owns baseline approval?
Percy (hosted snapshot comparison) Percy’s official Playwright example demonstrates uploading snapshots and reviewing visual differences in Percy. Does its hosted workflow fit your existing browser tests and PR process? Who handles review feedback and snapshot maintenance?

Chromatic describes UI Review and UI Tests as separate workflows: tests compare story snapshots with accepted baselines, while review compares branch changes for pull requests. Its documentation explains that UI Review shows what will change on the base branch when the PR merges. See Chromatic’s pull request workflow, branching and baselines, and review documentation. For Percy’s demonstrated Playwright flow, see the official example repository.

Choose by fit rather than by a feature checklist alone. Evaluate integration with your test stack and CI, where baselines live, how designers and engineers review changes, the browsers and states you need, and who will maintain snapshots and resolve noisy differences. The cited documentation does not establish current pricing or exact feature parity across these services, so verify those details directly before selecting a plan.

6. Add visual review to the PR checklist

A short checklist makes visual review repeatable without turning every PR into a full design audit. Tailor it to the risk and scope of the change.

  • [ ] The PR description explains the intended user-visible change.
  • [ ] Relevant pages, components, and interaction states were identified.
  • [ ] A preview link or screenshots are attached when code alone does not make the result clear.
  • [ ] The changed UI was inspected at relevant viewport sizes and states.
  • [ ] Screenshot checks ran where the project has visual baselines.
  • [ ] Every diff was reviewed; intentional baseline changes are visible in the PR.
  • [ ] Required design, product, and engineering reviewers have approved before merge.

For a low-risk copy adjustment, that may be a quick preview check. For a shared navigation redesign, it may mean checking multiple routes, viewport sizes, themes, and interaction states. Match the review effort to the surfaces and users affected.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A one-call capture can provide an image to attach to a PR or inspect in a review workflow; it does not replace deciding whether the UI change is intended. The API accepts a URL and returns an image or PDF. See the ScreenshotNeo API documentation for parameters and options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets AI agents use screenshot and page-information tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

Troubleshooting visual review

Symptom Likely cause Fix
Snapshots fail repeatedly with small differences Rendering varies because of fonts, animation, timing, dynamic data, or environment. Stabilize test data and browser environment; wait on meaningful readiness conditions; disable irrelevant animation; isolate only genuinely dynamic regions.
The first run creates snapshots, but the PR has no meaningful visual review New references were accepted without inspecting them. Open the generated images, compare them with the intended UI, and include the baseline changes in code review.
A broad diff appears after a small code change A shared style, font, viewport, or browser environment changed, or the capture includes unstable content. Inspect the diff pattern, confirm the environment and assets, then determine whether the scope is expected before changing baselines.
Screenshot is blank or incomplete The capture may occur before the application or relevant content is ready, or the route/test data may be wrong. Verify the route and fixture, wait for a visible landmark or component, and confirm the app server is ready before capture.
Snapshot changes differ across developer machines and CI Browser or operating system rendering differs between environments. Run baseline generation and comparison in a consistent environment and document where snapshots should be updated.
Reviewers cannot tell whether the change is intended The PR omits visual context or a clear statement of expected behavior. Add a preview link or screenshots and describe the expected differences, affected states, and responsive behavior.
Tests pass but a visual bug reaches review Automated coverage did not include the affected route, state, viewport, or interaction. Map coverage to actual user-visible surfaces and add a focused check for the missing state; keep human review for intent and usability.

Performance, reliability, and cost considerations

Screenshot review adds browser work to local development or CI. Keep the suite focused on high-value routes and states, and avoid capturing every possible combination without a reason. Broader coverage across browsers, viewports, themes, and locales can catch more issues, but it also creates more snapshots for a team to review and maintain.

Reliability depends on repeatable setup: deterministic data, a ready application, stable assets, consistent capture environments, and deliberate handling of dynamic regions. A passing comparison means the rendered output matched the stored reference under that test configuration; it does not establish that the reference is good or that untested states are correct.

Local comparison uses the test workflow and stored references your team maintains. Hosted workflows add a shared review surface and snapshot handling, so evaluate the service workflow and current pricing directly. Do not approve a changed baseline merely to clear a failing check; first establish whether the changed rendering is intended.

FAQ

Does a screenshot diff tell me whether the UI is correct?

No. It identifies a difference from a reference. A reviewer must compare that difference with the intended behavior and design.

Should every component have a visual test?

Prioritize components and states where visual regressions affect important user journeys or shared UI. Add coverage where it catches a risk the existing review process could miss.

When should a baseline be updated?

After a reviewer confirms that the rendered change is intentional and correct. Review the updated image and the code together.

Can hosted UI review replace pull request approval?

It can make visual changes and feedback visible to collaborators, but the team still needs its own required approvals and merge checks.

Sources