ScreenshotNeo

BlogHow-to

How to Visually Test Every GitHub Pull Request

Build a pull request workflow that captures stable UI states, compares reviewed Playwright baselines, and gives reviewers useful visual evidence.

By the ScreenshotNeo team4 October 20269 min read

To visually test every GitHub pull request, run browser screenshot assertions in a required pull request check: capture meaningful UI states, compare them with reviewed reference images, and upload the report and images so reviewers can inspect failures. A pixel difference is a signal for review; it does not decide whether the change is a bug or an intentional redesign.

This guide uses Playwright Test and GitHub Actions for a repository-owned workflow. The same CI job can run for every pull request targeting the branches you select. “Every PR” means each matching workflow run; it does not automatically cover every route, browser, viewport, or interaction. Add tests for the states your team considers important.

1. Choose the workflow that fits your project

For projects already using Playwright, screenshot assertions are a direct way to keep baselines beside the code and review changes through ordinary pull requests. Your team owns the reference images, capture environment, and approval process. If you already use Storybook or want a hosted visual review flow, Chromatic may fit; if your existing Playwright setup should send its visual comparisons to a hosted service, consider Percy. Hosted tools can provide pull request review and status checks, but are optional. Confirm current service setup, plans, and limits with their documentation before adopting them.

Approach Fits when Tradeoffs
Playwright Test snapshots You want repository-owned baselines and a native CI job. Your team maintains stable execution conditions and reviews baseline updates.
Chromatic You want hosted visual review, particularly with a supported existing workflow. Requires service setup and a project token; check current plan and limits.
Percy with Playwright You already use Playwright and want hosted comparison or an optional CI gate. Adds a hosted service dependency and token management.

Use the tool that matches how your team builds and reviews UI. A visual test suite is useful only when it captures stable, meaningful states and someone can inspect and approve the resulting changes.

2. Add a stable Playwright screenshot assertion

Install Playwright Test and its browser binaries, then define a test for a high-value route or component state. The first run creates a baseline; review that image before treating it as expected behavior. Subsequent runs compare the rendered page against the baseline.

npm init playwright@latest
npx playwright install

Example test file tests/home.visual.spec.ts:

import { test, expect } from '@playwright/test';

test('home page desktop visual state', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 900 });
  await page.goto('http://127.0.0.1:3000/', { waitUntil: 'networkidle' });
  await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
  await expect(page).toHaveScreenshot('home-desktop.png', {
    fullPage: true,
    animations: 'disabled',
    maxDiffPixels: 0,
  });
});

The exact route and heading depend on your app. Keep the assertion after the page reaches the intended state, not merely after navigation begins. Playwright documents toHaveScreenshot(), baseline generation and updates, and options such as maxDiffPixels and stylePath in its visual comparisons guide.

Cover states that matter

  • Capture the routes and component states where a visual regression has meaningful user impact.
  • Test responsive layouts with explicit viewport sizes. A desktop capture does not cover mobile.
  • Exercise important interaction states, such as an open menu or validation error, before taking the screenshot.
  • Use descriptive snapshot names so reviewers can identify route, viewport, and state.
  • Consider element-level screenshots when the page contains unrelated, frequently changing regions.

A full-page capture helps find layout shifts below the fold; a viewport capture keeps the comparison focused on what is visible initially. Pick deliberately: neither implies coverage of states your test never creates.

3. Create and review the baseline

  1. Run the visual test locally in the same browser and operating system intended for CI.
  2. Inspect the generated expected image. Confirm it represents the intended design and that the test reached the correct state.
  3. Commit the approved baseline with the test and relevant UI code.
  4. When a design intentionally changes, run npx playwright test --update-snapshots, inspect the resulting image changes, and commit only the reviewed updates.

Do not accept a baseline update merely because the test failed. First determine whether the actual image shows an intended UI change, a defect, or capture noise. Keep baseline changes visible in code review so the expected appearance has an owner and rationale.

4. Run visual tests on pull requests with GitHub Actions

Create .github/workflows/visual-tests.yml. This example assumes the application can be started with npm run start -- --port 3000 and that npm run test:e2e runs Playwright. Adapt those commands to the project. GitHub Actions supports the pull_request trigger, and Playwright’s CI guide documents installing dependencies and browsers, running tests, and uploading reports or results.

name: Visual tests

on:
  pull_request:
    branches: [main]
    types: [opened, synchronize, reopened]

jobs:
  visual:
    timeout-minutes: 30
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: npm run build
      - run: npm run start -- --port 3000 &
      - run: npx wait-on http://127.0.0.1:3000
      - run: npm run test:e2e
      - name: Upload Playwright report and results
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: playwright-visual-results
          path: |
            playwright-report/
            test-results/
          if-no-files-found: ignore
          retention-days: 14

This workflow is illustrative: add the required wait-on dependency or replace that step with your app’s readiness check. If your tests need a development server, configure Playwright’s webServer setting instead of starting it in a background shell. Choose the pull request branches and activity types to match your repository policy. Protect the relevant branch by requiring the visual job’s check before merge if visual results must gate merging.

Playwright’s CI guide includes GitHub Actions examples and discusses its container image for a consistent runner. Pin tool and browser versions through your dependency lockfile and CI setup, and keep the baseline-generation environment aligned with CI.

5. Make screenshot output stable

Rendered pixels can vary across operating systems, browser versions, browser settings, hardware, power source, and headless mode. Differences in fonts, antialiasing, timing, or data can turn an otherwise useful check into noisy review. Keep the baseline and CI capture conditions consistent: use the same operating system and browser family, pin dependencies, and avoid regenerating references on a developer’s different machine if CI uses another environment.

Control common sources of noise

  • Animations and transitions: disable them in the screenshot assertion or with a test stylesheet.
  • Dates and relative time: freeze the clock or provide deterministic dates.
  • Random content: seed the data or replace it with stable fixtures.
  • External data: mock API responses so the page does not depend on live content changing between runs.
  • Asynchronous assets: wait for a visible, meaningful readiness condition, such as a heading or image, before capturing.
  • Dynamic regions: use a custom screenshot stylesheet, documented as stylePath, to hide or normalize volatile elements where appropriate.

Do not hide broad page regions just to make the check pass: that can conceal real regressions. Begin with strict comparisons, inspect representative diffs, and adjust maxDiffPixels only when known rendering noise justifies it. There is no universal safe threshold; it depends on the page and capture conditions.

6. Give reviewers useful evidence

Make the visual job visible as a pull request check and retain its report and test artifacts even when an assertion fails. A reviewer should be able to inspect the expected image, actual image, and diff, then decide whether to request a UI fix or approve a deliberate baseline change. Playwright’s CI documentation covers uploading reports and test-result artifacts.

For a large suite, --only-changed can act as an early heuristic to run likely affected test files. Playwright cautions that it can miss tests, so run the full suite afterward; do not use changed-test selection as the only required visual check.

7. Keep pull request CI secure and dependable

  • Do not expose visual-service tokens or other secrets to untrusted pull request code. Align third-party contributor behavior with repository security settings.
  • Give workflow credentials only the access they need. A screenshot check usually should not receive broad write permissions.
  • Use deterministic fixtures and a readiness check to reduce intermittent failures from live dependencies or race conditions.
  • Retain artifacts long enough for the team’s review process, while respecting repository storage policy.
  • Keep required checks focused and predictable. If a job is flaky, identify whether the cause is the page, data, runner, or assertion before making it a merge gate.

8. Troubleshoot common failures

Symptom Likely cause Fix
Snapshot differs on every run Dynamic content, animation, timing, external data, or inconsistent environment. Stabilize fixtures and time, disable animations, wait for a real ready state, and align OS/browser versions.
Snapshot differs only in CI Runner OS, browser build, fonts, or headless rendering differs from the baseline environment. Generate baselines in the CI-compatible environment and pin the browser and dependencies.
Page is blank or incomplete in capture Test captures before app readiness, server startup, or asset loading. Wait for the app server and a meaningful page condition; inspect network failures and console output.
Every pixel changes after a dependency update Browser, font, CSS, or rendering behavior changed. Inspect the diff and dependency change; update the baseline only if the rendered change is intended.
Baseline missing First run has no approved reference, or the baseline was not committed for this platform/project. Generate it in the correct environment, review it, and commit the expected image.
Artifact is absent after a failure Upload step did not run, path is wrong, or files were written elsewhere. Use if: always(), verify reporter output paths, and inspect the upload step log.
Check is not blocking merge Branch protection does not require the job, or the check name changed. Require the current visual job check in repository branch protection settings.
CI takes too long Too many redundant states, browsers, or serial captures; excessive setup work. Prioritize high-value states, use supported parallelism carefully, and cache dependencies where appropriate. Preserve a complete required suite.

9. Performance, reliability, and cost

Runtime depends on the number and complexity of routes, viewports, browsers, and interactions. Start with high-value paths, then expand coverage where visual defects are costly. Parallel workers can shorten wall-clock time but increase runner resource use; confirm that the app and test data behave consistently under concurrency. Browser installation and application builds can also dominate CI time, so use dependency caching and a suitable runner without changing capture conditions unexpectedly.

Repository-owned Playwright baselines add image files to version control and use your CI capacity. Hosted comparison services introduce service setup, token handling, and dependence on that provider; verify current plans and limits before choosing one. No option removes review work: humans still decide whether a diff is acceptable. Treat flaky results as reliability debt because reviewers may learn to ignore an unreliable gate.

10. Or skip the browser setup

If you need page screenshots for a CI workflow, report, or review artifact without managing browser capture code, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. It is useful for capturing pages, but a screenshot API response alone is not a substitute for versioned UI-state assertions and reviewed baselines in a visual regression suite.

See the ScreenshotNeo API documentation. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before capture; known consent platforms, newsletter popups, and chat widgets are also removed. Each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month, no card required.

Frequently asked questions

Does a visual test tell me whether a UI change is wrong?

No. It identifies a difference from the approved reference. A reviewer decides whether to fix the UI or approve a deliberate design change.

Do I need a hosted service to test pull requests visually?

No. Playwright Test can compare screenshots against repository baselines in CI. Hosted tools are optional workflows.

Will one screenshot prove the whole site is unchanged?

No. It covers only the route, browser conditions, viewport, and UI state that the test captured.

Should baseline images be committed?

For Playwright’s repository-owned snapshot workflow, yes: the reviewed references are the comparison inputs and should travel with the code they describe.

Primary documentation