How to Integrate Visual Tests with GitHub
Run Playwright visual tests on GitHub pull requests, inspect screenshot diffs, and choose a merge policy that fits your team.
Integrate visual tests with GitHub by adding a GitHub Actions workflow in .github/workflows that runs on pull requests, installs the project and browser dependencies, executes screenshot comparisons, and uploads the report and failure screenshots. For reliable results, keep the browser, operating system, fonts, viewport, and test data consistent with the environment used to create or update baselines.
This guide uses Playwright Test for browser-driven visual checks. It covers baseline setup, a runnable GitHub Actions workflow, reports and artifacts, merge gating, hosted review options, and common CI failures.
1. Choose what to compare
Visual tests compare rendered pixels against an accepted reference. First decide which pages or states are important enough to maintain as baselines. Good candidates include stable, high-value UI states such as a checkout summary, navigation menu, or a responsive layout at a supported viewport. Keep test data deterministic: dates, randomized content, rotating promotions, and user-specific values can create noise unrelated to a code change.
For browser-driven pages and user flows, Playwright screenshot assertions fit naturally into an existing Playwright suite. A basic assertion captures the page and compares it with a stored baseline:
import { test, expect } from '@playwright/test';
test('home page matches its visual baseline', async ({ page }) => {
await page.goto('http://127.0.0.1:3000');
await expect(page).toHaveScreenshot('home.png', {
fullPage: true,
animations: 'disabled',
});
});
On the first run, Playwright may create a reference image, depending on the command and configuration. Review that image before accepting it as the expected state. Subsequent runs compare the current rendering with the reference. Keep baseline changes in the same pull request as the intentional UI change, and review the image diff rather than blindly updating snapshots.
For component-focused visual review, Storybook-centered teams may prefer Chromatic. Chromatic also documents a Playwright integration for end-to-end states. Percy offers a Playwright client for sending snapshots to its hosted review workflow. These hosted approaches overlap with native assertions but add managed review and pull-request integration; select based on framework fit, baseline ownership, environment control, and how changes should be approved.
2. Add a pull-request workflow
GitHub Actions workflows are YAML files stored under .github/workflows. The following example assumes the repository has an npm lockfile, a Playwright configuration, and a test script. It starts the application, waits for it to respond, runs the visual suite, and uploads the Playwright HTML report even when a test fails.
name: Visual tests
on:
pull_request:
push:
branches: [main]
permissions:
contents: read
jobs:
visual-tests:
timeout-minutes: 30
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v4
- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
- name: Install project dependencies
run: npm ci
- name: Install Playwright browsers and system dependencies
run: npx playwright install --with-deps chromium
- name: Start application
run: npm run start -- --host 127.0.0.1 &
- name: Wait for application
run: npx wait-on http://127.0.0.1:3000
- name: Run visual tests
run: npx playwright test
- name: Upload Playwright report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@v4
with:
name: playwright-report
path: playwright-report/
if-no-files-found: ignore
retention-days: 14
Adjust the example to your project. If your application is already served by a test fixture or Playwright’s webServer setting, remove the explicit start and wait steps. If the suite uses Firefox or WebKit, install those browsers too. The action versions shown are explicit examples; follow your repository’s security and action-update policy, pinning actions to the level of specificity that policy requires.
Configure Playwright’s report and server
You can let Playwright start the app and configure the HTML report in playwright.config.ts. This avoids background-process handling in the workflow:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
reporter: [['html', { outputFolder: 'playwright-report', open: 'never' }]],
use: {
baseURL: 'http://127.0.0.1:3000',
browserName: 'chromium',
},
webServer: {
command: 'npm run start -- --host 127.0.0.1',
url: 'http://127.0.0.1:3000',
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
});
With this configuration, tests can use page.goto('/'). Install the wait-on package if you use the workflow’s separate wait step; otherwise, rely on Playwright’s webServer support as shown here. Keep one server-start strategy to avoid starting the app twice.
3. Make the baseline environment reproducible
A screenshot can differ because the rendering environment changed, even when the application did not. Playwright calls out containers as one way to keep screenshot-testing environments consistent across operating systems. At minimum, control these inputs:
- Browser and OS: use the same browser engine and a consistent browser build and runner environment for baseline creation and CI.
- Fonts: install or bundle the fonts used by the page. Font substitution changes glyph widths, wrapping, and layout.
- Viewport and scale: set the viewport and device scale factor deliberately. A different width can change responsive breakpoints; a different scale can change rasterization.
- Data and time: use stable fixtures and freeze dates or other time-sensitive values where the test framework allows it.
- Animations and loading: disable animations when appropriate and wait for the page state your assertion is meant to capture.
- External content: mock or stabilize third-party content such as ads, remote avatars, and live feeds instead of relying on their timing or current contents.
Playwright’s toHaveScreenshot supports options such as fullPage, animations, maxDiffPixels, maxDiffPixelRatio, threshold, and mask. Use tolerances and masks narrowly: broad allowances can hide real regressions. Consult the Playwright screenshot assertion documentation for the full option set and semantics.
4. Inspect and retain failures
When a visual assertion fails, the report and image evidence help distinguish a real UI change from rendering noise. The workflow uploads the HTML report as an artifact, and Playwright’s report can include test output and attachments. Set retention according to how long your team needs to investigate failures and any applicable repository policy. Artifacts are useful for debugging, but they do not replace reviewing the diff and deciding whether a new baseline is intentional.
For easier local diagnosis, run the failing test with Playwright’s UI or headed mode, then compare the current screenshot, expected image, and actual image. When the change is intentional, update the baseline using the documented snapshot-update workflow, inspect the changed image files, and include them with the code change. Avoid updating snapshots solely to make CI green.
5. Choose a pull-request review and merge policy
Decide what a changed screenshot means for merging. Teams commonly choose one of these policies:
- Blocking check: a mismatch fails the job, and the author must fix it or update the reviewed baseline.
- Human visual approval: a hosted review service presents diffs for approval and reports a status to the pull request.
- Informational run: visual results are retained for inspection but do not block merging. This can help while a suite is being established, but it leaves review responsibility to the team.
Native Playwright comparisons keep the baseline lifecycle in the repository and the comparison close to the test runner. Hosted services such as Chromatic with GitHub Actions and Percy for Playwright can add hosted visual review and pull-request status integration. Chromatic documents its GitHub Actions integration, Playwright support, and CI behavior; Percy documents a Playwright client and optional failure behavior. Confirm each tool’s current configuration and behavior in its official documentation before making merge policy depend on it.
Secrets and permissions
Store hosted-service project tokens in GitHub repository or environment secrets, then pass them to the relevant CLI or action through the workflow’s env. Never commit tokens to source files or print them in logs. Use the narrowest workflow permissions the job needs; the native example only requests read access to repository contents. If a workflow runs for contributions from forks, account for GitHub’s restrictions on secrets available to those runs and choose a safe review process.
6. Or skip the browser setup
If you need a rendered screenshot for a pull-request check, preview, or visual review without managing browser installation and capture code, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF; its screenshot API can capture pages for a workflow step. It complements visual-test assertions: you still need to define what should match and how a diff affects merging.
See the ScreenshotNeo API documentation for request options. This cURL example saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Equivalent Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
- Cookie banners are accepted and removed before capture; newsletter popups and chat widgets are removed too. Each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status in headers.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable is missing | The workflow installed project packages but not Playwright’s browser binaries, or the browser package does not match the installed Playwright version. | Run npx playwright install --with-deps chromium after npm ci. Install every browser engine the suite uses. |
| Page navigation times out or shows connection refused | The app did not start, the workflow used the wrong port or host, or tests ran before the server was ready. | Check the server command and URL, bind to a reachable address, and use Playwright’s webServer setting or a readiness wait. |
| Many unrelated pixel diffs appear in CI | Fonts, browser build, OS, viewport, device scale, animation, or data differs from baseline generation. | Align the environment and test inputs; disable unnecessary animation and make dynamic content deterministic. |
| Only a small changing region differs | The page may include timestamps, randomized data, live content, or a third-party widget. | Stabilize or mock the source. Use a targeted mask only when that region is intentionally outside the test’s purpose. |
| Report artifact is missing | The report was written to a different directory, no report was generated, or the upload step only ran after success. | Set the reporter output folder, match the artifact path, and use an upload condition that runs after a test failure. Inspect the Actions job log for the actual report path. |
| Hosted service reports an authentication error | The token is absent, named incorrectly, unavailable to a forked pull request, or invalid for the project. | Check the secret name and project configuration, confirm the event has access to secrets, and keep the credential out of logs and source. |
| Visual changes do not block the pull request | The workflow is informational, a check is not required by branch protection, or the hosted integration is configured not to fail on changes. | Review the CI exit behavior and configure the intended check as required in repository branch protection. Follow the chosen human-approval process for baseline changes. |
| Baseline update makes every test pass but review is unclear | Snapshot updates were accepted without inspecting the changed references. | Review the baseline diff alongside the code change and only commit intentional changes. |
8. Performance, reliability, and cost
Visual tests consume runner time for application startup, browser installation, page rendering, and image comparison. Start with the most valuable stable cases, avoid redundant captures of identical states, and parallelize only after checking runner capacity and test isolation. Browser installation can be a noticeable part of a job; caching dependencies may help, but a cache must not create a mismatch between the browser version expected by the installed Playwright package and the binary on disk.
Reliability depends more on controlling inputs than on increasing pixel tolerance. Keep test data isolated, wait for the required page state instead of arbitrary long delays, and retain enough report evidence to investigate failures. A flaky visual check that is routinely retried or ignored stops providing dependable merge feedback; track and fix its source.
GitHub Actions usage and hosted visual-review costs depend on your repository plan, runner configuration, and each vendor’s current terms. This research does not establish current pricing for Chromatic or Percy, so check their official pricing pages before choosing a service. Native Playwright avoids a separate hosted snapshot-review product but still uses CI compute and requires your team to own baselines and review. ScreenshotNeo offers 1,000 free shots monthly, then paid plans from $5 for 3,000; its billing rules treat bot checks, blank pages, timeouts, failed loads, and cache hits as unbilled.
FAQ
Should visual tests run on every pull request?
Run them on pull requests when the feedback is useful before merge and the suite is stable enough to review. Teams can begin with a focused set of important screens and expand as they establish reliable baselines.
Do I need a hosted visual-testing service?
No. Playwright can compare screenshots against repository baselines. A hosted service is useful when your team wants managed diff review or pull-request status integration.
Can visual checks work alongside ordinary tests?
Yes. They can run in the same Playwright suite or in a separate workflow job. A separate job can make status and runtime easier to see, while a shared suite can reuse setup; choose the arrangement that fits your repository.
Can an API screenshot replace a visual regression assertion?
An API can produce screenshots without a locally managed browser, but a regression assertion still needs a reference image, comparison rules, and a decision about whether a difference blocks merging.


