How to add screenshot regression tests to a GitHub Actions workflow
Add Playwright screenshot assertions to GitHub Actions, keep visual baselines reliable, and inspect failures with uploaded reports.
Use Playwright Test’s toHaveScreenshot() to compare rendered pages with approved reference screenshots, then run those tests in a GitHub Actions job. Commit the reference images to the repository, keep the browser environment consistent, and upload the Playwright report even when a test fails so a reviewer can inspect the difference.
1. Add a visual assertion
Install Playwright Test if the project does not already use it:
npm install --save-dev @playwright/test
npx playwright install
Create a test such as tests/homepage.visual.spec.ts. This example assumes the application is available at http://127.0.0.1:3000 while tests run:
import { test, expect } from '@playwright/test';
test('homepage matches its approved appearance', async ({ page }) => {
await page.goto('http://127.0.0.1:3000');
await expect(page).toHaveScreenshot('homepage.png', {
fullPage: true,
animations: 'disabled',
});
});
On the first run, Playwright creates the reference screenshot. Inspect it and commit it along with the test. Later runs compare a new capture with that reference. A changed screenshot is a review signal: check the diff, decide whether the UI change is intended, and update the reference only through the project’s normal review process.
2. Configure the project for repeatable captures
Set a stable base URL, browser project, and output locations in playwright.config.ts. This example starts a local development server for the test run:
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: Boolean(process.env.CI),
retries: process.env.CI ? 2 : 0,
reporter: [['html', { open: 'never', outputFolder: 'playwright-report' }]],
use: {
baseURL: 'http://127.0.0.1:3000',
trace: 'on-first-retry',
screenshot: 'only-on-failure',
},
projects: [
{
name: 'chromium',
use: { ...devices['Desktop Chrome'] },
},
],
webServer: {
command: 'npm run start -- --host 127.0.0.1',
url: 'http://127.0.0.1:3000',
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
});
Adapt the server command and URL to the application. A production build served locally is often more representative than a development server; whichever you choose, use the same mode for baseline creation and CI. If the app is already deployed in the workflow, remove webServer and set baseURL to the test target.
Choose the comparison scope
toHaveScreenshot()with no options compares the page screenshot using the configured defaults.fullPage: truecaptures the full scrollable page; leave it off for viewport-only checks.- Pass an element locator to compare one component:
await expect(page.locator('header')).toHaveScreenshot('header.png'). - Use explicit names so references remain understandable and stable. The first argument is the expected screenshot filename.
- Set a project viewport and device consistently. A responsive layout at one viewport does not validate other breakpoints; add separate projects or tests if those are important.
Playwright supports screenshot assertion options such as pixel difference thresholds, masks for dynamic regions, and animation handling. Use tolerances only for known rendering noise; broad thresholds can hide real regressions. Mask a region only when its content is inherently variable and its appearance is not the subject of the test. See the snapshot assertion reference for the current options and defaults.
3. Run the tests in GitHub Actions
Add .github/workflows/visual-tests.yml. Replace the action-version placeholders with reviewed versions and use the Node version supported by your project and current Playwright release.
name: Playwright visual tests
on:
pull_request:
push:
branches: [main]
jobs:
visual-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@<reviewed-version>
- uses: actions/setup-node@<reviewed-version>
with:
node-version: '<project-version>'
cache: npm
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
- uses: actions/upload-artifact@<reviewed-version>
if: ${{ !cancelled() }}
with:
name: playwright-report
path: playwright-report/
retention-days: 30
The sequence checks out the code, installs Node dependencies, installs Playwright’s browser and operating-system dependencies, runs the tests, and saves the HTML report as an Actions artifact. Choose an artifact retention period that fits your repository. The Playwright CI guide has a live example; its action and runtime versions can change, so consult it when updating this workflow.
4. Create and review screenshot baselines
- Run the application and tests locally with the same browser project and viewport as CI.
- On the initial run, inspect the generated expected screenshots. Confirm they show the intended UI and contain no accidental personal data or secrets.
- Add the generated reference files to version control. Playwright stores snapshots alongside the test by default; consult its snapshot path configuration if you want a different layout.
- Push the test and references. GitHub Actions will compare subsequent captures with the committed references.
- When a pull request intentionally changes the design, inspect the actual, expected, and diff images in the report or downloaded artifacts. Update and commit the reference in the same reviewed change.
Do not use --update-snapshots unconditionally in ordinary pull-request CI: that would replace the expected image instead of asking whether the change is acceptable. Run it deliberately when approving a change, review the generated files, and include those files in the commit.
5. Keep comparisons stable
Screenshot comparisons can vary with the environment. Playwright recommends using a consistent environment for screenshot and visual-regression tests; a container is one documented way to achieve this. Keep the browser version, operating-system libraries, fonts, viewport, color scheme, locale, timezone, and relevant test data stable where practical. Avoid comparing baselines made on one operating system with captures from another.
- Pin Playwright in the lockfile and install its matching browser with the documented install command.
- Wait for the page to reach a meaningful state before capturing. Prefer a locator assertion or an application-ready signal over an arbitrary long sleep.
- Disable or mask animations and variable content only when appropriate. Time, rotating promotions, randomized data, remote avatars, and live counters can create noisy diffs.
- Use deterministic fixtures and stable test accounts. Avoid depending on live third-party services for critical visual baselines.
- Start with one browser and a small set of high-value pages. Add browsers and viewports when they cover a real risk, since each additional combination adds CI work and baseline files.
Retries can help reveal intermittent failures, but a retry does not make an unstable screenshot trustworthy. Investigate tests that pass only on retry. Playwright’s CI guidance and snapshot guidance explain the recommended setup and comparison workflow.
6. Inspect a failed run
Open the failed GitHub Actions run and download the playwright-report artifact. The HTML report provides test results and failure details; traces can help diagnose what happened before the screenshot assertion. The configuration above records a trace on the first retry. See Playwright’s CI introduction for viewing logs, reports, and traces.
When a run fails, determine whether the page changed intentionally, rendered differently because the environment drifted, or failed to reach the expected state. Review the expected, actual, and diff images before changing a baseline.
Options beyond repository snapshots
Playwright-native snapshots keep assertions and reference images in the test project. If hosted review is a better fit for your team’s process, Percy documents a Playwright integration and an optional reporter gate, while Chromatic documents GitHub Actions workflows for visual tests and Storybook publishing. These are different review and operating models; compare baseline management, review flow, gating, runtime, and repository complexity for your own needs. The cited documentation establishes integration capabilities, not current prices or universal suitability.
Or skip the browser setup
If the goal is to capture a page image in a workflow without installing and managing a browser, ScreenshotNeo provides a screenshot API and MCP server. Its API can capture a target URL in one request; see the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Replace the example URL with a page your workflow is authorized to capture, and store the API key as a GitHub Actions secret rather than committing it. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These are page captures rather than Playwright’s committed reference-image assertions, so use the API when its capture workflow fits the job.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
Performance, reliability, and cost
A Playwright job’s time and compute use grow with the number of tests, browsers, viewports, and retries. Start with a focused set of important pages, avoid redundant full-page captures, and add matrix coverage only where it catches a meaningful class of defects. Installing browsers and system dependencies is part of the CI setup; use the documented Playwright approach and keep dependencies aligned with the package version.
Reliability depends on repeatable rendering and reviewable baselines. Upload reports after execution, preserve enough artifact history to investigate failures, and treat baseline updates as code changes. GitHub Actions artifact retention and runner usage affect repository operations and CI consumption; choose settings according to the project’s needs. The research sources do not establish a universal runtime, price, or savings figure.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| First run fails because the snapshot is missing | No reference has been created yet. | Run the test deliberately, inspect the generated image, and commit the approved baseline. |
| Screenshot differs on every run | Dynamic content, animations, fonts, browser, or other environment details vary. | Stabilize inputs and environment; wait for the intended page state; disable or mask only known irrelevant variation. |
| Navigation times out or the page is blank | The server did not start, the configured URL is wrong, or the app is not ready. | Check the server command and readiness URL, inspect job logs, and ensure the app is reachable from the runner. |
| Browser executable or shared library is missing | Playwright’s browser or operating-system dependencies were not installed. | Run npx playwright install --with-deps in the Linux job and keep Playwright and installed browsers aligned. |
| Large unexpected diff after a dependency update | Rendering changed with a browser, font, or system-library change. | Review the environment change and diff; regenerate references only if the new rendering is intended. |
| Artifact is missing after a failure | Upload step did not run, or its path does not match the reporter output. | Keep if: ${{ !cancelled() }}, check the configured report directory, and inspect the upload step logs. |
| Test passes only on retry | Race, flaky page readiness, or unstable content. | Wait on a meaningful readiness condition, make test data deterministic, and investigate rather than relying on retries. |
FAQ
Do screenshot regression tests replace functional tests?
No. They detect rendered visual changes; retain assertions for behavior, accessibility, and application logic.
Should every page have a full-page baseline?
No. Cover high-value layouts and components. Full-page snapshots are useful for page-level structure, while element snapshots can make focused checks easier to review.
Can a workflow approve screenshot updates automatically?
It can run an explicit update command, but automatic baseline replacement removes the review gate. Keep approval in the pull request process when the screenshot is intended to detect regressions.
Can I use the same baseline on every operating system?
Comparisons are most reproducible when baseline generation and CI use the same browser and rendering environment. Use separate baselines when distinct environments are a requirement.


