Playwright Screenshot Testing in GitHub Actions: Setup and Artifacts
Run Playwright visual tests in GitHub Actions, keep screenshot baselines stable, and download the reports and traces you need to debug failures.
Run visual tests with Playwright Test’s expect(page).toHaveScreenshot(), commit the generated baseline images, and run the suite in GitHub Actions with the same browser and operating system used to create those baselines. Upload playwright-report/ after the test step so you can download the HTML report when a run fails. Start with one worker for reproducibility; shard the suite when it grows and merge the shard reports.
This guide sets up a small runnable project, explains baseline updates and stability controls, and shows where to find failure artifacts. It also covers sharding, troubleshooting, artifact security, and an API option for capturing reference screenshots without managing a browser in CI.
1. Create a Playwright screenshot test
Playwright’s screenshot assertions are part of Playwright Test. The first run creates a reference image; later runs compare the rendered page against it. Commit the generated snapshot directory alongside the test so CI can compare against the same expected image. See the official visual comparisons guide.
For an existing Node.js project, install Playwright Test and its Chromium browser:
npm install --save-dev @playwright/test
npx playwright install chromium
Create tests/homepage.spec.ts:
import { test, expect } from '@playwright/test';
test('homepage visual baseline', async ({ page }) => {
await page.goto('http://127.0.0.1:4173/', { waitUntil: 'networkidle' });
await expect(page).toHaveScreenshot('homepage.png', { fullPage: true });
});
The test assumes your application is available at http://127.0.0.1:4173/. Change that URL to the local address and route for your app. The fullPage assertion captures the full document; omit it when you want the viewport screenshot. The initial run has no reference to compare with, so Playwright reports the missing baseline and writes the actual image. Review the file it creates in the test’s snapshot directory, then commit it.
A small playwright.config.ts can make the local server and CI behavior explicit:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
workers: process.env.CI ? 1 : undefined,
retries: process.env.CI ? 1 : 0,
reporter: process.env.CI ? 'html' : 'list',
use: {
baseURL: 'http://127.0.0.1:4173',
browserName: 'chromium',
trace: 'retain-on-failure',
},
webServer: {
command: 'npm run preview -- --host 127.0.0.1 --port 4173',
url: 'http://127.0.0.1:4173',
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
});
This example expects an npm script named preview; replace the server command and readiness URL with those used by your framework. If the app needs environment variables or a seeded test database, arrange those before the test command and avoid placing secrets in committed files. Retries can collect more diagnostics for intermittent failures, but they do not make an unstable screenshot test deterministic.
2. Generate, review, and commit the baseline
- Run the test locally in the same browser project you intend to use in CI:
npx playwright test. - When Playwright reports that the snapshot does not exist, inspect the generated actual image. Confirm that the app reached the expected state and that the captured image is the intended one.
- Commit both the test and its generated
*-snapshots/directory. - Run the test again. It should compare the current rendering with the committed reference.
- When a deliberate UI change alters the expected image, regenerate with
npx playwright test --update-snapshots, inspect every changed image, and commit the reviewed baseline changes with the code change.
Snapshot names and locations can be customized with Playwright configuration, but keep the references version controlled and review image diffs in code review. Playwright’s default snapshot naming includes context such as test, project/browser, and platform, which matters if you test multiple browsers or operating systems.
3. Add the GitHub Actions workflow
Save this as .github/workflows/playwright.yml. It checks out the repository, installs Node dependencies and the browser’s Linux dependencies, runs the suite, and uploads the HTML report even when tests fail. The !cancelled() condition skips artifact upload when the workflow has been cancelled, while still allowing it after an ordinary test failure. Playwright’s CI documentation recommends one worker in CI for stability and reproducibility.
name: Playwright visual tests
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
test:
timeout-minutes: 60
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v6
- name: Set up Node.js
uses: actions/setup-node@v6
with:
node-version: lts/*
cache: npm
- name: Install dependencies
run: npm ci
- name: Install Playwright browsers and Linux dependencies
run: npx playwright install --with-deps chromium
- name: Run Playwright tests
run: npx playwright test
- name: Upload HTML report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@v5
with:
name: playwright-report
path: playwright-report/
retention-days: 14
Action versions and Playwright’s supported Node and browser versions can change. Check the current CI example and your repository’s action policies when adopting this file. The 14-day retention here is a choice for this example, not a required setting; set a period that fits your team’s access and retention policy. If you use Firefox or WebKit, install those browsers as well (or install all supported browsers with npx playwright install --with-deps).
If your test setup does not start the app via Playwright’s webServer config, add an explicit build and start step before testing. The app must be reachable from the runner at the configured URL. For deployments requiring secrets, use repository or environment secrets and account for the fact that workflows triggered by forked pull requests generally do not receive repository secrets.
4. Keep screenshot comparisons stable
Visual output can differ across operating systems, browser versions, settings, hardware, power state, and headless mode. Generate and update baselines in an environment matching CI; otherwise harmless renderer differences can appear as failures. The Playwright screenshot documentation explains these sources of variation.
Control the environment first
- Use the same browser engine and Playwright version locally and in CI. Browser binaries are tied to the installed Playwright version.
- Use the same operating system for baseline generation and CI where practical. A Linux runner and a developer’s macOS workstation can render fonts and pixels differently.
- Keep viewport size, device scale, locale, timezone, color scheme, and browser project consistent for the test and its baseline.
- Wait for meaningful app state, not an arbitrary long delay. Prefer a locator assertion or application-ready signal before capturing.
- Remove sources of randomness: freeze test data, use fixed dates and seeded records where possible, and disable or stabilize animation and rotating content.
- Be deliberate with network-dependent resources such as remote fonts, avatars, and ads. Stub or host deterministic test assets if they make the rendering unreliable.
Playwright’s screenshot assertion retries the capture until consecutive screenshots match or the assertion times out. This helps with short-lived rendering changes, but cannot solve a page whose content keeps changing.
Mask volatile regions or apply a screenshot stylesheet
If a timestamp, randomized avatar, or other region is intentionally variable, mask only that locator rather than loosening the entire comparison:
await expect(page).toHaveScreenshot('account.png', {
mask: [page.locator('[data-testid="last-updated"]')],
});
You can also apply a stylesheet for volatile or animated elements. For example, save tests/screenshot.css:
[data-testid="rotating-promo"],
iframe {
visibility: hidden !important;
}
*,
*::before,
*::after {
animation-duration: 0s !important;
transition-duration: 0s !important;
}
Then pass it to the screenshot assertion:
await expect(page).toHaveScreenshot('homepage.png', {
fullPage: true,
stylePath: './tests/screenshot.css',
});
Use this to hide known noise, not meaningful page content. Playwright supports maxDiffPixels, maxDiffPixelRatio, and threshold to adjust image comparison sensitivity. These settings are available per assertion and through expect.toHaveScreenshot configuration. A narrow, evidence-based tolerance may help with known renderer noise; a broad tolerance can hide a real regression.
import { defineConfig } from '@playwright/test';
export default defineConfig({
expect: {
toHaveScreenshot: {
maxDiffPixels: 20,
threshold: 0.2,
},
},
});
Choose a limit based on the specific image and investigate the diff when it fails. Avoid setting a large global allowance simply to make CI green. See the test configuration reference for the supported assertion options.
5. Find and inspect artifacts after a failure
- Open the failed GitHub Actions workflow run and select the test job to read the failing assertion and call log.
- On the run’s summary page, find the Artifacts section and download
playwright-report. - Unzip the artifact and open
index.htmllocally in a browser to inspect tests, expected and actual images, and failure details. - If a trace was recorded, open it with
npx playwright show-trace path/to/trace.zip. The trace viewer shows recorded actions and screenshots and can display image diffs. - For locally generated reports, run
npx playwright show-reportfrom the project directory.
The trace viewer is particularly useful when the final screenshot differs because an earlier action or navigation left the page in the wrong state. See the official Trace Viewer guide.
If no artifact appears, check that the report directory exists, that the upload step ran, and that the workflow was not cancelled. The reporter must produce a report even when tests fail; reporter: 'html' does so. If you changed its output directory, update the artifact’s path too.
6. Shard a larger screenshot suite
A single job is easier to understand and is a good starting point. When the suite takes too long, sharding distributes test files across jobs. Each shard should upload its own blob report; a dependent job downloads the reports and creates one HTML report. Playwright documents this workflow in its sharding guide.
Set the reporter in playwright.config.ts so CI creates mergeable blob reports while local runs retain the normal HTML report:
import { defineConfig } from '@playwright/test';
export default defineConfig({
reporter: process.env.CI ? 'blob' : 'html',
workers: process.env.CI ? 1 : undefined,
use: {
browserName: 'chromium',
trace: 'retain-on-failure',
},
});
Replace the single test job with a matrix and add a merge job. Preserve the same dependency installation, browser installation, and app setup used in the single-job workflow:
jobs:
test:
name: Tests (shard ${{ matrix.shard }})
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3, 4]
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
cache: npm
- run: npm ci
- run: npx playwright install --with-deps chromium
- run: npx playwright test --shard=${{ matrix.shard }}/4
- name: Upload shard report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@v5
with:
name: blob-report-${{ matrix.shard }}
path: blob-report/
retention-days: 1
merge-reports:
if: ${{ !cancelled() }}
needs: [test]
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
cache: npm
- run: npm ci
- name: Download shard reports
uses: actions/download-artifact@v5
with:
path: all-blob-reports
pattern: blob-report-*
merge-multiple: true
- name: Create combined HTML report
run: npx playwright merge-reports --reporter html ./all-blob-reports
- name: Upload combined report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@v5
with:
name: playwright-report
path: playwright-report/
retention-days: 14
Use the action versions approved by your organization and verify current download-artifact syntax against GitHub’s action documentation when you adopt the example. If you use different test environments or operating systems across shards, configure the merge accordingly and keep environment identity clear. Sharding reduces elapsed time when enough runner capacity is available, but adds jobs, artifact handling, and coordination. It does not make an individual test faster.
7. Choose a runner and manage CI cost
ubuntu-latest is convenient, but its environment can change over time. A container can provide a more controlled operating environment; use a Playwright image corresponding to your project’s Playwright version and verify the currently supported image tag. The official CI guidance describes Docker use and browser installation.
Playwright’s CI documentation says caching browser binaries is generally not recommended because restoring the cache can take comparable time to downloading them, and Linux system dependencies still need installation. If you choose to cache, key it to the Playwright version and measure whether it helps your own workflow.
For a longer suite, first inspect which tests dominate runtime and whether setup is repeated unnecessarily. A single worker uses less concurrent runner capacity and tends to favor reproducibility; parallel workers or shards can reduce wall-clock time but consume more concurrent capacity. Start with a stable single-worker job, then increase parallelism only after checking that the screenshot tests and shared test data remain isolated.
8. Protect reports, traces, and screenshots
HTML reports, traces, screenshots, console logs, and test output may contain credentials, tokens, personal data, or internal application details. Playwright advises uploading reports and traces only to trusted artifact stores or encrypting them before upload. Set artifact access and retention to match your repository’s data policy, and avoid adding sensitive values to test names or logs. Read the official CI setup security guidance.
Keep pull-request permissions minimal. In particular, do not expose secrets to untrusted pull request code. If screenshots include real user or production data, change the test data or sanitize the captured page before storing its artifact.
9. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Snapshot does not exist in CI | The baseline was generated locally but not committed, or the runner is looking under a different project or snapshot path. | Commit the test’s snapshot directory. Check the browser project, test name, snapshot path configuration, and platform context. |
| Passes locally, fails on GitHub Actions | Different OS, browser binary, fonts, viewport, device scale, headless mode, app data, or timing. | Generate and review baselines in a CI-matching environment. Pin the Playwright dependency through the lockfile and stabilize page state and assets. |
| Many unrelated pixels differ | A renderer or environment change, a different font, a shifted layout, or a changed viewport. | Inspect actual and expected images first. Verify environment and fonts before changing tolerances; update baselines only for an intentional change. |
| Only a small region changes between runs | Dynamic content such as a timestamp, animation, rotating banner, or remote asset. | Control the content, wait for a ready state, or mask/hide the narrowly identified volatile region with a locator or stylePath. |
page.goto times out or the page is blank |
The app server did not start, the configured URL is wrong, or the page is waiting on an external dependency. | Confirm the workflow starts the app, align its port and baseURL, use webServer, and check the job log for server startup errors. |
| Browser executable or shared library missing | The browser matching the installed Playwright package is not installed, or Linux dependencies are absent. | Run npx playwright install --with-deps chromium (or install the browsers you use) after npm ci. |
| Report artifact is missing | The reporter did not write to playwright-report/, the upload path differs, or the workflow was cancelled. |
Set the HTML reporter, check its output directory, align the upload path, and inspect whether the artifact step ran. |
| Shard merge says no reports were found | The shard job uploaded the wrong directory, artifacts were not downloaded into the merge directory, or reporter mode was not set to blob. |
Confirm each shard writes and uploads blob-report/; verify the download pattern and run merge-reports against the directory containing the report ZIP files. |
| Tests fail only when parallelized | Tests share mutable data, accounts, or backend state, or the runner has insufficient resources. | Isolate test data and accounts, start with one worker, and increase parallelism gradually. Sharding also requires tests to be independently runnable. |
| Artifact contains sensitive information | The page, trace, or logs captured real data or credentials. | Use test-only data, restrict artifact access and retention, and follow your trusted-storage or encryption policy. |
For browser launch diagnostics, Playwright documents DEBUG=pw:browser npx playwright test. For headed Linux debugging, a display server such as Xvfb is needed; GitHub’s hosted runner and Playwright’s Docker image include it for the documented setup.
10. Or skip the browser setup
If your task is to capture a reference image rather than run a committed visual regression test, ScreenshotNeo provides a website screenshot API and MCP server. Its one-call API returns an image or PDF, and the same request can be made with cURL, Python, or Node.js. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners are accepted and removed before capture; newsletter popups and chat widgets are removed too. Each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing status in headers.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.
ScreenshotNeo is useful for producing clean captures without installing and maintaining browser binaries for that capture task. It does not replace Playwright’s committed baselines and assertions when you need to detect visual regressions in your own application. Sign up for 1,000 free screenshots a month, with no card required.
Frequently asked questions
Where do screenshot baselines live?
By default, Playwright stores screenshot snapshots in a directory beside the test file, typically named after the test file with -snapshots. Commit that directory so CI has the expected images.
Should I upload the baseline images as GitHub Actions artifacts?
Usually the baseline belongs in version control, where code review can inspect changes. Upload the HTML report and failure diagnostics as artifacts so maintainers can investigate a run.
Can I use screenshot assertions with more than Chromium?
Yes. Configure Playwright projects for the browsers you need, install their browser binaries in CI, and maintain baselines for each browser context you test. Renderings can differ between browsers and platforms.
What does the first screenshot test run do?
It creates the reference image because none exists yet. Review and commit that generated file, then subsequent runs compare against it.
Can a screenshot API replace a visual regression test?
An API can capture a page and return an image, but visual regression testing also needs a known baseline, comparison rules, and a reviewable failure diff. Use Playwright assertions when those checks are the goal.


