Playwright Screenshot Tests Fail on CI but Pass Locally: Fixes
Compare the CI diff, align its browser environment with your baseline, and control dynamic page state before changing thresholds or snapshots.
If a Playwright screenshot test passes locally but fails on CI, first inspect the expected, actual, and diff images and the failed run’s trace. Then compare the environments that produced them and make the page state deterministic. A local macOS or Windows baseline compared with a Linux CI render is a common possibility, but the failure itself does not identify its cause. Playwright warns that rendering can vary with the host OS, browser version, settings, hardware, power source, and headless mode. For repeatable comparisons, generate and check baselines in the same environment. Playwright visual comparisons.
1. Read the failure before changing the test
Download the expected image, actual image, and diff from the CI report or artifacts. Look at the shape and location of the mismatch:
- Broad changes across text and layout: compare operating system, fonts, browser build, viewport, and device scale.
- A small changing region: check timestamps, rotating content, avatars, ads, counters, and other live data.
- A substantially different page: check whether navigation, loading, authentication, or test data differed.
- No screenshot diff because the browser or test did not run: diagnose a launch, dependency, timeout, or test-state failure instead.
Keep the artifacts from the failing run. Regenerating a baseline before understanding the difference can turn an unexplained failure into an accepted change.
2. Compare the local and CI rendering environments
Record these values for both runs. OS, browser version, hardware, and headless mode are documented sources of rendering variation; viewport, scale, fonts, locale, test data, and network responses are useful additional diagnostic variables.
| Check | What to compare | Typical next step |
|---|---|---|
| Operating system and runner | Local OS versus CI OS and container image | Generate and compare baselines in the same OS/container. |
| Playwright and browser | Installed package version, browser project, and browser build | Install browsers using the project’s Playwright CLI; align container tags with the package version. |
| Display settings | Viewport, device scale factor, headless/headed mode | Set viewport and device scale explicitly; keep the mode consistent. |
| Fonts and content | Installed fonts, font loading, locale, timezone, and test data | Install required fonts and use controlled data where output depends on them. |
| Network and dependencies | Responses, third-party resources, and page readiness | Stabilize test data and wait for the relevant page state rather than an arbitrary delay. |
Playwright snapshot names include browser and platform information by default. If you intentionally test multiple operating systems or browser projects, keep project-specific baselines rather than comparing renders from unlike environments. See snapshot configuration and test configuration.
3. Reproduce the CI setup consistently
Install the locked Node dependencies and the browsers and OS dependencies expected by that Playwright version. For Linux CI, Playwright documents this basic setup:
npm ci
npx playwright install --with-deps
npx playwright test
A container can make the operating system and browser environment more consistent. Pin the Playwright container image to a tag aligned with the package version installed by the project; do not assume a floating or mismatched image has the same browser build.
Example GitHub Actions workflow:
name: Playwright tests
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
- uses: actions/upload-artifact@v5
if: ${{ !cancelled() }}
with:
name: playwright-report
path: playwright-report/
retention-days: 30
Use the action versions and Node version appropriate for your repository and CI policy. The install and artifact pattern follows the official Playwright CI guide.
4. Make the page state deterministic
toHaveScreenshot() waits for two consecutive screenshots to match before comparing the last capture to the baseline. It disables CSS animations, transitions, and Web Animations by default. Finite animations are fast-forwarded; infinite animations are canceled for capture and resumed afterward. These defaults help, but they cannot make changing data or unstable responses deterministic. See the API reference.
Use controlled fixtures or mock responses for data that changes between runs. Wait for the user-visible state the screenshot depends on. Mask only known, intentionally variable regions, or use a screenshot stylesheet to hide them. Avoid masking large areas: that can conceal a real regression.
import { test, expect } from '@playwright/test';
test('account page matches its visual baseline', async ({ page }) => {
await page.goto('/account');
await expect(page.getByRole('heading', { name: 'Account' })).toBeVisible();
await expect(page.locator('[data-testid="account-summary"]')).toBeVisible();
await expect(page).toHaveScreenshot('account.png', {
fullPage: true,
animations: 'disabled',
caret: 'hide',
scale: 'css',
mask: [page.locator('[data-testid="live-clock"]')],
maskColor: '#777777',
});
});
This is runnable in a project using @playwright/test with an /account route, an “Account” heading, and the indicated test IDs. Remove or adapt those selectors to match your application. If a changing area should not appear at all, a stylesheet is another option:
/* tests/visual.css */
[data-testid="live-clock"] {
visibility: hidden !important;
}
await expect(page).toHaveScreenshot('account.png', {
stylePath: 'tests/visual.css',
});
Other useful options include fullPage for the complete scrollable page; clip for a specified capture rectangle; scale: 'css' for one image pixel per CSS pixel, or 'device' for device pixels; omitBackground for transparency (not JPEG); and mask/maskColor for variable locators. Screenshot assertion options include timeout, threshold, maxDiffPixels, and maxDiffPixelRatio. Verify option availability against the API docs for the installed Playwright version.
5. Capture traces and CI artifacts
A screenshot shows what differed; a trace can help explain why. Playwright recommends Trace Viewer for CI failures because it exposes the action timeline, DOM snapshots, and network requests. Configure tracing on retry so failed tests leave diagnostic evidence without recording every successful test. The exact setting is part of Playwright Test configuration; consult the CI debugging guidance.
npx playwright test --trace on
This records traces for each test in that run and is useful for a focused reproduction. Open them through the HTML report. For browser launch failures, set DEBUG=pw:browser to see launch logs:
DEBUG=pw:browser npx playwright test
Upload the HTML report, traces, and expected/actual/diff images as CI artifacts, including when a job fails. This lets a developer inspect the failure without trying to recreate the original runner locally.
6. Use stable CI settings without assuming the cause
Playwright recommends one worker in CI to prioritize stability and reproducibility. This can reduce resource contention, but it does not prove parallelism caused a particular pixel diff. Teams with powerful self-hosted runners can enable parallel tests; sharding across jobs is another documented way to scale. A starting configuration is:
import { defineConfig } from '@playwright/test';
export default defineConfig({
workers: process.env.CI ? 1 : undefined,
retries: process.env.CI ? 1 : 0,
use: {
trace: 'on-first-retry',
viewport: { width: 1280, height: 720 },
deviceScaleFactor: 1,
},
});
Keep only settings appropriate to the project: changing a viewport or device scale can change the expected image and may require reviewed baselines. Retries can collect evidence and reveal intermittent behavior, but a test that passes only on retry still deserves investigation.
Browser binary caching is generally not recommended in Playwright’s CI guide: restoring the cache can take as long as downloading, and Linux OS dependencies are not cacheable. If you cache browser binaries anyway, key the cache to the Playwright version and install required OS dependencies separately. See caching browsers.
7. Treat diff thresholds and baseline updates as reviewed changes
threshold allows a perceived color difference; maxDiffPixels and maxDiffPixelRatio allow a number or proportion of differing pixels. They change what the test accepts. Use them only when the difference is understood and acceptable, and scope them to the specific assertion where possible. Raising a global limit to silence an unexplained failure can hide real changes. See the visual comparison options.
When a visual change is intentional, review the new image and update the baseline deliberately:
npx playwright test --update-snapshots
Review the resulting snapshot changes and commit them with the code change that explains them. Do not have CI silently rewrite expected images to make a failing job pass.
8. Troubleshooting common CI failures
| Symptom | Likely cause to investigate | Fix |
|---|---|---|
| Text wraps differently everywhere | Different OS, fonts, browser build, viewport, or scale | Align the baseline and CI environment; install needed fonts; set viewport and scale. |
| Only a timestamp, avatar, or live tile differs | Dynamic content or uncontrolled data | Use fixed test data, mock the response, or narrowly mask/style the changing locator. |
| The page is blank or partially loaded | Navigation, readiness, authentication, or network response differs | Inspect the trace and network; assert the expected content before capturing. |
| Failure appears only under load | Resource contention or test interference may be involved | Try one CI worker, inspect shared state, and compare traces. Treat worker count as a diagnostic, not a proven root cause. |
browserType.launch fails |
Missing browser binary or Linux system dependency, or a mismatched installation | Run npx playwright install --with-deps and inspect DEBUG=pw:browser output. |
| Snapshot missing or unexpected platform snapshot | Baseline was not committed or the run uses a different project/platform snapshot path | Check snapshot artifacts and configured path template; create and review the intended project-specific baseline. |
| Test passes only after retry | Intermittent page state, network behavior, or runner pressure | Use the retry trace to find the divergence; stabilize the relevant state rather than relying on retries. |
| Many unrelated diffs after dependency change | Playwright or browser version changed | Confirm the lockfile and installed browser version; review whether a deliberate baseline refresh is needed. |
9. Performance, reliability, and cost considerations
Visual tests capture and compare image data, and full-page captures cover more content than viewport captures. Keep assertions focused on meaningful pages or components. A single CI worker favors reproducibility but can increase total suite time; use sharding or carefully measured parallelism when throughput matters. Tracing every test can add overhead, so capture on retry or enable full tracing for a focused debugging run. Browser caching may not save time, and cached binaries do not replace Linux system dependencies. These are operational tradeoffs; the dossier provides no benchmark for a particular project.
Playwright is an open-source browser automation framework; CI cost depends on the CI provider, runner size, job duration, and concurrency. No particular provider price or savings can be inferred here. The most reliable first investment is preserving failure artifacts and aligning the execution environment, so each failure can be diagnosed.
Or skip the browser setup
If your immediate need is a rendered website image rather than a Playwright visual assertion, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. It is not a replacement for Playwright’s baseline assertions, but it can avoid maintaining browser capture code for screenshot generation.
Install Python’s requests package with python -m pip install requests, set YOUR_API_KEY, and run:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
See the ScreenshotNeo API docs for request options. Equivalent calls:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with response headers indicating page verdict and billing. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. This API generates captures; keep Playwright for checking your application against visual baselines.
Get 1,000 free screenshots a month with no card.
FAQ
Should I commit screenshots generated on my laptop?
Commit reviewed baselines produced in the same environment used for comparison, or maintain distinct baselines for each intended platform/project.
Does toHaveScreenshot() already wait for the page to settle?
It waits for consecutive screenshot captures to match and disables animations by default. It does not make external data, network responses, or application state deterministic.
Should I increase maxDiffPixels until CI passes?
Only after inspecting and accepting the difference. A larger allowance changes the assertion’s sensitivity; it does not explain the mismatch.
Can ScreenshotNeo diagnose my Playwright diff?
No. It can capture a website through an API or MCP server, while Playwright traces and screenshot assertions are the tools for diagnosing and validating your test run.


