Run Website Screenshot Tests in Continuous Integration
Set up reliable Playwright screenshot comparisons in CI, manage baselines and diffs, and troubleshoot common failures.
Use Playwright Test’s built-in toHaveScreenshot() assertion to compare a rendered page with a saved baseline in continuous integration. For reliable results, make the page state and rendering environment repeatable, install the same browser dependencies in CI, start with one worker, and review every new baseline or image diff before accepting it. Screenshot comparisons catch visual changes; keep functional assertions for behavior.
1. Make the page reproducible
A visual test is only useful when the page renders in a predictable state. Before writing a screenshot assertion, decide what the test should capture and control the inputs that affect rendering.
- Navigation: use a stable URL and ensure the app is running before tests start.
- Viewport: set an explicit width and height for the test. Keep the viewport the same when creating and reviewing baselines.
- Test data: seed or mock data that changes the page, including dates, randomized content, and user-specific state.
- Readiness: wait for a meaningful page element or application-ready state. Avoid relying on an arbitrary delay when a selector or app signal can tell you the page is ready.
- External content: prevent ads, rotating recommendations, or other uncontrolled third-party content from changing the image.
- Fonts and assets: make sure fonts and images have loaded before capture. Missing fonts can change line wrapping and create widespread diffs.
Use functional assertions to confirm important page behavior and state. A screenshot assertion complements those checks; it does not prove that buttons, links, forms, or application logic work.
2. Add a Playwright screenshot test
Install Playwright Test using the project’s package manager, then create a test such as tests/homepage.spec.ts:
import { test, expect } from '@playwright/test';
test('homepage matches its visual baseline', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.goto('http://127.0.0.1:3000/');
await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
await expect(page).toHaveScreenshot('homepage.png');
});
Replace the URL and heading with ones from your app. The heading assertion checks that a meaningful piece of content appeared; use additional functional checks for the behavior the test is meant to protect.
On first use, Playwright creates a reference screenshot. Its documented capture process waits for two consecutive screenshots to match before saving the result. Inspect this initial baseline to confirm it represents the intended page and state. Do not accept it blindly. On later runs, Playwright compares the current rendering with that reference.
Run the test locally with npx playwright test. When the design intentionally changes, review the new rendering and diff before updating references with npx playwright test --update-snapshots. Commit accepted baseline changes together with the design change so reviewers can see both.
See the official Playwright visual comparisons documentation for assertion behavior and snapshot options.
3. Run the test in CI
The basic Playwright CI sequence is: install project packages from the lockfile, install the browsers and operating-system dependencies, then run the test command. The following GitHub Actions example assumes a Node project with a committed lockfile and a build script. Adapt the setup and server command to your application.
name: Visual tests
on:
pull_request:
push:
branches: [main]
jobs:
visual:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm
- run: npm ci
- run: npx playwright install --with-deps
- run: npm run build
- run: npx playwright test
env:
CI: true
This example uses GitHub Actions syntax. The install and test commands are the key steps and can be used with other CI providers. Pin the Node version and dependency versions as appropriate for your project. Playwright recommends one worker in CI as the default for stability and reproducibility; its documentation describes setting workers: process.env.CI ? 1 : undefined in Playwright configuration. A single worker can make a job slower, but it reduces variation from resource contention on shared runners.
If your tests navigate to a local application, start the server in CI before the test command. Playwright Test’s webServer configuration can start a command and wait for a URL to become available:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
workers: process.env.CI ? 1 : undefined,
use: {
baseURL: 'http://127.0.0.1:3000',
},
webServer: {
command: 'npm run start -- --host 127.0.0.1',
url: 'http://127.0.0.1:3000',
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
});
With that configuration, the test can navigate using await page.goto('/'). Make sure the command and readiness URL match how your app actually starts. If the build must run first, include it in your CI steps or in a project-specific startup command.
The official Playwright continuous integration guide covers CI setup, browser installation, workers, and running tests.
4. Keep baselines and rendering environments consistent
Screenshot references depend on more than application code. Operating system, browser version, fonts, rendering libraries, viewport, and device scale can all affect pixels. Create, compare, and review baselines in a consistent environment. A container can help standardize screenshot testing across operating systems; keep its browser and dependency versions aligned with the Playwright version used by the project.
Keep baseline files under version control so changes are reviewable alongside code. If a visual change is intended, update the references deliberately and include the changed images in the review. If CI reports a mismatch, inspect both the expected and actual images and the diff before deciding whether to fix the page or update the baseline.
5. Choose coverage and scaling deliberately
Begin with a small set of high-value pages and states: for example, a landing page, a key form, and a critical signed-in screen. Add coverage for meaningful viewport sizes or states where layout differs. Each additional page, viewport, and state adds browser work and baselines to maintain.
One worker is a stable starting point for CI. If the suite grows, teams with suitable runner capacity can increase workers or shard tests across CI jobs. Sharding can reduce elapsed time, but it uses more concurrent CI capacity and requires a way to inspect results across jobs. Store Playwright reports and failure evidence as CI artifacts when they help reviewers diagnose a failed run; choose retention based on your team’s workflow and the data captured in screenshots.
Visual assertions compare rendered pixels, so tiny differences may arise from rendering or timing changes even when the page is functionally correct. Keep the capture state deterministic, review diffs, and tune comparison settings only when you understand which differences should be tolerated. Avoid loosening thresholds just to make a flaky test pass.
6. Built-in snapshots or hosted review
| Approach | Where references and review live | Operational considerations |
|---|---|---|
| Playwright built-in assertions | Reference images are managed in the Playwright test workflow and repository. | Requires baseline review and consistent rendering environments. No external snapshot service is needed for the basic workflow. |
| Percy integration | Playwright snapshots can be uploaded to Percy for hosted review. | Uses a service, an account, and a project token workflow. Review service fit, token handling, data access, and current plan terms independently. |
Percy’s Playwright integration documentation describes running snapshots through percy exec with a project token. A hosted review flow may suit teams that want snapshot review outside their source repository; the built-in workflow keeps comparison in Playwright’s existing test and baseline process. Decide based on review needs, data handling, and operational requirements. Verify current commercial terms directly with the service.
7. Troubleshoot common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Browser executable or system library is missing | The CI job installed the package but not its browser binaries or OS dependencies. | Run npx playwright install --with-deps in the job and make sure it uses the same Playwright package version as the project. |
| Navigation fails with connection refused | The app server was not started, is listening on a different address or port, or was not ready before the test. | Start the server in CI and configure Playwright’s webServer URL and command to match it. Bind the server to an address reachable from the test process. |
| Large or repeated image diffs appear unexpectedly | Rendering environment, browser version, fonts, viewport, device scale, test data, or page readiness differs from baseline creation. | Compare environment and inputs, wait for required content and assets, then recreate the baseline only if the new rendering is intended. |
| Only a changing region fails | The page includes a timestamp, rotating content, animation, or other dynamic element. | Make the test data deterministic or remove the changing source from the test state. Use Playwright’s documented screenshot options only when they fit the intended assertion. |
| Test passes locally but fails in CI | The local and CI operating systems, fonts, browser dependencies, resource limits, or configuration differ. | Use a consistent container or runner environment, install dependencies explicitly, and run with one worker while diagnosing. |
| Visual test is slow or unstable under parallel execution | Workers compete for CPU or memory on a shared runner, or the app has shared mutable state. | Start with one CI worker. Add parallelism or shard only after checking runner capacity and test isolation. |
| Updated references produce noisy review | Many baselines changed because a shared font, browser, or layout changed. | Review representative diffs, identify the shared cause, and separate broad intentional updates from unrelated changes where practical. |
8. Performance, reliability, and cost
CI time grows with the number of test cases, viewport variants, browser projects, and screenshot captures. Keep the suite focused on pages and states where visual regressions matter. One worker favors repeatability; additional workers or shards can shorten elapsed time when the runner has capacity, while increasing resource use and coordination.
Reliability depends on controlled state and matching environments. Keep browser installation tied to the project’s Playwright version, avoid unstable external page content, and treat baseline updates as reviewed code changes. A passing pixel comparison is one signal, not a substitute for functional coverage or accessibility checks.
Built-in Playwright comparisons do not require a hosted snapshot service. CI still consumes runner time and storage for artifacts. Hosted review adds an external service and token workflow; check its current pricing, retention, access controls, and data handling before adopting it. The referenced technical materials do not establish current service pricing.
9. Or skip the browser setup
If you need a rendered screenshot from an API or an AI agent rather than a repository-managed visual regression baseline, ScreenshotNeo is a website screenshot API and MCP server. Its API returns an image or PDF from one GET request. This does not replace Playwright’s baseline comparison workflow; it is an option for capturing pages without setting up browser automation in your own job.
For a complete CI visual regression test, Playwright should capture the page in your controlled test environment and compare it with a reviewed baseline. For a standalone capture, the cURL, Python, and Node.js calls below request a screenshot of Stripe. The ScreenshotNeo API documentation describes its request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(async fs => {
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
});
ScreenshotNeo accepts cookie banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
10. FAQ
Do screenshot tests replace end-to-end tests?
No. They show whether a rendered result changed. Keep assertions for navigation, interactions, data, and other behavior your application must provide.
Should every page have a visual test?
No. Prioritize pages and states where a visual regression would matter, then expand when the coverage is useful to maintain.
Can different operating systems share the same baselines?
Rendering can differ across operating systems and dependencies. Keep baseline creation and comparison on a consistent environment, or manage separate references when your supported environments need separate coverage.


