ScreenshotNeo

BlogHow-to

How to Run Playwright Screenshot Tests in GitLab CI

Set up reproducible Playwright screenshot tests in GitLab CI with a compatible browser image, stable baselines, failure artifacts, and optional sharding.

By the ScreenshotNeo team4 October 20267 min read

Run Playwright screenshot tests in GitLab CI with Playwright Test, a container image whose browser version matches your project, and npx playwright test. Keep the baseline and CI environments consistent, commit screenshot references, and save the HTML report and test output as GitLab artifacts so failures are reviewable.

1. Add a visual comparison test

Use Playwright Test’s toHaveScreenshot() assertion after navigating to the page and waiting for the state you want to compare. On its first run, Playwright creates the reference image; later runs compare against it.

import { test, expect } from '@playwright/test';

test('homepage visual appearance', async ({ page }) => {
  await page.goto('/');
  await expect(page).toHaveScreenshot('homepage.png');
});

For this to run locally, configure the base URL or use an absolute URL. For example, in playwright.config.ts:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  use: {
    baseURL: 'http://127.0.0.1:3000',
  },
  webServer: {
    command: 'npm run start -- --host 0.0.0.0',
    url: 'http://127.0.0.1:3000',
    reuseExistingServer: !process.env.CI,
  },
});

Adapt the start command and readiness URL to your application. If the app is deployed separately for the pipeline, omit webServer and point baseURL at that deployment.

2. Run the tests in GitLab CI

The following npm job uses the Playwright container image shown in Playwright’s GitLab CI example. The tag is version-specific: verify the current documented tag and align it with the Playwright version in your lockfile before adopting it. Playwright’s GitLab CI guide explains the image and installation choices.

stages:
  - test

playwright-screenshots:
  stage: test
  image: mcr.microsoft.com/playwright:v1.63.0-noble
  variables:
    CI: "true"
  script:
    - npm ci
    - npx playwright test
  artifacts:
    when: always
    paths:
      - playwright-report/
      - test-results/
    expire_in: 1 week

npm ci installs the locked dependencies. For another package manager, use its lockfile-respecting install command. This job assumes your Playwright config writes its report and test output to the listed directories; artifact paths are relative to $CI_PROJECT_DIR.

Configure those directories and use one worker in CI as a stable starting point:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  reporter: [['html', { outputFolder: 'playwright-report', open: 'never' }]],
  outputDir: 'test-results',
  workers: process.env.CI ? 1 : undefined,
  use: {
    trace: 'on-first-retry',
  },
});

GitLab’s when: always setting retains artifacts after a test failure. GitLab does not upload artifacts when a job times out, so set a practical job timeout and investigate timeouts rather than relying on artifacts to diagnose them. See the GitLab artifacts reference.

3. Keep screenshot baselines reproducible

Visual diffs are meaningful only when rendering conditions are controlled. Playwright notes that screenshots can vary with operating system, browser version, browser settings, hardware, power source, and headless mode. Generate baselines in the same environment used in CI where possible. The Playwright visual comparisons guide recommends running in the same environment as the baseline generation.

  • Use the same Playwright package and browser image version for baseline creation and CI.
  • Prefer a consistent container and browser configuration across developer and CI workflows.
  • Commit snapshot files alongside the tests and review image diffs in code review.
  • Control app data, timestamps, rotating banners, and embedded content so the page reaches a stable state.

For intentionally volatile regions, Playwright supports a stylesheet via stylePath. For example, hide a cursor or timestamp during a screenshot:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  expect: {
    toHaveScreenshot: {
      stylePath: './tests/visual-stability.css',
    },
  },
});
/* tests/visual-stability.css */
.test-only-timestamp,
.blinking-cursor {
  visibility: hidden !important;
}

Prefer stabilizing the underlying test data when that is practical. A stylesheet should target only the content that cannot be made deterministic.

4. Create and update reference images intentionally

The first run produces snapshots. Inspect them, then commit the generated reference files in the snapshot directories with the test change. When a UI change is expected, regenerate deliberately:

npx playwright test --update-snapshots

Review the resulting images before accepting them. Do not make routine CI failures auto-update their own expected snapshots: that would replace the comparison baseline rather than tell you whether the rendered page changed unexpectedly.

5. Install browsers only when needed

The Playwright container image is intended to provide browsers and system dependencies. If you use a different job image or environment without the needed browser binaries and operating system packages, Playwright’s CI instructions show this after dependency installation:

npm ci
npx playwright install --with-deps
npx playwright test

Choose one coherent environment setup. Check that the browser binaries correspond to the installed Playwright package; adding an install step blindly to a prebuilt image can add unnecessary work or version mismatches.

6. Scale the suite with GitLab sharding

Start with one worker for stable, reproducible CI runs. If the suite is too slow and runners are available, split it across GitLab jobs with Playwright sharding:

playwright-screenshots:
  stage: test
  image: mcr.microsoft.com/playwright:v1.63.0-noble
  parallel: 4
  variables:
    CI: "true"
  script:
    - npm ci
    - npx playwright test --shard=$CI_NODE_INDEX/$CI_NODE_TOTAL
  artifacts:
    when: always
    paths:
      - playwright-report/
      - test-results/
    expire_in: 1 week

GitLab provides CI_NODE_INDEX and CI_NODE_TOTAL for parallel jobs. Ensure that sharded jobs have compatible report handling and do not overwrite shared outputs. More shards can reduce elapsed time only when the runner capacity is available; they also use more concurrent CI resources.

7. Cache dependencies with care

Cache package-manager downloads when that helps your pipeline, using a key tied to the lockfile so dependency changes invalidate the cache. Playwright does not recommend caching browser binaries by default: restoring them can take about as long as downloading them, and Linux system dependencies cannot be cached. If you choose to cache browser binaries, key that cache to the Playwright version so it cannot silently pair stale browsers with a changed package.

8. Troubleshoot common failures

Symptom Likely cause Fix
Browser executable is missing or fails to launch The job image and Playwright package do not match, or the environment lacks browser binaries or system dependencies. Align the container tag with the installed Playwright version. If the environment requires it, run npx playwright install --with-deps after installing dependencies. For launch diagnostics, run DEBUG=pw:browser npx playwright test.
Many visual tests fail only in CI The baseline was created on a different OS, browser, or rendering configuration, or the page contains dynamic data. Regenerate baselines in the CI-compatible environment and stabilize app data, fonts, and volatile page regions before updating snapshots.
Screenshot assertion fails intermittently The capture happens before the intended page state is stable, or content changes between runs. Wait for an app-specific readiness signal or required selector; control test data and use a narrowly scoped stylePath for unavoidable volatile regions.
HTML report or trace is missing after a failure Reporter/output paths differ from artifact paths, or the job timed out before artifact upload. Make outputFolder and outputDir match artifacts.paths. Keep when: always; investigate timeout duration because timed-out jobs do not upload artifacts.
CI job is slow after adding browser installation Browsers may already be present in the selected image, or downloads repeat each run. Use the compatible Playwright image or install only what is missing. Cache lockfile-based package downloads first; measure whether browser caching helps before keeping it.
Shards produce confusing or conflicting outputs Parallel jobs may write to overlapping locations or reports may be aggregated incorrectly. Use shard-aware output/report handling and separate job artifacts where needed; confirm each job’s evidence is retained.

9. Cost and reliability choices

The main CI trade-offs are runner minutes, concurrency, and the time spent debugging non-reproducible failures. A single worker is the simpler stability baseline. Sharding trades more concurrent runner capacity for shorter wall-clock time. Package caches can reduce repeated dependency downloads, while browser caches may cost as much time to restore as to download. Artifacts consume storage according to the report volume and retention period, so keep only the evidence your team needs and choose an expiry that supports review.

For reliability, use a lockfile, pin compatible Playwright tooling, keep baselines reviewed in version control, and retain failure evidence. Treat a visual diff as a signal to inspect the page and test environment; accept a new baseline only after deciding the rendering change is intended.

Or skip the browser setup

If you need a screenshot of a live page outside your regression suite, ScreenshotNeo offers a website screenshot API and MCP server. Its API accepts one GET request and returns an image or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace the example URL with the page you need and keep the API key private. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These captures are useful for page capture workflows; Playwright’s test runner remains the workflow for asserting that your own application matches committed visual baselines.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Should visual snapshots be committed?

Yes. They are the expected images future runs compare against, so commit and review them with the tests.

Can I generate snapshots automatically in CI?

The initial run can create them, but inspect and commit them intentionally. CI should compare against reviewed references rather than silently replacing them.

Should I run screenshots on every shard?

Yes, if the test suite is sharded, each shard should run its assigned tests against the same application state and compatible rendering environment.