How to Compare Website Screenshots in a GitLab CI Pipeline
Capture a stable baseline and candidate screenshot, fail GitLab CI when visual differences exceed your threshold, and save useful evidence for review.
To compare website screenshots in a GitLab CI pipeline, capture a reference image and a candidate image under the same browser and page conditions, run an image comparator that exits nonzero when the difference exceeds your chosen threshold, and upload both images plus the diff as job artifacts. GitLab runs the job and stores its evidence; it does not perform pixel-level screenshot comparison for you.
This guide uses Playwright to capture screenshots and a small Node.js script with pngjs to calculate changed pixels. The baseline is a screenshot committed to the repository. The CI job fails when the changed-pixel ratio exceeds the limit you set.
1. Set up the baseline and candidate capture
First, make a baseline from the page state you intend to protect. Use a stable local environment and data. Save the baseline at visual-baselines/home.png and commit it. On each pipeline run, the job captures the candidate into screenshots/current.png.
Use the same route, browser build, viewport, device scale factor, locale, timezone, fonts, and application data in both captures. Wait for the page to reach the state that matters. Disable animations and hide genuinely dynamic content such as a clock if it would create irrelevant pixel changes. Do not mask areas simply because they are inconvenient: a mask can conceal real regressions.
Install the dependencies in your project:
npm install --save-dev playwright pngjs
npx playwright install chromium
Create scripts/capture.mjs:
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
const url = process.env.VISUAL_URL ?? 'http://127.0.0.1:4173/';
const output = process.env.SCREENSHOT_PATH ?? 'screenshots/current.png';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
locale: 'en-US',
timezoneId: 'UTC',
colorScheme: 'light',
});
await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
await page.evaluate(() => document.fonts.ready);
await page.addStyleTag({ content: `
*, *::before, *::after {
animation: none !important;
transition: none !important;
caret-color: transparent !important;
}
` });
await mkdir(output.substring(0, output.lastIndexOf('/')) || '.', { recursive: true });
await page.screenshot({ path: output, fullPage: true, animations: 'disabled' });
} finally {
await browser.close();
}
This example captures the full page. For a viewport-only check, remove fullPage: true. If networkidle never occurs because the application uses long polling or analytics, wait for a specific selector that marks the page ready instead, or use an explicit delay as a last resort. Capture the same page state for the baseline and candidate.
2. Compare the images and set a failure threshold
Create scripts/compare.mjs. The comparator below checks dimensions and the fraction of pixels that differ. CHANNEL_TOLERANCE allows small per-channel variation before a pixel is counted as changed; MAX_CHANGED_RATIO sets the maximum allowed changed-pixel fraction. It also writes a red-overlay diff image for review.
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import { PNG } from 'pngjs';
const baselinePath = process.env.BASELINE_PATH ?? 'visual-baselines/home.png';
const currentPath = process.env.CURRENT_PATH ?? 'screenshots/current.png';
const diffPath = process.env.DIFF_PATH ?? 'visual-diff/diff.png';
const tolerance = Number(process.env.CHANNEL_TOLERANCE ?? 16);
const maxRatio = Number(process.env.MAX_CHANGED_RATIO ?? 0.001);
if (!Number.isFinite(tolerance) || tolerance < 0 || tolerance > 255) {
throw new Error('CHANNEL_TOLERANCE must be a number from 0 to 255');
}
if (!Number.isFinite(maxRatio) || maxRatio < 0 || maxRatio > 1) {
throw new Error('MAX_CHANGED_RATIO must be a number from 0 to 1');
}
const baseline = PNG.sync.read(await readFile(baselinePath));
const current = PNG.sync.read(await readFile(currentPath));
if (baseline.width !== current.width || baseline.height !== current.height) {
console.error(`Image dimensions differ: baseline ${baseline.width}x${baseline.height}, candidate ${current.width}x${current.height}`);
process.exit(1);
}
const diff = new PNG({ width: baseline.width, height: baseline.height });
let changed = 0;
for (let i = 0; i < baseline.data.length; i += 4) {
const differs = Math.abs(baseline.data[i] - current.data[i]) > tolerance
|| Math.abs(baseline.data[i + 1] - current.data[i + 1]) > tolerance
|| Math.abs(baseline.data[i + 2] - current.data[i + 2]) > tolerance
|| Math.abs(baseline.data[i + 3] - current.data[i + 3]) > tolerance;
if (differs) {
changed++;
diff.data[i] = 255;
diff.data[i + 1] = 0;
diff.data[i + 2] = 0;
diff.data[i + 3] = 255;
} else {
diff.data[i] = current.data[i];
diff.data[i + 1] = current.data[i + 1];
diff.data[i + 2] = current.data[i + 2];
diff.data[i + 3] = 90;
}
}
const pixels = baseline.width * baseline.height;
const ratio = changed / pixels;
await mkdir(diffPath.substring(0, diffPath.lastIndexOf('/')) || '.', { recursive: true });
await writeFile(diffPath, PNG.sync.write(diff));
console.log(`Changed pixels: ${changed}/${pixels} (${(ratio * 100).toFixed(4)}%); allowed ${(maxRatio * 100).toFixed(4)}%`);
if (ratio > maxRatio) process.exit(1);
This is a deliberately simple comparator, not a perceptual image metric. Alpha handling, color profiles, browser rasterization, font rendering, and anti-aliasing can affect pixel results. Tune the tolerance and ratio using representative changes from your own pages, inspect the diff, and keep the threshold strict enough to catch defects. A permissive threshold can hide small but important changes.
3. Add package scripts and capture a baseline
Add scripts to your package.json, adapting the app’s own build and start commands. The commands below assume the built site listens on port 4173 and has a start:test script. Keep the server running until capture finishes.
{
"scripts": {
"start:test": "your-existing-test-server-command",
"screenshots:capture": "node scripts/capture.mjs",
"screenshots:compare": "node scripts/compare.mjs"
}
}
Replace the placeholder server command with your actual command, such as the preview command for your framework. Start the app, wait until it is ready, capture a candidate, review it, then copy the approved image to the baseline path and commit it:
mkdir -p visual-baselines
SCREENSHOT_PATH=visual-baselines/home.png npm run screenshots:capture
git add visual-baselines/home.png
git commit -m "Add homepage visual baseline"
For a site with multiple important routes, capture one baseline per route and compare each candidate against its matching baseline. Keep baseline updates reviewable in merge requests so an expected design change does not silently normalize an unintended one.
4. Run the visual check in GitLab CI
GitLab jobs are configured in .gitlab-ci.yml. This example runs the project-defined build and start commands, waits for the local server, captures and compares the page, and saves diagnostic images even when the comparator fails. The server readiness loop is important: a background process starting does not mean the site is ready to capture.
stages:
- test
visual-regression:
image: node:22
stage: test
variables:
VISUAL_URL: "http://127.0.0.1:4173/"
MAX_CHANGED_RATIO: "0.001"
CHANNEL_TOLERANCE: "16"
script:
- npm ci
- npx playwright install --with-deps chromium
- npm run build
- npm run start:test > server.log 2>&1 &
- SERVER_PID=$!
- |
ready=0
for attempt in $(seq 1 60); do
if curl --fail --silent "$VISUAL_URL" > /dev/null; then
ready=1
break
fi
sleep 1
done
if [ "$ready" -ne 1 ]; then
cat server.log
exit 1
fi
- npm run screenshots:capture
- npm run screenshots:compare
artifacts:
when: always
expire_in: 1 week
paths:
- visual-baselines/
- screenshots/
- visual-diff/
- server.log
The example’s image tag and install command are illustrative configuration; choose a Node and browser setup supported by your runner and keep browser dependencies consistent. The capture and compare script names, server command, port, and output paths are project-defined. For an application that needs a database or service containers, start those dependencies and seed deterministic data before capture. A real pipeline may split build, serve, capture, and comparison work into jobs; pass generated files between jobs as artifacts where necessary. GitLab documents stages, jobs, and dependencies in its [pipeline quick start](https://docs.gitlab.com/ci/quick_start/) and [pipeline architecture guide](https://docs.gitlab.com/ci/pipelines/pipeline_architectures/).
5. Make failures useful in merge requests
The comparator’s nonzero exit status is what fails this job. Merely uploading a diff or publishing a test report does not make a failed visual check. The example sets artifacts:when: always because artifacts upload only on success by default. Artifact paths are relative to the repository checkout; expire_in controls the configured lifetime, while GitLab can retain artifacts from the latest pipelines on a ref beyond that interval. Restrict downloads with artifacts:access if needed. See [GitLab job artifacts](https://docs.gitlab.com/ci/jobs/job_artifacts/).
To link an image to a JUnit test detail, include the screenshot path in that testcase’s system-out and upload the file as an artifact:
<testcase classname="VisualRegression" name="homepage" time="1.2">
<failure>Changed pixel ratio exceeded configured threshold</failure>
<system-out>[[ATTACHMENT|visual-diff/diff.png]]</system-out>
</testcase>
Configure the job’s artifacts:reports:junit to point to the generated XML and include the screenshot in artifacts:paths. GitLab can display JUnit results and attachments in merge request test details, but JUnit reports do not set job status; the script must still exit nonzero when the visual check fails. See [GitLab unit test reports](https://docs.gitlab.com/ci/testing/unit_test_reports/).
GitLab’s browser performance feature uses sitespeed.io to compare performance metrics between branches. It reports performance measurements; it is not a pixel-level screenshot comparator. See [GitLab browser performance testing](https://docs.gitlab.com/ci/testing/browser_performance_testing/).
6. Choose a comparison approach
The example’s local script is one option. When evaluating another comparator, compare setup and maintenance effort, browser and framework support, pixel versus perceptual behavior, tolerance and masking controls, clarity of failure diffs, runtime and CI resource needs, and whether images stay in your own CI or are sent to a hosted service. Choose based on your page types and review process; a tool’s report only helps if its failure status blocks the pipeline when intended.
Keep baselines in version control when you want changes reviewed alongside code. A hosted comparison service may centralize baselines or reporting, but introduces its own integration and data-handling decisions. Whatever you choose, preserve enough image evidence to explain a failure and make threshold updates reviewable.
7. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Connection refused or blank capture | The app server has not started or is listening on another host/port. | Wait for a successful readiness request, check server.log, and set VISUAL_URL to the actual address reachable from the job. |
| Navigation times out | The page never reaches network idle due to long-lived requests, or a dependency is unavailable. | Wait for a stable page selector instead of network idle, verify services and test data, and set a timeout appropriate to the environment. |
| Every run differs slightly | Fonts, animations, locale, time, remote content, or application data is unstable. | Pin the browser environment, wait for fonts and app readiness, fix locale/timezone/data, disable animations, and mask only known dynamic regions. |
| All pixels differ or dimensions mismatch | Viewport, device scale factor, full-page mode, route, or baseline version changed. | Match capture settings and confirm the candidate is compared to the correct route’s baseline. Review intentional layout changes before updating the baseline. |
| The job passes despite a large visual change | The configured changed-pixel ratio is too permissive, or the comparator script does not propagate failure. | Lower MAX_CHANGED_RATIO, inspect the threshold logic, and confirm the command returns exit code 1 for a deliberate test change. |
| The job fails but artifacts are missing | Artifacts use the default success-only upload behavior or paths do not match generated files. | Set artifacts:when: always and check paths relative to the repository directory. See [artifact upload conditions](https://docs.gitlab.com/ci/jobs/job_artifacts/). |
| JUnit attachment is not shown | The path is not in testcase-level system-out, the file is not uploaded, or the XML is invalid. |
Use a repository-relative attachment path, upload the image as an artifact, and validate the JUnit report configuration. |
| Browser installation fails in the runner | Required system libraries are missing or the runner cannot download browser packages. | Use Playwright’s dependency installation in a compatible Linux runner, or use a container image with matching browser dependencies already present. |
8. Performance, reliability, and cost
Each route and viewport adds browser startup, page load, image storage, and comparison time. Start with routes that represent important user journeys, then expand as the signal proves useful. Reuse installed dependencies through the CI cache when appropriate; use artifacts for generated screenshots and diffs that reviewers need. Caches are for reusable dependencies, while artifacts preserve job outputs and can be passed to later jobs.
Stability matters more than raw capture speed. Prefer local fixtures or controlled test data over remote content that changes outside your code. Use the same capture environment for baseline refreshes and pipeline runs. Set artifact expiry to fit the review window and project storage policy. Hosted comparison services can add service costs; a local script consumes your runner time and storage. Review screenshots for credentials, private data, or personal information before uploading, and restrict artifact access where necessary. GitLab warns against saving tokens, passwords, or other sensitive information in artifacts; see [debugging CI/CD pipelines](https://docs.gitlab.com/ci/jobs/job_artifacts/).
Or skip the browser setup
[ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server. You can use its API to generate candidate screenshots, then compare them with your versioned baseline using the same kind of local comparator shown above. It does not replace the baseline comparison or decide your pipeline threshold. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. [Create a free account](https://screenshotneo.com/account/sign-up/).
FAQ
Does GitLab compare screenshots automatically?
No. Your project must run a comparator or use a separate service. GitLab provides jobs, reports, and artifact storage.
Can a JUnit report fail the pipeline?
The report itself does not set job status. Make the comparison command return a nonzero exit status when the change exceeds your policy.
Should every pixel difference fail?
That depends on the page and browser stability. Tune channel tolerance and changed-pixel ratio against real expected and unintended changes, and review the diff before accepting threshold changes.


