How to Compare Puppeteer Screenshots for Visual Regression Testing
Capture repeatable Puppeteer screenshots, compare them with reviewed image baselines, and diagnose visual differences without masking real regressions.
To compare Puppeteer screenshots, use Puppeteer to capture a repeatable PNG and an image matcher to compare it with a reviewed baseline. Puppeteer captures the page; it does not decide whether two screenshots match. In a Jest project, jest-image-snapshot provides baseline storage and a matcher. For a custom test runner, Pixelmatch compares decoded pixel arrays and can produce a diff image.
A reliable test controls the browser and page state, captures the same region at the same dimensions, compares the result with a baseline, and makes a person review any proposed baseline update. Puppeteer’s screenshot API returns image bytes by default. [c001] [c002]
1. Choose how to compare screenshots
| Approach | Use it when | What you manage |
|---|---|---|
jest-image-snapshot |
You already run Jest and want an image matcher with baseline and diff artifacts. | Reviewing and committing baselines; choosing comparison and failure thresholds. |
| Pixelmatch directly | You want to integrate pixel comparison into another Node.js test runner or script. | Image decoding, baseline lifecycle, artifact output, and failure reporting. |
| Playwright screenshot assertions | You are open to using Playwright’s runner rather than Puppeteer. | Adopting another framework. It is an alternative framework path, not a Puppeteer comparator. |
Start with pixel comparison if you need a simple, inspectable diff. jest-image-snapshot uses Pixelmatch by default and also offers SSIM comparison. Its README describes SSIM as experimental; check the installed release’s documentation before relying on it. [c003] Pixelmatch requires the images to have equal dimensions and exposes a per-pixel threshold. [c004]
2. Install a repeatable Jest and Puppeteer setup
The following is a small JavaScript example using Jest, Puppeteer, and jest-image-snapshot. It uses the current CommonJS-style matcher setup. Match package versions to your project and check the matcher’s peer dependency before installing; its documentation lists supported Jest versions. [c003]
npm install --save-dev jest puppeteer jest-image-snapshot
Create jest.config.cjs:
module.exports = {
testEnvironment: 'node',
testMatch: ['<rootDir>/test/**/*.test.cjs'],
};
Create test/visual.test.cjs:
const puppeteer = require('puppeteer');
const { toMatchImageSnapshot } = require('jest-image-snapshot');
expect.extend({ toMatchImageSnapshot });
describe('home page appearance', () => {
let browser;
beforeAll(async () => {
browser = await puppeteer.launch({ headless: true });
});
afterAll(async () => {
await browser?.close();
});
it('matches the reviewed desktop baseline', async () => {
const page = await browser.newPage({
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 1,
});
try {
await page.goto(process.env.APP_URL || 'http://127.0.0.1:3000', {
waitUntil: 'networkidle0',
timeout: 30000,
});
await page.evaluate(() => document.fonts.ready);
await page.addStyleTag({
content: `
*, *::before, *::after {
animation: none !important;
transition: none !important;
caret-color: transparent !important;
}
`,
});
const image = await page.screenshot({
type: 'png',
fullPage: true,
});
expect(image).toMatchImageSnapshot({
failureThreshold: 0,
failureThresholdType: 'pixel',
storeReceivedOnFailure: true,
});
} finally {
await page.close();
}
});
});
Start the app separately, then run the test:
npm run start -- --port 3000
# In another terminal:
npx jest test/visual.test.cjs
The app command depends on your project’s scripts. The example waits for network idle, then for fonts, and suppresses CSS animation and transitions. Network-idle waiting can hang or time out on pages with persistent polling or long-lived requests; in that case wait for a meaningful page selector or application-ready signal instead.
First run and reviewed baseline
On its first run, the matcher creates a PNG baseline. Review that image as the intended appearance before accepting it into version control. Later runs compare the received screenshot with that baseline. On mismatch, inspect the baseline, received image, and diff before deciding whether the change is a regression or an intentional design update. [c003] Commit approved image baselines with the code so reviewers can see what changed. Jest’s snapshot guidance also recommends reviewing snapshot changes. [c007]
3. Make Puppeteer captures deterministic
A visual comparison is only meaningful if the test renders the same state under the same conditions. Differences in operating system, browser version, settings, hardware, power source, and headless mode can change screenshots. Playwright’s visual-testing guide documents these sources of variation; the same caution applies to Puppeteer. [c005]
- Pin the environment: use the same browser build, OS or container image, headless setting, viewport, and device scale factor for baselines and CI.
- Control page data: use fixed fixtures, stable accounts, deterministic dates, and predictable feature flags.
- Wait for a real ready condition: prefer an app-specific selector or readiness flag when network idle is unreliable. Wait for fonts and images that matter to the design.
- Disable motion and blinking: inject CSS to stop animations, transitions, caret blinking, and other known dynamic effects.
- Stabilize volatile content: freeze timestamps, rotate content to fixed test data, or mask a genuinely irrelevant dynamic region. Do not mask areas where regressions need to be caught.
- Set capture scope explicitly: use
fullPage: truefor the document, the viewport default for above-the-fold behavior, orclipto capture a defined region. Puppeteer documents PNG as the default screenshot type andfullPage: falseas the default. [c002]
Viewport, full-page, and clipped screenshots
Choose scope to match what the test is asserting. A full-page capture can catch lower-page changes but may be tall and more sensitive to content length. A viewport capture is smaller and useful for a specific responsive breakpoint. A clip is useful for a component or region, but establish its coordinates only after the page has reached its stable layout.
// Stable viewport screenshot
const viewportImage = await page.screenshot({ type: 'png' });
// Entire document
const fullPageImage = await page.screenshot({ type: 'png', fullPage: true });
// A fixed page region; CSS pixels
const clippedImage = await page.screenshot({
type: 'png',
clip: { x: 0, y: 0, width: 800, height: 600 },
});
Do not compare screenshots captured at different viewport sizes or device scale factors as if they were the same baseline. Keep one baseline per relevant viewport or device configuration.
4. Set thresholds without hiding real changes
There are two different thresholds to understand:
- Per-pixel sensitivity: Pixelmatch’s
thresholdcontrols how different a pixel must be before it counts as changed. Lower values are more sensitive. Pixelmatch documents a range from 0 to 1 and a default of 0.1. [c004] - Allowed total difference:
failureThresholdandfailureThresholdTypeinjest-image-snapshotdecide how many changed pixels or what percentage of the image can differ before the assertion fails. [c003]
For example, the matcher’s customDiffConfig.threshold is not the same as failureThreshold. The former changes per-pixel sensitivity; the latter allows a total mismatch allowance. Start strict, inspect the diffs produced in your pinned environment, and then set only the smallest justified allowance. There is no universal threshold that works across applications and renderers.
expect(image).toMatchImageSnapshot({
// Total allowed changed pixels, not per-pixel color sensitivity.
failureThreshold: 20,
failureThresholdType: 'pixel',
// Pixelmatch per-pixel sensitivity.
customDiffConfig: { threshold: 0.05 },
storeReceivedOnFailure: true,
});
Use a percentage when screenshot dimensions vary intentionally and percentage is the meaningful policy. Prefer consistent dimensions when possible: a size mismatch often signals a real test setup problem. Avoid increasing a threshold just to turn a noisy test green; first find and remove the source of noise.
5. Compare with Pixelmatch directly
Use Pixelmatch when you want to own the baseline and failure behavior. This complete script captures a Puppeteer page, decodes a checked-in PNG baseline, writes a visual diff, and exits with a failure code when the mismatch allowance is exceeded.
npm install --save-dev puppeteer pixelmatch pngjs
Save a known-good baseline as baseline.png, then create compare.cjs:
const fs = require('node:fs');
const path = require('node:path');
const puppeteer = require('puppeteer');
const pixelmatch = require('pixelmatch');
const { PNG } = require('pngjs');
async function main() {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage({
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 1,
});
await page.goto(process.env.APP_URL || 'http://127.0.0.1:3000', {
waitUntil: 'networkidle0',
timeout: 30000,
});
await page.evaluate(() => document.fonts.ready);
const actualBytes = await page.screenshot({ type: 'png', fullPage: true });
await page.close();
const expected = PNG.sync.read(fs.readFileSync('baseline.png'));
const actual = PNG.sync.read(actualBytes);
if (expected.width !== actual.width || expected.height !== actual.height) {
throw new Error(
`Image dimensions differ: expected ${expected.width}x${expected.height}, ` +
`received ${actual.width}x${actual.height}`
);
}
const diff = new PNG({ width: expected.width, height: expected.height });
const mismatchedPixels = pixelmatch(
expected.data,
actual.data,
diff.data,
expected.width,
expected.height,
{ threshold: 0.1 }
);
fs.writeFileSync('diff.png', PNG.sync.write(diff));
const allowedPixels = 0;
console.log(`${mismatchedPixels} changed pixels; allowance ${allowedPixels}`);
if (mismatchedPixels > allowedPixels) process.exitCode = 1;
} finally {
await browser.close();
}
}
main().catch((error) => {
console.error(error);
process.exitCode = 1;
});
Pixelmatch expects equal-size images, returns a mismatch count, accepts an optional diff buffer, and lets you tune its per-pixel threshold. [c004] The script intentionally fails on different dimensions rather than resizing, since an unexpected size change can itself be a regression.
6. Review diffs and update baselines safely
- Run the test and save the baseline, received image, and diff as CI artifacts when available.
- Open the three images at the same scale. Check whether the diff corresponds to a meaningful layout or content change.
- Fix the application or test setup if the difference is unintended or caused by unstable data, fonts, animations, or environment mismatch.
- If the design change is intended, update the baseline deliberately, inspect the replacement image, and commit it alongside the code change.
- Remove obsolete baselines only after verifying that the corresponding tests were actually removed. The matcher documents that obsolete image snapshot cleanup requires extra handling; avoid cleanup based on a partial test run. [c003]
With Jest, run npx jest -u to update snapshots, then inspect every changed PNG before committing. The update flag is a file operation, not an approval decision.
7. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Every run creates a different diff | Unstable data, animation, fonts, image loading, or browser/OS differences. | Pin the render environment, use fixtures, wait for fonts and meaningful app readiness, and suppress known motion. |
| “Image dimensions do not match” | Viewport, device scale factor, page content height, or capture scope differs. | Set these values explicitly and investigate content changes. Do not enable size mismatch as a way to conceal unexpected layout changes. |
| Navigation times out waiting for network idle | The page polls continuously, opens long-lived connections, or loads third-party requests. | Wait for a stable selector or an application-ready signal; set a bounded timeout and diagnose genuinely stalled assets. |
| Screenshot contains a fallback font or missing image | The capture happened before fonts or images finished loading, or the CI environment lacks the expected font. | Wait for document.fonts.ready, verify required fonts exist in the runner, and wait for important images to be complete. |
Matcher says toMatchImageSnapshot is undefined |
The matcher was not registered with Jest, the setup file did not run, or the test runner is incompatible. | Call expect.extend({ toMatchImageSnapshot }) in the loaded test or setup file and check the installed package’s Jest peer requirements. |
| Jest updates the baseline but the test still looks wrong | The update command accepted a new expected image without visual review. | Open the changed baseline and diff, confirm the product change is intended, then commit it with the implementation. |
| Pixelmatch reports color noise over a large image | Small rendering differences affect many pixels, often due to environment or scaling variance. | Align the render environment and image dimensions first; only then calibrate per-pixel sensitivity or a small total allowance. |
| Test passes after raising threshold but misses visible regressions | The total allowed difference or per-pixel sensitivity is too permissive. | Restore stricter settings and control nondeterminism or isolate a truly volatile region. |
8. Performance, reliability, and cost
Visual tests spend time launching a browser, loading the application, waiting for a stable render, capturing pixels, and decoding/comparing images. Full-page and high device-scale screenshots contain more pixels and generally take more resources to capture and compare than a smaller viewport or clipped region. Keep the tested surface no larger than the behavior you need to protect, reuse a browser process when the test runner safely supports it, and avoid capturing the same page repeatedly without a reason.
Reliability depends more on deterministic rendering and review discipline than on the particular matcher. Keep CI and baseline generation on the same pinned image, preserve failure artifacts, and treat the screenshot baseline as a reviewed test artifact. Pixel thresholds can reduce noise but cannot tell whether a visual change is acceptable.
The open-source path has no ScreenshotNeo API charge, but it uses your CI time and storage for browser runs and PNG artifacts. Hosted screenshot capture can be useful when you need captures outside a local browser test process; it is a different workflow from comparing app builds against checked-in visual baselines.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It can capture a URL as PNG, JPEG, WebP, or PDF; for visual regression, keep your own reviewed baseline and compare the returned image bytes in your test.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs').then(({ writeFileSync }) => writeFileSync('shot.webp', bytes));
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers say which page verdict and billing status applied. An MCP server lets AI agents use screenshot tools. One thousand screenshots a month are free with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month, with no card required.
FAQ
What is the difference between snapshot testing and visual regression testing?
Ordinary Jest snapshots compare serialized values such as text or object representations. Visual regression testing compares rendered page images. They share the word “snapshot,” but need different comparison logic. [c007]
Should a visual test capture the whole page?
Only when the full document is in scope. Use a viewport or clip for focused assertions; give each viewport or device configuration its own baseline.
Should I use SSIM or pixel comparison?
Use pixel comparison when you need a direct pixel diff and precise mismatch counts. Consider SSIM when structural similarity better fits your review needs, while accounting for the matcher’s documented experimental status and verifying behavior with your own images. [c003]
When should I accept a new baseline?
After reviewing the actual image and diff and confirming the visual change is intended. A passing test after baseline replacement is not evidence by itself that the change is correct.


