ScreenshotNeo

BlogHow-to

Puppeteer Visual Regression Testing with BackstopJS

Set up BackstopJS with Puppeteer, create stable screenshot references, review visual diffs, and run the checks in CI.

By the ScreenshotNeo team4 October 20268 min read

BackstopJS uses Puppeteer by default to capture web pages and compare them with approved reference screenshots. Define viewports and scenarios, prepare each page in a repeatable state, create references with backstop reference, then run backstop test and review the visual report. A mismatch is a signal to investigate; approve it only when the change is intentional.

This guide builds a small, runnable setup, explains the choices that affect repeatability, and shows how to run the comparison in CI. BackstopJS describes itself as a tool that automates visual regression testing by comparing screenshots over time. See the BackstopJS project documentation for version-specific configuration details.

1. Install BackstopJS and initialize a project

Run these commands from the frontend project root. BackstopJS documents a default backstop.json configuration and supports JavaScript configuration files as well.

npm install --save-dev backstopjs
npx backstop init

Initialization creates a starter configuration and the supporting folders. The example below uses JavaScript so that setup can be version controlled alongside the application. Save it as backstop.config.js in the project root and invoke BackstopJS with --config=backstop.config.js.

2. Define scenarios, viewports, and capture behavior

A scenario needs a label and URL; the configuration needs at least one viewport. This example captures the page viewport at desktop and mobile widths and waits for a product heading before capture. Replace the sample URL and selector with an environment and readiness signal your team controls.

module.exports = {
  id: 'storefront-visual-tests',
  viewports: [
    { label: 'desktop', width: 1440, height: 900 },
    { label: 'mobile', width: 390, height: 844 }
  ],
  scenarios: [
    {
      label: 'Product page',
      url: 'http://localhost:3000/products/example',
      readySelector: '[data-testid="product-title"]',
      delay: 250,
      selectors: ['document'],
      selectorExpansion: false,
      requireSameDimensions: true,
      misMatchThreshold: 0.1
    }
  ],
  paths: {
    bitmaps_reference: 'backstop_data/bitmaps_reference',
    bitmaps_test: 'backstop_data/bitmaps_test',
    engine_scripts: 'backstop_data/engine_scripts',
    html_report: 'backstop_data/html_report',
    ci_report: 'backstop_data/ci_report'
  },
  report: ['browser', 'CI'],
  engine: 'puppeteer',
  engineOptions: {
    args: ['--no-sandbox']
  },
  asyncCaptureLimit: 2,
  asyncCompareLimit: 20,
  debug: false
};

The --no-sandbox launch argument is often needed in containerized CI environments, but it changes Chromium’s security posture. Use an appropriately isolated runner and check the installed BackstopJS/Puppeteer documentation before adding or removing browser flags. The project documents Puppeteer as its default engine and supports configurable engine options.

Choose what to capture

  • document covers the full page and is useful for page-wide layout changes.
  • viewport limits the capture to the visible viewport, which can keep a scenario focused on the initial screen.
  • A CSS selector focuses on a component. BackstopJS captures the first matching element by default. Set selectorExpansion: true when repeated matches should each be captured.

Prefer the narrowest scope that still protects the behavior you care about. A focused component capture can reduce unrelated diffs; full-page coverage can reveal interactions between sections. Capture scope affects coverage and review effort, so document the reason for each scenario.

Set comparison tolerance deliberately

BackstopJS documents misMatchThreshold with a default of 0.1 percent and requireSameDimensions defaulting to true. The example makes both explicit. A threshold can absorb small rendering noise, but a higher value can also hide real changes. Dimension checking catches layout shifts that change the image size. Keep defaults or tune them only after reviewing representative diffs.

3. Make the page state repeatable

Visual comparisons are meaningful only when the reference and test runs represent the same state. BackstopJS supports setup scripts, interaction scripts, and readiness controls.

  • Wait for application readiness. Use readySelector when a stable element marks completion, or readyEvent when the application can emit a signal. Prefer these state-based conditions to an arbitrary delay.
  • Set up cookies and authentication. Use a before script for cookies or other browser preparation. Keep credentials out of committed configuration; supply secrets through your CI secret store.
  • Perform interactions. Ready scripts can click, hover, or otherwise prepare the page. Keep the same sequence for reference creation and test runs.
  • Control changing data. Use known static fixtures or stubs when possible. If a volatile region cannot be stabilized, mask or remove it only when that region is deliberately outside the behavior under test. Masking pixels can conceal a regression.
  • Account for animation. Wait for the application to reach its settled state, then add a small delay only if a known transition needs time to finish.

For custom preparation, BackstopJS scripts receive the browser page and scenario context. The project documentation describes using them for cookies, user agents, and viewport-specific behavior. Check the installed version’s script interface before copying an example from another version.

4. Create references, compare, and approve changes

  1. Start the application at the configured URL and ensure its data and assets are available.
  2. Create approved baseline images: npx backstop reference --config=backstop.config.js.
  3. Run a comparison: npx backstop test --config=backstop.config.js.
  4. Open the generated browser report. Inspect each mismatch and its difference view; determine whether it is an actual defect, rendering noise, or an intended design change.
  5. When a change is intentional, update the baseline with npx backstop approve --config=backstop.config.js. Use filtering carefully if approving only a subset of scenarios.

Approval promotes the most recent test captures into the reference collection. Do not approve a failing run just to make the pipeline green: the reference set is the expected state against which later changes are judged.

5. Run BackstopJS in CI

BackstopJS can produce browser and CI reports; its CI report uses JUnit format by default. The documented CLI exits with 0 when tests pass and 1 when a test fails, so a normal command step can gate the build.

# Example CI steps; adapt these to your CI provider.
npm ci
npm run start:test &
npx wait-on http://localhost:3000
npx backstop test --config=backstop.config.js

If your project does not already provide wait-on, install it as a development dependency. Ensure the server command remains alive for the duration of capture, and configure the CI provider to retain the BackstopJS report and test images when a job fails. Keep reference images under source control or another deliberate versioned artifact workflow so reviewers can see baseline changes.

Use the same rendering environment for reference generation and CI comparisons. Fonts, browser versions, operating systems, and data can all affect pixels. BackstopJS documents Docker rendering as an option for improving consistency; weigh the extra setup and runtime against the reproducibility your team needs.

6. Puppeteer and browser-engine choices

Puppeteer is BackstopJS’s default engine and is sufficient for many Chrome or Chromium screenshot comparisons. Engine flags and navigation parameters are configurable through engineOptions; consult the versions installed in your project before copying flags because defaults can change.

If your visual tests must cover Firefox or WebKit, BackstopJS documents Playwright as an alternative rendering engine. Choose an engine based on the browser coverage you need. Adding another engine solely for basic screenshot comparison adds setup and maintenance without expanding the tested requirement.

7. Troubleshooting common failures

Symptom Likely cause Fix
Navigation fails or the page is blank The app server is not running, the URL is wrong, or the service is not ready when capture starts. Check the URL from the same runner, start the app before BackstopJS, and wait for a reachable health endpoint or page.
Intermittent diffs on the same commit Dynamic data, animations, late-loading content, or inconsistent fonts and browser environments. Stub data, wait for a readiness selector/event, settle animations, and align the reference and test environments.
A component capture is missing or captures the wrong content The selector does not match, matches a hidden instance first, or is asynchronous. Use a stable unique selector, wait for it to appear, and enable selector expansion only when repeated matches are intended.
CI reports a browser launch error Chromium dependencies or launch permissions differ in the runner. Use a supported environment, install required browser dependencies, and review Puppeteer launch flags for that runner. Avoid blindly copying flags across versions.
Every image fails the dimension check A viewport, content height, or responsive breakpoint differs between runs. Confirm the configured viewport and data are identical. Keep requireSameDimensions enabled when size changes matter; change it only when variable dimensions are expected.
CI exits with code 1 At least one visual comparison failed, or capture/report generation encountered a failure. Inspect the report and logs to distinguish a genuine mismatch from a setup error. Fix the page or deliberately approve an intended visual change.
Text differs across developer machines Font availability, operating system rendering, or browser version differs. Standardize the rendering environment, including fonts and browser, or use the documented Docker rendering option.

8. Performance, reliability, and maintenance

Capture cost grows with the number of scenarios and viewports, the amount of page content, and how long each scenario waits for readiness. Start with the pages and states that carry the most risk, then expand coverage where diffs have useful diagnostic value. Tune asynchronous capture limits only after observing resource pressure in your own runner; no universal concurrency value fits every application.

Reliability comes from deterministic inputs and deliberate baseline review. Pin dependencies through the project lockfile, preserve reports when a run fails, and keep reference and test images tied to the same browser and operating environment. A visual mismatch is evidence for review, not proof of a defect.

The BackstopJS repository README currently says, “BackstopJS needs a new maintainer/owner.” The available documentation does not establish a release cadence, support lifetime, or vulnerability response process. For long-lived infrastructure, review recent project activity and verify compatibility with your required browser and Node.js versions before adopting or upgrading.

Or skip the browser setup

If your goal is to capture a page screenshot rather than maintain a local comparison harness, ScreenshotNeo is a website screenshot API and MCP server. Its API accepts one GET request with a URL and returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Sign up for 1,000 free screenshots a month, with no card required.

Frequently asked questions

Does a visual mismatch automatically mean the page is broken?

No. It means the captured image differs from its approved reference. Review the report and decide whether the difference is a defect or an intended change.

Can BackstopJS capture only one component?

Yes. Configure a CSS selector in the scenario. By default the first matching element is captured; use selector expansion when each repeated match should be covered.

Should I use a delay or a readiness selector?

Use an application readiness selector or event when possible. A delay is best reserved for a known transition or final settling period.

Do I need Playwright to use BackstopJS?

No. Puppeteer is the default engine. BackstopJS documents Playwright as an alternative when browser coverage such as Firefox or WebKit is required.