ScreenshotNeo

BlogHow-to

How to Test Web Pages with Screenshot Diffs in BackstopJS

Set up BackstopJS visual regression tests, stabilize browser captures, read screenshot diffs, and safely approve updated reference images.

By the ScreenshotNeo team4 October 20269 min read

BackstopJS tests web pages with screenshot diffs by capturing a reference image, capturing the same scenario again after a change, and showing the visual differences for review. Initialize a project with backstop init, define viewports and scenarios in backstop.json, run backstop test, inspect the report, then run backstop approve only when the new appearance is intentional.

A diff is evidence to inspect, not automatic proof of a bug. Reliable results depend on repeatable page state, viewport, browser environment, and data. The examples below follow the BackstopJS project documentation; check the current README and your installed version for any changes.

1. Install and initialize BackstopJS

Use a Node.js project so the BackstopJS version is recorded with the project. From the repository root:

npm init -y
npm install --save-dev backstopjs
npx backstop init

The initialization command creates a starter configuration and supporting scripts. Commit the configuration and any scenario scripts your team maintains. Keep the same BackstopJS and browser dependencies in local development and CI to reduce rendering variation.

2. Define viewports and scenarios

A scenario represents one URL and page state that you want to compare. Give every scenario a descriptive label and URL. Add at least one viewport. This compact example captures a desktop and mobile product page, waits for an application readiness selector, masks a variable timestamp, and checks a selected region:

{
  "id": "storefront",
  "viewports": [
    { "label": "desktop", "width": 1440, "height": 900 },
    { "label": "mobile", "width": 390, "height": 844 }
  ],
  "scenarios": [
    {
      "label": "Product detail",
      "url": "http://localhost:3000/products/widget",
      "selectors": ["#product-detail"],
      "readySelector": "#product-detail[data-ready='true']",
      "delay": 250,
      "hideSelectors": [".last-updated-time"],
      "misMatchThreshold": 0.1
    }
  ],
  "paths": {
    "bitmaps_reference": "backstop_data/bitmaps_reference",
    "bitmaps_test": "backstop_data/bitmaps_test",
    "html_report": "backstop_data/html_report",
    "ci_report": "backstop_data/ci_report"
  },
  "engine": "puppeteer"
}

The threshold shown is BackstopJS’s documented default, not a general recommendation. Tune it only after you have stable captures and understand what the threshold means for your pages.

Choose what the screenshot represents

  • selectors captures only the specified DOM element or elements; omit it to capture the document. Prefer a stable selector that identifies the region whose appearance matters.
  • readySelector waits for an element that indicates the page is ready. Use a real application state signal rather than guessing how long rendering takes.
  • readyEvent can wait for an application event when your page can signal readiness explicitly.
  • delay adds a wait after readiness. Use it for known delayed rendering, not as the only synchronization mechanism.
  • onBeforeScript and onReadyScript let you set browser state or perform interactions. Use the scripts that match your configured engine.
  • hideSelectors hides selected elements while retaining their layout space. removeSelectors removes elements from the DOM, which can change layout flow.
  • misMatchThreshold sets the percentage of different pixels allowed. Lower values are more sensitive. Select a value based on the repeatability and visual importance of the page.

For pages with rotating ads, live counters, or changing timestamps, first decide whether the content’s presence and layout are part of the behavior under test. Hide a changing fixed-size region when its space should remain; remove it only when removing its space is the intended test state.

3. Make page state repeatable

Visual tests are meaningful when the same scenario produces the same intended page state. A practical setup checklist:

  1. Control the data. Use fixtures, seeded records, or static content stubs instead of live data that changes between runs. Cover content of different lengths so the layout is tested under realistic conditions.
  2. Control authentication. Establish the same signed-in or signed-out state every run. BackstopJS documents Playwright storage state for cookies and local storage when using Playwright.
  3. Wait for the application. Prefer a readiness selector or event tied to actual application state. Add a short delay only for a known remaining animation or deferred render.
  4. Control interactions. Use scenario scripts to navigate menus, dismiss onboarding, or select a stable tab before capture. Make the actions deterministic.
  5. Control motion and time-dependent content. Disable or settle animations where appropriate, and freeze or mask clocks, random values, and continuously changing regions in the test environment.
  6. Keep external services out of the critical path. Stub requests whose responses can vary or fail independently of the UI change under test.

BackstopJS specifically recommends known static content stubs for dynamic applications, including content of varying lengths. A stable empty-state-only fixture can miss wrapping and overflow regressions that appear with longer content.

4. Run the comparison and inspect the report

Run the first test to create captures and a report:

npx backstop test

Review the reference, test, and difference views for each failed scenario. Look for shifted containers, missing assets, font changes, wrapping changes, unexpected overlays, and legitimate design updates. Use --filter to rerun selected scenarios while iterating:

npx backstop test --filter="Product detail"

Confirm the filter matches the scenario labels used in your configuration. A filtered run is useful for focused debugging; run the full suite before accepting a change so unrelated scenarios are still checked.

5. Approve only intentional visual changes

When a change is expected, review the affected captures and then promote the latest test captures to references:

npx backstop approve

Approval updates the baseline used by later comparisons. Treat it as a review decision: confirm the changed appearance matches the design or product requirement, and make sure the report covers the scenarios you intend to promote. The BackstopJS documentation also describes filtering which captures are promoted. Do not approve merely to turn a failing run green.

6. Select the rendering engine and environment

BackstopJS documents Puppeteer as the default configuration and says Puppeteer and Playwright are installed by default. Playwright provides a route to Chromium, Firefox, or WebKit; select an engine and browser that match the compatibility questions your team needs to answer. For Playwright, use the corresponding Playwright onBefore and onReady scripts. Use Playwright storage state when the capture needs a prepared authenticated context.

The same page can render differently in different environments, particularly text. Keep reference and test runs on the same operating system, browser and dependency versions, fonts, and rendering configuration. BackstopJS offers a --docker mode to render in a container, which can reduce cross-environment differences when the container matches your workflow; it does not remove all sources of nondeterminism.

7. Tune sensitivity without hiding meaningful changes

The documented default misMatchThreshold is 0.1, a percentage of differing pixels. It is a configurable tool default, not a universal setting. A noisy page may need its state stabilized before threshold changes. A threshold that is too permissive can hide small but important changes; an overly strict one can make runs noisy when rendering has harmless variation.

The README notes that default mismatch reporting does not detect mismatches below 0.01%. It documents usePreciseMatching for cases that need a threshold below that level. Enable more precise matching only when those tiny differences matter and the test environment is sufficiently stable.

8. Use BackstopJS in continuous integration

Run the same install, application setup, and BackstopJS command in CI that developers use locally. A typical job should:

  1. Install the locked project dependencies.
  2. Start the application with deterministic test data and wait for it to be reachable.
  3. Run npx backstop test using the chosen engine and consistent environment.
  4. Preserve the generated HTML report and captures as CI artifacts when the run fails.
  5. Require human review of changed references before running approval and committing updated baselines.

Store approved reference images in version control or another controlled artifact location that lets reviewers see the baseline change alongside the code change. Keep secrets out of scenario files and logs. Avoid running approval automatically on every branch: doing so can silently bless regressions.

9. Troubleshoot common failures

Symptom Likely cause What to do
Many diffs appear without a UI change Different browser, operating system, fonts, device scale, data, or timing between captures. Run both captures in the same environment and use deterministic fixtures, consistent dependencies, and app-specific readiness signals.
Screenshot is blank or incomplete The page was captured before client rendering or required data finished loading. Add a readySelector or readyEvent that represents application readiness; check the page and network behavior separately.
Only live widgets keep failing Ads, chat, counters, or timestamps vary between runs. Stub the source where possible; otherwise use hideSelectors to preserve layout or removeSelectors when removal of layout is appropriate.
Text wraps differently in CI Font availability or rendering environment differs. BackstopJS notes that text can vary between environments. Use consistent fonts and browser environment; consider --docker when a container fits your workflow.
Navigation to localhost fails in Docker The application address inside the container may not refer to the host machine. For Mac and Windows Docker setups, the README points to a host-accessible address such as host.docker.internal. Check the exact networking behavior of your host and container configuration.
Browser times out or the container runs out of memory Headless Chrome memory use can contribute to timeout trouble in Docker, according to the README. Check container memory allocation and concurrency, reduce the workload, and inspect whether the page is waiting on an external request or readiness condition.
Test fails after changing scenario names or filters The filter may not match a configured scenario label. Use the exact scenario label with --filter, then run the complete suite after debugging.
A small difference is not reported The difference may fall below default precision or configured threshold. Review the documented precision limit and consider usePreciseMatching only if sub-0.01% differences matter.
Approval changes more baselines than expected Approval operates on captures from the most recent test batch. Inspect the report and use the documented filtering options to limit promoted captures before approving.

10. Performance, reliability, and maintenance

Screenshot suites multiply work by scenarios and viewports, and browser startup and page readiness add time to each capture. Start with representative high-value pages and states, then expand coverage where visual regressions would matter. Use --filter for local iteration; keep a full run in CI. Reduce avoidable waits and external dependencies, but do not remove readiness checks that protect against capturing an unfinished page.

For reliability, keep test data and browser environments repeatable, retain failure reports, and make baseline updates reviewable. More parallel browser work can shorten elapsed time but may increase memory pressure; tune concurrency against the resources available to the runner. The project README mentions headless Chrome memory use as a possible Docker timeout cause.

The BackstopJS README contains a notice that the project needs a new maintainer or owner. That notice reflects the README when accessed for this article and does not establish current release cadence or whether a maintainer has since been appointed. Before adopting it for a long-lived test system, check the current project repository for activity and verify compatibility with your dependency and browser requirements.

Or skip the browser setup

If you need a screenshot for a check, report, or agent workflow without managing a browser runner, ScreenshotNeo provides a website screenshot API and MCP server. The API can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for its options and setup.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 screenshots; all features are available on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does BackstopJS decide whether a design change is acceptable?

No. It identifies visual differences against a reference; a developer or reviewer decides whether those differences are intended.

Can I compare only part of a page?

Yes. Configure selectors to capture specific DOM elements instead of the whole document.

Should I approve references after every passing run?

No. Approval is for reviewed, intentional changes that should define the future expected appearance.

Does using Docker guarantee identical screenshots?

No. It can reduce differences between rendering environments, but page data, timing, fonts, browser versions, and other factors can still affect captures.

Is BackstopJS suitable for checking functional behavior?

Screenshot comparison checks appearance. Use functional tests for behavior such as whether a form submits or navigation works; visual captures can complement those checks.