ScreenshotNeo

BlogHow-to

Why BackstopJS reports false visual differences and how to fix them

BackstopJS diffs can reflect capture timing or environment changes, not just UI regressions. Stabilize the screenshot first, then tune comparison settings.

By the ScreenshotNeo team4 October 20268 min read

A BackstopJS failure means the test and reference images differ. It does not by itself prove that a user-facing regression occurred: timing, dynamic content, browser state, and rendering environment can all change the pixels. To fix false visual differences, first stabilize what the browser captures and when; next make the reference and test environments consistent; only then adjust comparison tolerances.

Use this order when asking, “Why is BackstopJS failing when nothing changed?” It avoids hiding real regressions by raising thresholds before you understand the diff. BackstopJS captures test screenshots and compares them with reference screenshots; its official workflow and troubleshooting documentation discusses rendering variation such as text appearing slightly differently between Linux and Mac.

1. Inspect the diff and establish what changed

  1. Open the report and identify the changed region. Determine whether the difference is page-wide, isolated to text, an animation, an ad, or an asynchronously loaded section.
  2. Compare the reference and test screenshots at the same scenario, viewport, and scroll position. Check image dimensions as well as pixels.
  3. Run the scenario again without changing configuration. If the changing pixels move between runs, suspect nondeterministic content, timing, or state. If they remain consistent, investigate the application change and capture setup.
  4. Check browser console output. BackstopJS documents scenarioLogsInReports for including browser console logs in reports.

A diff is a signal to investigate, not automatic proof of a user-visible defect. Keep a real change in scope when the scenario is meant to verify it; only suppress content that is deliberately outside the test’s purpose.

2. Wait for the page state the scenario needs

Single-page apps, Ajax requests, and progressive rendering can leave the browser capturing a partial page. Prefer a readiness condition that represents the content this scenario actually verifies:

{
  "readySelector": "#results-loaded",
  "readyTimeout": 30000
}

Choose a selector that appears only after the relevant view is ready. A selector that exists before important data or layout updates finish will still allow an early capture. BackstopJS also supports readyEvent for an explicit console signal from the application. A fixed delay can help when a known extra wait is appropriate, but it is less precise than waiting for the actual ready state.

The scenario properties documentation lists readyTimeout with a 30,000 ms default. If the page does not meet its readiness condition before the timeout, inspect the selector or event and application logs rather than continually extending the timeout.

3. Control dynamic and personalized content

For changing content, decide whether its layout space matters to the scenario. BackstopJS offers two different treatments:

Option What happens Use it when
hideSelectors The selected content is hidden for image comparison while its layout space remains. The content changes, but surrounding positions and reserved space are part of what you want to verify.
removeSelectors The selected element is removed from the DOM before capture. The content is outside the scenario and its size or presence is unpredictable.

Example scenario properties:

{
  "hideSelectors": ["#rotating-promotion"],
  "removeSelectors": ["#unpredictable-widget"]
}

Use the actual selectors from the page. If a fixture, cookie, or test account can make the content deterministic, that is often a better test when the content itself matters. Hiding or removing it can conceal a defect if it is part of the intended behavior.

4. Make reference and test capture environments match

Use the same browser, operating system or container, fonts, viewport, and relevant rendering configuration when creating references and running tests. A font fallback or platform text-rendering difference can change many pixels even if the page code is unchanged. The BackstopJS README explicitly gives Linux-versus-Mac text rendering as an example of visual variation and points to Docker-based sanity-test commands as an option for environment consistency.

If references were created on one machine and CI runs another environment, regenerate references in the intended test environment or make both environments consistent. Do not approve a new reference until you have established that its differences are expected.

5. Confirm interactions and application state

A scenario can differ because it captures a different state, not because rendering is unstable. Review scenario interactions, scripts, and any onReadyScript that runs after readiness conditions. Verify that clicks, hovers, and other actions target the intended element, and that resulting asynchronous updates finish before capture.

When state depends on authentication, cookies, locale, or test data, make those inputs repeatable between reference and test runs. If an interaction triggers a transition or animation, arrange for the scenario to capture the same settled state each time.

6. Tune comparison settings after capture is stable

misMatchThreshold sets the percentage of differing pixels tolerated before a screenshot fails. The project README describes threshold values as percentages from 0.00% to 100.00%. There is no universal correct value: it depends on the page, rendering stability, and the risk of missing a regression. Raise it only for understood residual noise, and inspect representative diffs before doing so.

requireSameDimensions determines whether a change in image dimensions is itself treated as a failure. Disabling that check may allow a test to pass despite different dimensions, but the size change can indicate a genuine layout regression. Keep dimension enforcement when page size is part of the behavior under test.

Symptom First setting or check Risk to consider
Incomplete content in the test image readySelector, readyEvent, or a suitable delay A selector that appears too early does not wait for the important content.
Changing widget pixels hideSelectors or removeSelectors Suppressing in-scope content can hide a real defect.
Text differs across machines Match browser, OS/container, and fonts Changing tolerance may conceal other differences too.
Small, understood residual rendering noise Narrowly adjust misMatchThreshold A higher threshold can let actual changes pass.
Dimensions differ Check viewport and page layout; review requireSameDimensions Ignoring dimensions can conceal a layout regression.

Runnable configuration pattern

BackstopJS configuration structure can vary by version and setup. The following is a scenario-properties pattern showing the relevant documented options; merge it into the scenario in your project’s existing configuration and verify property support against the installed version’s BackstopJS package documentation.

{
  "scenarios": [
    {
      "label": "Results page stable state",
      "url": "http://localhost:3000/results",
      "readySelector": "#results-loaded",
      "readyTimeout": 30000,
      "hideSelectors": ["#rotating-promotion"],
      "removeSelectors": ["#unpredictable-widget"],
      "scenarioLogsInReports": true
    }
  ]
}

Set a mismatch threshold and dimension policy at the configuration level used by your project only after reproducing a stable capture. The exact placement and supported options depend on the BackstopJS version and configuration format. The published BackstopJS 6.3.25 Playwright example is useful for version-specific configuration context.

Troubleshooting common failures

What you see Likely cause Fix
Test image shows a loading state or missing results Capture happened before asynchronous content was ready. Wait for a meaningful readySelector or readyEvent; use a delay only when a fixed wait is suitable. Check whether the readiness marker appears too early.
Diff changes from run to run Dynamic content, personalized state, animation, or inconsistent test data. Stabilize the fixture, account, cookies, and interactions. Hide or remove out-of-scope changing content using the option that matches whether its layout space matters.
Text changes across local and CI runs Different OS, browser, font availability, or rendering configuration. Run reference and test capture in the same environment, such as the same container and browser setup.
Whole page shifts after a widget is removed The element occupied layout space, and removal changed the flow. Use hideSelectors if the space must remain; use removeSelectors only when removing the element and its space is appropriate.
Scenario times out waiting to become ready The selector or event never occurs, is misspelled, or is not emitted in this state. Check the selector against the rendered page and verify the application emits the event. Increase the timeout only if the state is valid but predictably slower.
Threshold change makes failures disappear, but confidence falls The tolerated mismatch is broad enough to mask meaningful changes. Restore a stricter threshold, identify the noisy region, and stabilize or narrowly exclude only out-of-scope content.
Image dimensions differ unexpectedly Viewport, page content, or layout height changed. Compare capture dimensions and investigate the layout. Do not disable same-dimension checks just to silence an unexplained failure.

Performance, reliability, and maintenance

  • Wait on a meaningful condition. A readiness selector or event can avoid both premature captures and unnecessary fixed waiting. Choose a timeout that accommodates legitimate load time while still surfacing stalled pages.
  • Keep the environment repeatable. Consistent browser and container configuration reduces unrelated diffs and makes failures easier to reproduce.
  • Keep exclusions narrow. Selector rules are easiest to trust when they target specific unstable regions and have a clear reason. Review them when the page changes.
  • Preserve useful evidence. Reports with console logs can help distinguish application errors from pixel changes. Keep representative diffs when deciding whether residual variation warrants a threshold adjustment.
  • Account for maintenance cost. Readiness selectors and fixtures need updates when the application changes. Broad thresholds and broad exclusions reduce maintenance in the short term but also reduce the test’s ability to catch regressions.

Or skip the browser setup

If you need a clean page capture outside your BackstopJS regression workflow, ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL as PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
  • An MCP server lets AI agents, including Claude, Cursor, and other MCP clients, take screenshots.
  • The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

FAQ

Why is BackstopJS failing when nothing changed?

The compared pixels may have changed because capture timing, dynamic page state, fonts, browser, or operating system changed. Check the diff and capture conditions before treating it as a product regression.

How do I ignore dynamic content in BackstopJS?

Use hideSelectors to keep the element’s layout space while hiding its changing pixels, or removeSelectors when the element and its space should be absent from the capture.

How do I stop screenshot tests from changing between runs?

Make application state and test data repeatable, wait for a meaningful ready condition, and run reference and test captures in the same browser and operating-system environment.

Should I raise misMatchThreshold?

Only after you have stabilized captures and identified small residual rendering noise. A higher threshold can allow genuine changes to pass, so inspect diffs and adjust narrowly.

Does a pixel diff always mean users will see a problem?

No. A pixel difference can be caused by rendering or capture conditions. It still deserves investigation because it may also represent a real visual regression.