How to debug a Happo visual test that fails only in CI
Find out whether a CI-only Happo failure is a real visual change, a target mismatch, or a screenshot capture problem—and debug it systematically.
A Happo visual test that fails only in continuous integration (CI) does not automatically mean the application is broken. First determine whether Happo completed the screenshot comparison and found changed pixels, or whether the snap request failed or captured incomplete content. A completed diff needs visual review; an errored or incomplete capture needs investigation of the run, target, resources, and page readiness.
This guide gives you a repeatable way to compare a failing CI run with a passing local run, inspect report logs, isolate target-specific differences, and decide whether a change is a real regression or harmless rendering noise. Without the report, CI log, config, and failing target, there is no way to identify a particular project’s root cause.
1. Classify the failure before changing code
Open the Happo report and locate the failing component and target. Separate these cases:
| What you see | What it means | First action |
|---|---|---|
| A completed comparison with changed pixels | The capture completed, but the rendered image differs from its baseline. | Inspect the diff and the target context. Decide whether the change is intended. |
| A snap request error or timeout | The comparison may not have produced a usable screenshot. | Read the request log and the surrounding CI output. |
| A screenshot with missing or incomplete content | The page rendered, but resources or asynchronous work may not have been ready. | Check image, font, network, and readiness evidence in the report logs. |
Happo’s [visual testing overview](https://admin.happo.io/) describes side-by-side comparisons and visual diffs. Its [report logs announcement](https://happo.io/blog/report-logs) explains that logs can surface snap-request errors and quieter problems such as missing images, delayed fonts, or asynchronous work that finishes after capture.
Do not immediately accept a diff, increase tolerance, or change shared CSS. First establish which of the three cases you have.
2. Reproduce the CI run as closely as possible
- Record the failing evidence. Save the report URL, commit SHA, failing component, target name, viewport, and whether the report shows a completed image comparison or an errored request.
- Compare the commands. Check the CI working directory, package manager, install command, build and start steps, selected Happo config, and exact Happo command against your local run.
- Check the configuration and secrets. Confirm CI loads the intended config, project, and target set. Happo’s public package repository shows credentials supplied as
HAPPO_API_KEYandHAPPO_API_SECRET. Check that the variables exist without printing their values. - Rerun the same commit and target. Keep the command, config, browser, and viewport fixed. If possible, compare a local run from the same commit, rather than comparing unrelated branches or dependency states.
- Change one suspected input at a time. A single controlled change makes it easier to tell whether the evidence supports a cause.
Happo’s [public repository](https://github.com/happo/happo) documents the CLI, automatic configuration-file discovery, and target configuration. The repository lists discovery for happo.config.js, .mjs, .cjs, .ts, .mts, and .cts. Make sure the file you intend to use is in the working directory from which the CI command runs.
Example CI command and config checks
The following is a diagnostic shell snippet for a Node-based project. It checks that the credential variables are present without exposing them, prints the working directory, and invokes Happo. Adapt the install and build commands to the project; do not print secret values.
set -eu
pwd
node --version
npm --version
: "${HAPPO_API_KEY:?HAPPO_API_KEY is missing}"
: "${HAPPO_API_SECRET:?HAPPO_API_SECRET is missing}"
npx happo
A minimal TypeScript config, based on Happo’s public repository example, can make the intended targets and viewport explicit:
import { defineConfig } from 'happo';
export default defineConfig({
apiKey: process.env.HAPPO_API_KEY!,
apiSecret: process.env.HAPPO_API_SECRET!,
targets: {
'chrome-desktop': {
type: 'chrome',
viewport: '1280x720',
},
'firefox-desktop': {
type: 'firefox',
viewport: '1280x720',
},
},
});
Use your existing project config as the source of truth; this example is not a drop-in replacement for every integration. Check the repository for current configuration details before changing target names or options.
3. Compare browser targets and viewports
Determine whether the failure is limited to one target or appears across all targets. Happo’s repository documents desktop targets including Chrome, Firefox, Edge, Safari, and accessibility, plus mobile Safari targets and viewport configuration. A failure isolated to one target is a reason to compare that target’s browser and viewport settings with a passing target before changing shared application CSS.
- One browser fails: inspect the diff at that target and compare its viewport and rendering context with a passing browser.
- One viewport fails: check responsive breakpoints, wrapping, overflow, and content that changes at that width.
- Every target fails similarly: investigate shared inputs such as a changed component, missing asset, or wrong build/config before treating it as browser-specific.
- Only CI fails for every target: focus first on differences in command, dependency install, credentials, environment, and content readiness.
Keep the target fixed while testing a hypothesis. Changing browser, viewport, and CSS at the same time makes the result hard to interpret.
4. Read the report logs around the failing component
Use the report’s logs to search for the component name and inspect nearby lines from the relevant snap request. Happo announced a unified searchable report-log page that brings snap-request logs into one timeline, shows the originating target, and provides context around a match. See [Happo’s report-log guide](https://happo.io/blog/report-logs) for where to find it.
Look for evidence such as:
- Snap-request errors or timeouts near the component.
- Image requests that failed or returned a 404.
- Fonts that loaded late or did not become available before capture.
- Asynchronous work that completed after the screenshot was taken.
- A warning present only for the failing browser target.
A log message is a clue, not a diagnosis by itself. Relate it to the component and image: a missing image request matters when the image is visibly absent or changed in the captured region. Preserve the exact log line, target, and report link when escalating the issue.
5. Check content readiness, assets, and external requests
A local page that looks correct does not prove that CI rendered the same data or received the same assets. Compare what was available at capture time: the component data, images, fonts, and asynchronous work that affects the final frame.
- Open the affected component in the CI report and note what is visibly absent, stale, or shifted.
- Search the associated report log for the matching image, font, or asynchronous operation.
- Compare those observations with a local run of the same commit and target.
- If the evidence points to an external request, determine whether the request is expected and whether the rendered result should depend on it.
- Make one change that addresses the observed input, then rerun the same target and workflow.
Happo says its Playwright workflow silences animations and waits for asynchronous assets and fonts as part of flake reduction, but this does not remove the need to inspect the specific report logs. If the suite depends on external resources, investigate their availability and behavior in the reported run rather than assuming they matched local execution.
6. Separate real visual changes from capture noise
Inspect the changed area at full size. A consistent shift in layout, color, or typography may be a real regression even when functional or end-to-end tests pass. Happo presents visual testing as a complement to functional assertions; its [homepage](https://admin.happo.io/) includes an example where visual comparison caught a text color and positioning change while Playwright end-to-end tests passed.
Also check whether the component contains values or behavior that can change between runs:
- Current time, relative dates, or time-zone-dependent formatting.
- Randomized or unstable item ordering.
- Rotating content, transitions, or animations captured at different frames.
- Data that changes between the local and CI environment.
- External images or fonts that can vary or arrive at different times.
Happo documents animation-silencing options and color-delta tolerance for small rendering differences such as antialiasing or image-compression noise. Treat those as noise controls after reviewing the actual diff, not as root-cause fixes. A looser threshold can hide a small but meaningful change.
7. Troubleshooting common CI-only failures
| Symptom | Likely area to investigate | Next step |
|---|---|---|
| CI cannot authenticate or submit a run | Missing, misnamed, or unavailable credentials; wrong config or working directory. | Check that HAPPO_API_KEY and HAPPO_API_SECRET are present in the job environment and the intended config is loaded. Never log their values. |
| Happo cannot find the expected configuration | Config file absent from the job checkout or command launched from another directory. | Print pwd, check the file exists in the job workspace, and compare the invocation directory with local. |
| Snap request errors or times out | Execution failure or request did not complete; the report log is needed to narrow it down. | Read the complete snap-request log and surrounding CI output before changing visual thresholds. |
| Image is missing or different only in CI | Asset request failed, returned different content, or was not ready at capture. | Find the request in the report log and compare the captured image with the same commit locally. |
| Text wraps differently or falls back to another font | Different viewport or font availability/readiness. | Compare the target viewport and inspect font-related log lines; confirm the intended font appears in the completed capture. |
| Only one browser or viewport differs | Target-specific rendering or responsive behavior. | Compare the failing target’s browser and viewport with a passing target; reproduce without changing shared CSS first. |
| Diff moves between runs | Dynamic values, randomized order, rotating content, or motion. | Identify the changing input and stabilize the rendered state; use documented animation controls where applicable. |
| Many differences appear after changing tolerance | Threshold may be too strict or too permissive for the image, but the diff still needs judgment. | Review the changed pixels and tune tolerance only for known immaterial variation; verify it does not mask meaningful changes. |
8. Performance, reliability, and cost considerations
Debugging can become slower when you rerun a broad suite after every edit. First narrow the evidence to the failing component and target, then keep the same run conditions while checking a single hypothesis. Once the cause is understood, run the wider workflow required by your team to ensure the correction did not affect other targets.
For reliability, preserve the report URL and CI artifacts for intermittent failures, keep credentials out of logs, and compare like-for-like runs. Avoid using retries as proof that the failure is harmless: an intermittent diff can still point to unstable data or resource readiness.
The provided research does not establish Happo pricing or a general cost per rerun, so this guide does not estimate either. Review your Happo plan and CI runner usage for project-specific cost implications. The practical cost of broad reruns is additional CI time; narrowing the investigation can reduce unnecessary repeated work without changing the evidence standard.
9. When a screenshot API helps
Happo reports answer questions about your configured visual test suite. A separate website screenshot can help inspect a publicly reachable page or compare a capture outside that workflow, but it does not diagnose a Happo report by itself. For those one-off captures, ScreenshotNeo is a website screenshot API and MCP server for developers. Its clean-shot flow accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. It bills only clean shots, while bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the outcome identified in response headers.
Or skip the browser setup
For a one-call capture of a page you can access, use ScreenshotNeo’s API. See the ScreenshotNeo API documentation for request options and setup.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, timeouts, and failed loads are never billed; response headers identify the page verdict and billing outcome.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots. Yearly billing gives two months free, and every feature is on every plan.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Does a CI-only visual diff always mean a bug?
No. It means the captured image differs from its baseline. Inspect the changed region and target to decide whether it reflects an intended change, a regression, or a variable input.
Should I increase the color-delta tolerance to make CI pass?
Only after reviewing the actual diff and deciding the variation is immaterial. Tolerance can also suppress a small real regression.
Can a separate screenshot API tell me why Happo failed?
No. It can capture a webpage, but Happo’s report, snap-request logs, CI output, and target configuration are the evidence needed to diagnose a Happo run.
What information should I include when asking for help?
Share the report link, failing component and target, relevant log lines, CI command and config context, and whether the capture completed. Redact credentials and other secrets.


