How to Compare Website Screenshots with ImageMagick
Compare website screenshots with ImageMagick: create a difference map, choose a metric, and avoid false changes caused by mismatched captures.
Use ImageMagick’s magick compare command to compare two screenshot files pixel by pixel. To create a visual difference map and print the root mean square error (RMSE), run:
magick compare -metric RMSE before.png after.png difference.png
before.png and after.png should show the same page at the same viewport, zoom, scroll position, and interaction state. The command writes metric details to standard error; difference.png is the visual output. In ImageMagick’s example, changed areas are red and untouched areas are white. A difference tells you pixels changed; it does not tell you whether a change is correct or desirable.
This guide uses the ImageMagick 7 command form and official examples. See ImageMagick’s compare documentation for command behavior and output details.
1. Capture comparable screenshots
Before comparing files, make the capture conditions reproducible. A small change in browser state can produce a large difference map even when your code change did not affect the page.
- Capture both page versions in the same browser and environment.
- Use the same viewport width and height, browser zoom, device scale, scroll position, and page state.
- Wait for the same readiness condition on both captures: for example, a specific selector or a known delay after navigation.
- Use lossless PNG where possible. JPEG or other lossy compression can introduce pixel differences unrelated to the page change.
- Keep the original captures alongside the difference image so a reviewer can inspect what changed.
For full-page screenshots, make sure the capture method produces the same dimensions and scroll behavior for both versions. If the page contains time-dependent content, animations, rotating banners, randomized data, or live counters, stabilize or disable those inputs before capturing.
2. Install ImageMagick and compare the files
Install ImageMagick using the package or installer appropriate for your operating system, then confirm the magick command is available:
magick -version
Run a visual comparison and request RMSE:
magick compare -metric RMSE before.png after.png difference.png
Open difference.png to locate changed regions. The metric is reported separately on standard error, so it may appear in the terminal even when the output image is created successfully.
If you want only a score and do not need a difference image, use ImageMagick’s null: output:
magick compare -metric RMSE before.png after.png null:
To save the score in a shell variable or log, redirect standard error explicitly. For example, in a POSIX-compatible shell:
magick compare -metric RMSE before.png after.png null: 2>rmse.txt
ImageMagick documents another same-size form using -compare:
magick before.png after.png -metric RMSE -compare difference.png
For a straightforward two-file comparison, the standalone magick compare form is easier to read and automate.
3. Choose a metric that matches your question
There is no single best metric for every visual regression check. Decide whether you need to count changed pixels, summarize average error, or estimate structural resemblance. ImageMagick’s command-line options reference describes the metric and fuzz options.
| Metric family | Useful when | How to interpret it |
|---|---|---|
| AE / PDC | You want to count pixels considered different. | -fuzz determines how much small color variation is ignored. Keep its value consistent across runs. |
| MAE / RMSE | You want one summary of average channel error. | RMSE gives larger errors more weight. A scalar does not show where the changes are. |
| SSIM / DSSIM | Structural resemblance matters more than exact pixel equality. | Choose and validate your own threshold against representative captures. Do not assume a universal pass value. |
| NCC / PSNR | You need a correlation or signal-oriented measure. | These metrics have their own interpretations and perfect-match conventions. Do not compare their raw scores as though they used the same scale. |
For example, to request a pixel-difference count, use AE:
magick compare -metric AE before.png after.png null:
To ignore small color differences when evaluating AE or PDC, set a fuzz tolerance. A percentage value is easier to communicate than an absolute quantum value, but the right tolerance depends on your images:
magick compare -metric AE -fuzz 1% before.png after.png difference.png
Fuzz changes what counts as different. Calibrate it against real pages and known changes; do not increase it just to make a failing comparison pass.
4. Handle size, offsets, and rendering noise
ImageMagick compares corresponding pixels directly, starting from each image’s page offset unless subimage search is enabled. The official compare documentation describes this pixel-by-pixel behavior. If one screenshot is larger, ImageMagick aligns the smaller image with the larger and treats the extra region as virtual pixels, which can affect the metric.
For screenshot regression, matching viewport and crop dimensions is usually the clearest solution. If you intentionally need to compare unequal images, the documented definition below limits comparison to authentic pixels:
magick compare -define compare:virtual-pixels=false -metric RMSE before.png after.png difference.png
Check dimensions and page offsets before interpreting a surprising score. Cropping or repaging can leave offsets that shift the comparison. ImageMagick’s defines reference documents the virtual-pixel setting and other comparison-related definitions.
Anti-aliasing, font rendering, operating system differences, and GPU or browser variation can change many exact pixel values without a meaningful visual layout change. Keep the capture environment consistent, and consider a carefully calibrated fuzz tolerance or a structural metric when exact equality is too strict. A threshold should be based on repeated runs and representative pages, not a generic number.
5. Automate a visual regression check
Use a fixed ImageMagick version, metric, and fuzz setting across your comparison runs. Preserve both the numeric result and the difference image. This shell example records RMSE while allowing the command’s comparison status to remain visible:
#!/usr/bin/env bash
set -euo pipefail
before="before.png"
after="after.png"
diff="difference.png"
score_file="rmse.txt"
magick compare -metric RMSE "$before" "$after" "$diff" 2>"$score_file" || status=$?
status=${status:-0}
printf 'ImageMagick compare status: %s\n' "$status"
printf 'RMSE: '
cat "$score_file"
# ImageMagick compare uses 0 for similar images, 1 for dissimilar images,
# and 2 for an error. Apply your project's own validated threshold here.
ImageMagick documents return status 0 for similar images, a value between 0 and 1 when they are not similar, and 2 on error. Treat the exit status as a similarity signal, not a replacement for your project’s chosen metric threshold and visual review. Check the exact behavior against the version installed in your environment before using status alone as a CI gate.
For automated checks:
- Define what the metric means for your project and keep it fixed.
- Validate a threshold using repeated baseline captures and examples of both acceptable rendering noise and real regressions.
- Store the before image, after image, difference map, metric, ImageMagick version, and fuzz setting with the build artifacts.
- Use a visual review path for failures; a single score cannot identify whether the changed pixels are a bug.
6. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
magick: command not found |
ImageMagick is not installed or its executable is not on PATH. |
Install ImageMagick for your system and verify with magick -version. |
| The score is unexpectedly high or the whole image looks changed. | Viewport, dimensions, crop, page offset, scroll position, or browser state differs. | Compare dimensions and offsets, then recapture both versions under matching conditions. |
| Many tiny differences appear around text and edges. | Anti-aliasing or font rendering differs between captures. | Use the same browser and rendering environment; if appropriate, calibrate a small -fuzz tolerance. |
| The numeric result is missing from redirected output. | The metric is written to standard error rather than standard output. | Redirect file descriptor 2, such as 2>metric.txt. |
| A CI job fails even though a difference image was created. | The command reports dissimilar images through its exit status. | Handle the documented status intentionally; decide pass/fail using a project-specific policy and preserve the artifact. |
| The difference map shows a broad blank or extra band. | The images have unequal dimensions or virtual-pixel handling affects the result. | Recapture to equal dimensions, or use -define compare:virtual-pixels=false when authentic-pixel comparison is intended. |
| A score changes between machines or releases. | ImageMagick versions, metrics, fuzz values, or capture environments differ. | Pin or record the environment and use the same metric and settings for comparisons. |
7. Performance, reliability, and cost
Comparison work scales with the number of pixels being processed, so full-page captures and large batches require more work than small viewport images. Start with the smallest capture area that answers the regression question, and retain full-page images when below-the-fold changes matter. Avoid repeatedly running a comparison with different metrics unless each answers a defined question.
ImageMagick compares files you provide; it does not make browser capture reproducible by itself. The reliability of the result depends on stable page state, consistent dimensions, and a controlled rendering environment. Difference images make diagnosis easier, while a numeric metric helps track a consistent signal over time.
ImageMagick is a command-line image tool; the dossier does not establish a per-comparison service charge. Account for the runtime and storage your own capture and CI setup uses. If browser setup and capture are the costly parts, a screenshot API can return the images for the comparison step.
8. Or skip the browser setup
ImageMagick compares screenshots after you have captured them. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media; it returns an image or PDF from one GET request. Capture both versions using the same URL and relevant settings, then compare the resulting PNG files with the commands above. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For a pixel comparison, request or convert both captures to PNG and keep viewport and capture settings identical. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
9. FAQ
Does ImageMagick tell me whether a visual change is a bug?
No. It identifies image differences and reports a metric. A developer still needs to inspect the affected area and decide whether the change is expected.
Should I use RMSE or AE?
Use AE or PDC when you want a count of changed pixels under a defined fuzz setting. Use RMSE when an average error summary is useful and larger channel differences should count more. Keep the metric and settings consistent across runs.
Can I compare screenshots with different dimensions?
You can, but alignment and virtual pixels can affect the result. For regression checks, equal dimensions are easier to interpret. Use compare:virtual-pixels=false when you intentionally want to limit the comparison to authentic pixels.
Is there one correct visual-regression threshold?
No. Rendering noise and page content vary. Establish a threshold from repeated captures in your own environment and confirm that it catches the changes your project cares about.


