ScreenshotNeo

BlogHow-to

Compare Webpage Screenshots in Python with Pixel Differences

Capture consistent webpage screenshots with Playwright for Python, compare them with pixelmatch, and inspect pixel differences without mistaking rendering noise for regressions.

By the ScreenshotNeo team4 October 20269 min read

To compare webpage screenshots in Python, capture the same page state twice with Playwright for Python, then compare the images with a pixel-difference tool such as pixelmatch. Keep the browser, operating system, viewport, device scale factor, and page data consistent; save a diff image and review it before deciding whether a change is a regression. Pixel thresholds are project policy, not universal defaults.

This guide builds a runnable workflow for a reference image and a current image, explains exact versus tolerant comparisons, and covers noise, baselines, common failures, and cost. The comparison examples use pixelmatch; check the package’s current Python and Pillow compatibility before adopting it in a long-lived project. [pixelmatch on PyPI]

1. Install the tools

Use Python 3 and a Playwright-managed browser. Create a virtual environment if this is for a project, then install the Python packages and Chromium:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install playwright pixelmatch Pillow
python -m playwright install chromium

Playwright’s Python screenshot API can write screenshots to files or return image bytes. It supports viewport, full-page, and locator-based captures, so you can choose the scope that matches the change you want to detect. [Playwright Python screenshots]

2. Capture comparable screenshots

The capture conditions must match. Set the viewport and device scale factor explicitly, wait for a meaningful ready condition, and avoid comparing a reference from one browser environment with a current capture from another.

# capture.py
import asyncio
from pathlib import Path
from playwright.async_api import async_playwright

URL = "https://example.com"

async def capture(path: str) -> None:
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page(
            viewport={"width": 1440, "height": 900},
            device_scale_factor=1,
            color_scheme="light",
            locale="en-US",
            timezone_id="UTC",
        )
        await page.goto(URL, wait_until="networkidle", timeout=60_000)
        # Prefer an app-specific readiness signal when available.
        await page.locator("main").wait_for(state="visible", timeout=15_000)
        await page.screenshot(path=path, full_page=True, animations="disabled")
        await browser.close()

async def main() -> None:
    Path("artifacts").mkdir(exist_ok=True)
    await capture("artifacts/current.png")

if __name__ == "__main__":
    asyncio.run(main())

Run python capture.py. Capture the reference using the same function and environment, saving it as artifacts/baseline.png. For a first baseline, capture and review the image deliberately; do not treat the first render as an automatically approved expected result.

Choose the capture scope

  • Viewport: omit full_page=True to capture the visible viewport. This is faster and useful for above-the-fold layouts.
  • Full page: use full_page=True to capture the entire scrollable page. Long pages can consume more memory and capture time.
  • Element: target a locator and call its screenshot method, for example await page.locator("#pricing").screenshot(path="pricing.png"). This isolates a component from unrelated page changes.
  • In memory: call image_bytes = await page.screenshot() and pass the returned bytes to an image decoder or image-diff library. This avoids an intermediate screenshot file.

Playwright documents these capture forms and notes that screenshots can be post-processed or passed to third-party pixel-diff facilities. [Playwright Python screenshots]

3. Compare the images and write a diff

The following script loads both files as PIL images, checks that their dimensions match, and asks pixelmatch to produce a visual diff. Its Python package listing describes PIL image support, anti-aliased-pixel detection, and perceptual color-difference metrics. [pixelmatch on PyPI]

# compare.py
from pathlib import Path
from PIL import Image
from pixelmatch import pixelmatch

baseline_path = Path("artifacts/baseline.png")
current_path = Path("artifacts/current.png")
diff_path = Path("artifacts/diff.png")

baseline = Image.open(baseline_path).convert("RGBA")
current = Image.open(current_path).convert("RGBA")

if baseline.size != current.size:
    raise SystemExit(
        f"Image sizes differ: baseline={baseline.size}, current={current.size}. "
        "Use the same viewport, device scale factor, and capture scope."
    )

width, height = baseline.size
diff = Image.new("RGBA", baseline.size)

# threshold is a perceptual color-distance tolerance. The value is a policy
# choice for this page; don't assume one value suits every project.
differing_pixels = pixelmatch(
    baseline,
    current,
    diff,
    width,
    height,
    threshold=0.1,
    includeAA=False,
)
diff.save(diff_path)

pixel_count = width * height
ratio = differing_pixels / pixel_count if pixel_count else 0.0
print(f"Different pixels: {differing_pixels}/{pixel_count} ({ratio:.4%})")
print(f"Diff image: {diff_path}")

# Example only: choose and validate a project-specific allowed count.
max_allowed = 100
if differing_pixels > max_allowed:
    raise SystemExit("Visual comparison exceeded the configured pixel allowance")

Run python compare.py. Inspect artifacts/diff.png alongside both input images. A number alone does not tell you whether a changed region is a defect, an intended update, or harmless rendering variation.

Interpret threshold and allowed pixel count separately

The per-pixel threshold controls how much perceived color difference is tolerated before a pixel counts as changed. An allowed-difference count such as max_allowed controls how many changed pixels the whole image may contain. A higher color threshold can conceal subtle real changes; a larger count can permit a broad region to change. Playwright’s visual comparison documentation exposes analogous threshold and maxDiffPixels controls, but its values are not automatically appropriate defaults for this Python workflow. [Playwright visual comparisons]

For strict deterministic captures, use exact equality or a zero allowed-pixel count. If harmless variation is expected, measure known examples, inspect the diff, and set the smallest allowance that keeps the check useful. Keep threshold and pixel-count decisions explicit in code review.

4. Make captures repeatable

Screenshot comparison is only useful when the same input produces a sufficiently similar render. Control the page as well as the browser:

  • Use fixed test data and a stable account or fixture state.
  • Freeze clocks and avoid content driven by current time where the app permits it.
  • Disable or control carousels, random records, animations, and rotating promotions.
  • Wait for a meaningful application signal, such as a key locator becoming visible. Avoid relying only on a fixed sleep; it can be both slow and too short.
  • Use the same browser version, operating system, fonts, viewport, device scale factor, locale, timezone, color scheme, and headless configuration.
  • Hide only known volatile regions when controlling the source is impractical. Playwright Test’s visual comparison guide documents a style-path option for hiding volatile areas. The Python capture flow can instead inject a stylesheet before capture with page.add_style_tag, scoped to the known unstable element.
  • Wait for lazy-loaded content when capturing a full page. Scroll through the page or use an app-specific loading signal before taking the screenshot, then confirm that expected images have loaded.

Rendering can vary with host operating system, software versions, settings, hardware, power source, and headless mode. Consistent environments reduce noise; they do not guarantee identical output across every machine. [Playwright visual comparisons]

5. Use the comparison in a regression check

A practical CI check has three artifacts: the approved baseline, the newly captured image, and a diff image on failure. Store baselines with the code or in the team’s chosen artifact system, and make baseline updates an intentional review step.

  1. Capture the current page in the controlled browser environment.
  2. Compare it with the checked-in or otherwise approved baseline.
  3. On failure, retain both images and the diff so a reviewer can identify the changed area.
  4. If the design change is intentional, inspect the current capture and diff, then update the baseline in a reviewed change.

Playwright Test has an official toHaveScreenshot() visual assertion and a separate snapshot-update workflow. That assertion is part of Playwright Test; do not assume it is automatically available as a Python assertion. The Python documentation describes capturing images and passing them to a third-party diff facility. [Playwright visual comparisons] [Playwright Python screenshots]

6. Common problems and fixes

Symptom Likely cause Fix
Images have different dimensions Viewport, device scale factor, full-page state, or page content changed the capture size. Set viewport and device scale factor explicitly; ensure both captures use the same scope and stable page content.
Diff is noisy on every run Dynamic content, animation, font differences, asynchronous layout, or different browser/host environment. Control data and page state, disable animation, wait for app readiness, and run comparisons in the same environment.
Screenshot is blank or incomplete Navigation finished before the application rendered, a readiness locator was wrong, or the page requires authentication. Check the URL and login state, wait for an app-specific element, and save or inspect the capture before diffing.
Full-page image omits lazy images Images load only when scrolled into view. Trigger the page’s lazy-load behavior by scrolling through it and wait for expected images before capture.
pixelmatch import or argument error Installed package API or dependency versions differ from the example’s expected interface. Check the installed package version and its PyPI documentation, then pin compatible versions in the project environment.
Diff count is always zero but output looks different The files may be identical in the compared region, comparison inputs may be swapped or stale, or tolerance may be too permissive. Verify file paths and timestamps, lower the color threshold, and compare the saved inputs directly.
CI differs from a developer laptop Operating system, fonts, browser build, headless mode, or rendering hardware differ. Capture and compare in one pinned CI image or standard environment; avoid cross-environment baselines.

7. Performance, reliability, and cost

For a small set of pages, local Playwright capture and image comparison need no screenshot-service fee, though they use browser time, storage, and CI capacity. Full-page captures and large images cost more memory and processing than a small element capture. Reusing a browser process for multiple pages can reduce startup overhead, while keeping each page’s viewport and state explicit.

Reliability depends on deterministic inputs and retained failure artifacts. A screenshot that silently times out or captures before the page is ready can make a comparison misleading, so fail clearly when navigation or a readiness condition fails. Save the baseline, current image, and diff for diagnosis. Do not loosen thresholds simply to make a flaky test pass.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request can return an image or PDF; for this comparison workflow, capture the same URL and options for each run, then compare the returned images with your existing Python diff step. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o current.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("current.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('current.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Is pixel-by-pixel comparison the same as visual regression testing?

It is one comparison technique within a visual regression workflow. A useful workflow also controls capture conditions, retains baselines and diffs, and requires review of intentional changes.

Can I use Playwright’s screenshot assertion directly from Python?

The cited toHaveScreenshot() assertion is documented for Playwright Test. In Python, use Playwright to capture and connect the image output to a Python comparison step.

Should I compare whole pages or selected components?

Compare the smallest scope that covers the behavior under review. Element captures reduce unrelated changes; full-page captures catch changes outside the initial viewport.

What threshold should I start with?

There is no universal correct threshold. Start with strict comparison in a stable environment, inspect real diffs, then allow only measured rendering variance that your team considers harmless.