ScreenshotNeo

BlogGuides

Visual Regression Testing for Regulatory Compliance

Learn where visual regression tests fit in regulated software, what evidence to retain, and why screenshot diffs do not establish compliance on their own.

By the ScreenshotNeo team4 October 202612 min read

What is visual regression testing for regulatory compliance?

Visual regression testing compares a rendered page or component with an approved baseline to identify visual changes. In a regulated development process, those comparisons can contribute evidence that a particular change was reviewed and that selected visual behavior was checked. A screenshot diff alone does not establish software validation, regulatory compliance, accessibility conformance, or an electronic-record audit trail.

Whether visual testing is appropriate or required depends on the product’s intended use, regulated context, jurisdiction, applicable rules, and your organization’s documented risk rationale. FDA’s February 2026 Computer Software Assurance guidance addresses software used in medical-device production or quality management systems and recommends a risk-based approach; it supersedes the September 2025 final guidance. It does not state that every regulated organization must use screenshot comparison. Read the FDA Computer Software Assurance guidance.

How a visual regression check works

  1. Choose the page, component, document, or state whose appearance matters.
  2. Run the application in a controlled capture environment and save a baseline associated with an identified build and configuration.
  3. Capture the same target after a change using the same viewport, browser, fonts, data, and relevant state.
  4. Compare the new rendering with the baseline using a documented comparison method and pass/fail criteria.
  5. Review differences, record whether each is intended, and retain the result and disposition.
  6. Update the baseline only through a controlled review that identifies who approved the change and why.

A pixel diff highlights changed pixels. It does not explain whether a change is harmful, intended, or functionally correct; a reviewer or additional automated checks must make that determination. Dynamic content, animation, timestamps, remote fonts, and asynchronous loading can make otherwise equivalent renders differ.

Where visual regression evidence can contribute

A visual comparison can help detect rendering changes in screens that matter to a workflow, such as a changed control position, truncated warning, missing label, unexpected color, or altered page layout. Teams may include the capture, comparison result, review, and disposition in a broader verification or regression record.

FDA’s 2002 General Principles of Software Validation discusses documented test procedures, input data, results, objective pass/fail decisions, reporting, regression-suitable test material, and documentation for testing tools appropriate to their intended use. Those principles do not automatically validate a particular screenshot vendor or prescribe screenshot diffs. The guidance also says, “Testing at the user site is an essential part of software validation.” Its context is the role of user-site testing within the software validation process.

Use a screenshot comparison as one test with a defined purpose. It cannot replace requirements traceability, risk analysis, functional testing, user-site evaluation where applicable, change control, or the other activities your validation approach requires.

Is visual regression testing required for compliance?

There is no blanket requirement in the cited FDA materials for all organizations to use visual regression testing. First determine which rules apply to your product and process. FDA’s Computer Software Assurance guidance is specifically about software used in medical-device production or quality management systems, and it recommends assessing assurance activities based on risk. Other products, jurisdictions, and uses can have different requirements.

Before deciding whether to use visual tests, document:

  • The system’s intended use and the workflow or decision it supports.
  • The regulated context, applicable jurisdiction, and rules that apply to the system or process.
  • The risks of an undetected visual change, including whether it could affect safety, product quality, record integrity, or a user decision.
  • Which screens or states need coverage and why visual comparison is an appropriate control for those risks.
  • What other verification, validation, accessibility, and record controls are needed.
  • How the test, tool, environment, and evidence will be controlled for their intended use.

This is a decision framework, not a substitute for regulatory or quality-system advice specific to your organization.

How do you document visual regression tests for an audit?

The following evidence checklist is a practical synthesis of the FDA documentation principles cited above. It is not a verbatim FDA checklist or a universal legal requirement. Adapt it to your system, risk analysis, and quality procedures.

  • Test identity and purpose: test case or run identifier, covered requirement or risk, target page/component, and reason for selecting it.
  • Controlled environment: application build, test data or state, browser and version, operating environment, viewport/device scale, locale, timezone, fonts, and capture-tool configuration.
  • Baseline identity: baseline image or artifact identifier, creation and approval history, associated build, and any relevant configuration.
  • Written procedure and inputs: steps to reach the target state, test inputs, prerequisites, waits, and capture instructions.
  • Expected outcome and criteria: what must be visually checked, comparison method, thresholds or exclusions, and objective pass/fail rules.
  • Results: captured output, comparison artifact, run status, errors or environmental deviations, and enough information to reproduce the run.
  • Review and disposition: reviewer identity, date, interpretation of each relevant difference, whether it was intended, and the reason for accepting or rejecting it.
  • Baseline changes: retained history showing who changed or approved a baseline, when, and the rationale and change reference.
  • Summary: scope, outcome, unresolved deviations, and links or references to associated verification records.

Keep evidence under the organization’s record-retention and access controls. A screenshot or diff may show a result, but it does not by itself prove who approved a baseline, preserve an immutable history, or establish the meaning of the result.

Visual diffs, accessibility, and electronic records are separate controls

Does screenshot testing prove WCAG compliance?

No. A screenshot can reveal some visual issues, such as clipping or low apparent contrast, but it cannot establish WCAG conformance. Many accessibility requirements concern semantics, keyboard operation, focus behavior, accessible names, programmatic relationships, or assistive-technology behavior that a static image does not expose. Accessibility needs its own evaluation using appropriate methods. Section508.gov describes accessibility testing methods and tools, and the W3C Accessibility Conformance Testing (ACT) Rules document rules used for testing conformance against WCAG.

Does a screenshot diff create an electronic-record audit trail?

No. A diff compares renderings; an audit trail records relevant actions or changes to electronic records. FDA’s Part 11 Scope and Application guidance recommends a risk-based, documented decision about audit trails based on predicate rules and potential impacts on quality, safety, and record integrity. Determine which record controls apply to your records and processes. Do not treat an image comparison as a substitute for those controls.

Build a small, reproducible visual check with Python

This example uses Playwright to render a URL and Pillow to compare the resulting PNG with a saved baseline. It is a simple engineering example, not a validated test system or a compliance recipe. Pin and control dependencies and browser versions in your own environment; use synthetic or approved test data.

python -m pip install playwright pillow
python -m playwright install chromium

Save as visual_check.py. The first run with --update-baseline creates a baseline. Later runs compare against it and exit with status 1 when the changed-pixel ratio exceeds the chosen threshold. Baseline creation and updates should follow your review process.

import argparse
import asyncio
from pathlib import Path

from PIL import Image, ImageChops
from playwright.async_api import async_playwright


async def capture(url: str, output: Path, selector: str | None) -> None:
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        page = await browser.new_page(
            viewport={"width": 1365, "height": 900},
            device_scale_factor=1,
            color_scheme="light",
            reduced_motion="reduce",
        )
        response = await page.goto(url, wait_until="networkidle", timeout=60000)
        if response is None or not response.ok:
            status = "no response" if response is None else str(response.status)
            await browser.close()
            raise RuntimeError(f"Page load failed: {status}")
        if selector:
            target = page.locator(selector)
            await target.wait_for(state="visible", timeout=15000)
            await target.screenshot(path=str(output), animations="disabled")
        else:
            await page.screenshot(
                path=str(output), full_page=True, animations="disabled"
            )
        await browser.close()


def compare(baseline_path: Path, current_path: Path, threshold: float) -> int:
    baseline = Image.open(baseline_path).convert("RGBA")
    current = Image.open(current_path).convert("RGBA")
    if baseline.size != current.size:
        print(f"FAIL: image dimensions differ: {baseline.size} vs {current.size}")
        return 1
    diff = ImageChops.difference(baseline, current)
    changed = sum(1 for pixel in diff.getdata() if pixel != (0, 0, 0, 0))
    total = baseline.width * baseline.height
    ratio = changed / total if total else 0.0
    print(f"Changed pixels: {changed}/{total} ({ratio:.4%}); threshold={threshold:.4%}")
    diff.save(current_path.with_name(current_path.stem + "-diff.png"))
    return 1 if ratio > threshold else 0


async def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("url")
    parser.add_argument("--baseline", type=Path, default=Path("baseline.png"))
    parser.add_argument("--current", type=Path, default=Path("current.png"))
    parser.add_argument("--selector", help="Capture one element instead of the full page")
    parser.add_argument("--threshold", type=float, default=0.001,
                        help="Maximum changed-pixel fraction, from 0 to 1")
    parser.add_argument("--update-baseline", action="store_true")
    args = parser.parse_args()
    if not 0 <= args.threshold <= 1:
        parser.error("--threshold must be between 0 and 1")

    args.current.parent.mkdir(parents=True, exist_ok=True)
    await capture(args.url, args.current, args.selector)
    if args.update_baseline or not args.baseline.exists():
        args.baseline.parent.mkdir(parents=True, exist_ok=True)
        args.current.replace(args.baseline)
        print(f"Baseline saved: {args.baseline}")
        return 0
    return compare(args.baseline, args.current, args.threshold)


if __name__ == "__main__":
    raise SystemExit(asyncio.run(main()))

Example commands:

# Capture and establish an initial baseline; review before treating it as approved.
python visual_check.py https://example.com --update-baseline

# Compare a subsequent rendering with that baseline.
python visual_check.py https://example.com --threshold 0.001

# Compare just a selected element.
python visual_check.py https://example.com --selector "main .checkout-summary"

The default threshold is an example setting, not a recommended regulatory acceptance limit. Choose comparison criteria based on the test purpose and risk. This basic exact-pixel method can be noisy across browser, OS, font, and rendering changes; it also does not mask dynamic regions or assess whether a visual change is acceptable.

Or skip the browser setup

For a screenshot capture without maintaining your own browser setup, ScreenshotNeo provides a screenshot API and MCP server. See the ScreenshotNeo API documentation. This call captures a page; baseline management, test criteria, review, and record controls remain part of your process.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Capture consistency, reliability, and cost

Keep the capture environment stable

  • Pin the browser and operating environment used for comparisons. Browser updates, font substitution, and operating-system rendering can produce diffs without an application change.
  • Fix viewport dimensions, device scale factor, color scheme, locale, timezone, and relevant feature settings.
  • Use deterministic test data and application state. Avoid personal data and external production dependencies where possible.
  • Wait for the condition the test needs. A fixed delay is simple but can be slow or flaky; a specific selector or application-ready signal is usually more meaningful. Network-idle conditions can be unreliable on pages with long-lived requests.
  • Disable or control animation and other time-dependent content. Masking or hiding dynamic regions should be documented so it does not conceal meaningful changes.
  • For full-page capture, check lazy-loaded content and page length. A page may need scrolling or application-specific setup before every section is rendered.

Plan capacity and cost around reviewed coverage

Self-hosted browser capture shifts cost toward compute, browser maintenance, CI time, and artifact storage. Hosted capture changes the cost model to service usage and may reduce browser infrastructure work. For either approach, estimate the number of URLs, viewports, states, retries, and scheduled runs; retain only the artifacts your review and record policy calls for. A screenshot API call is not itself a full visual-regression workflow: you still need baseline comparison and controlled review.

ScreenshotNeo pricing is Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. The service bills only clean shots; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Review the response’s X-Page-Verdict and X-Billed headers when building usage or failure handling.

Choosing and governing a visual testing tool

Vendor documentation can illustrate workflows, not guarantee regulatory suitability. Chromatic documents snapshots, baseline comparisons, and accessibility tests; Applitools describes rendered-output visual testing and baseline comparisons. These descriptions do not establish FDA approval or validation for your use. Evaluate any tool in your own deployment and intended context.

  • Coverage: component, full page, PDF, responsive viewport, and mobile states needed for the intended workflow.
  • Environment control: browser/rendering consistency, data handling, and ability to reproduce a capture.
  • Baseline governance: permissions, approval flow, history, and ability to explain updates.
  • Dynamic content: waits, animation controls, masking, and whether exclusions are visible and reviewable.
  • Evidence: exportability of images, diffs, run details, reviewer dispositions, and history into your record process.
  • Integration and controls: CI behavior, access controls, retention, security, and fit with your organization’s validation and change-control procedures.

Primary workflow references: Chromatic snapshots, Chromatic accessibility tests, and Applitools visual testing.

Troubleshooting visual regression checks

Symptom Likely cause What to do
Large diff with no apparent product change Browser, OS, font, viewport, device scale, or color scheme changed. Restore the recorded capture configuration and compare in the same environment. Treat intentional environment changes as a reviewed baseline migration.
Intermittent diffs on the same build Asynchronous rendering, animation, rotating content, current time, random data, or unstable network resources. Use deterministic fixtures, disable animation, wait for an explicit ready condition, and document any masked regions.
Blank or incomplete capture Navigation failed, a selector was absent, authentication was missing, or content had not rendered. Check the response and browser logs, confirm access and selector state, and wait for the application’s actual ready signal before capture.
Full-page image omits lower content Lazy-loaded content did not load until scrolling, or the page uses a virtualized list. Scroll through the page or trigger the app’s supported load behavior before capture; verify that the test state contains the expected content.
Every run fails after a legitimate design change The baseline no longer represents the approved expected rendering. Review the diff against the intended change, retain the old baseline and disposition as required by policy, then update through the controlled approval path.
Diff passes despite an important defect The tested viewport/state missed the affected behavior, or threshold/exclusions were too permissive. Add coverage for the relevant state and review criteria. Pair visual checks with functional and accessibility tests.
Screenshot result is being used as accessibility evidence A static image cannot evaluate keyboard, semantic, or assistive-technology behavior. Run a separate accessibility evaluation with appropriate methods and retain its results.
Diff is being treated as an audit trail The comparison artifact does not record all relevant electronic-record actions or approvals. Assess record and audit-trail controls independently under the applicable rules and documented risk rationale.

Frequently asked questions

Can a visual regression test pass while the application is functionally broken?

Yes. A page can look unchanged while a button, workflow, or data operation is broken. Pair visual checks with tests that exercise behavior and requirements.

Should every screen have a screenshot baseline?

Not necessarily. Select pages and states based on intended use and risk, then document the coverage rationale and gaps.

Can I use a hosted screenshot API for regulated testing?

Potentially, if it fits your intended use and your organization’s data, security, validation, and record controls. Assess the actual deployment and retain the evidence your process requires; a provider’s feature description alone does not establish suitability.

Does approving a new baseline mean the change is safe?

No. Approval records a disposition of the visual difference. Safety, functionality, accessibility, and other risks require the evaluations applicable to the change.

Try ScreenshotNeo

If you need a capture endpoint for a visual testing workflow, start with the API documentation. ScreenshotNeo can remove cookie banners, popups, and chat widgets before a shot; bot checks, blank pages, and failed loads are never billed; and its MCP server lets AI agents take screenshots. You get 1,000 screenshots a month free with no card, and paid plans start at $5 for 3,000. Sign up for free and get 1,000 screenshots a month with no card.