ScreenshotNeo

BlogGuides

How to Do Advanced Automated Visual UI Testing With Selenium

Build repeatable visual regression checks with Selenium: choose checkpoints, capture screenshots, compare approved baselines, and review changes without hiding real defects.

By the ScreenshotNeo team4 October 202613 min read

Selenium drives the browser and puts your application into a state; a visual comparison layer captures that state, compares it with an approved baseline, and makes differences reviewable. To build dependable visual UI tests, choose meaningful checkpoints, wait for the interface to settle, keep capture conditions consistent, and review baseline changes instead of automatically accepting them.

This guide builds a small, runnable Python example using Selenium and Pillow for pixel comparison. It also explains how to choose screenshot scope, handle dynamic content, run checks in CI, and when to use a managed visual testing service. Selenium’s [waiting strategies](https://www.selenium.dev/documentation/webdriver/waits/) documentation explains why page navigation completing does not necessarily mean a dynamic interface is ready.

1. What visual UI testing with Selenium checks

A functional test asks whether an interaction or result is correct. A visual regression test asks whether the rendered interface changed from an approved reference in a way that needs attention. A useful visual test normally has four parts:

  1. Arrange: provide controlled data, viewport, browser, and application state.
  2. Drive: use Selenium WebDriver to navigate and interact as a user would.
  3. Capture and compare: take a screenshot at a named checkpoint and compare it with a baseline.
  4. Review: decide whether a difference is an intended change, a regression, or capture noise.

Selenium provides the browser automation layer. The comparison and baseline review policy are separate concerns. The first run can create a reference image, but that image is not automatically correct: inspect it before treating it as an approved baseline.

Choose checkpoints users recognize

Prefer a small set of meaningful states over a screenshot after every command. Common checkpoints include the initial page, a navigation menu after opening, a form with validation errors, a dialog, an empty state, and a responsive layout at a supported viewport. Give each checkpoint a stable, unique name such as checkout-empty-cart-desktop.

2. Set up a local Selenium visual test

The example below uses Python, Selenium 4, Chrome, and Pillow. Selenium Manager can manage browser drivers for common local setups; your CI image still needs a compatible browser and system libraries. Install dependencies and save the script as visual_check.py:

python -m pip install selenium pillow

Set TARGET_URL to a page your test environment can reach. The readiness selector must refer to an element that indicates the particular view is ready. The script creates a baseline when explicitly asked, then compares subsequent captures and exits nonzero when the images differ beyond the configured threshold.

import os
import sys
from pathlib import Path

from PIL import Image, ImageChops, ImageStat
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

TARGET_URL = os.environ.get("TARGET_URL", "https://example.com")
READY_SELECTOR = os.environ.get("READY_SELECTOR", "h1")
CHECKPOINT = os.environ.get("CHECKPOINT", "home-desktop")
BASELINE_DIR = Path(os.environ.get("BASELINE_DIR", "visual-baselines"))
ARTIFACT_DIR = Path(os.environ.get("ARTIFACT_DIR", "visual-artifacts"))
ACCEPT_BASELINE = os.environ.get("ACCEPT_BASELINE") == "1"
# Mean absolute channel difference, from 0 (identical) to 255 (max difference).
MAX_MEAN_DIFFERENCE = float(os.environ.get("MAX_MEAN_DIFFERENCE", "0.5"))

baseline_path = BASELINE_DIR / f"{CHECKPOINT}.png"
actual_path = ARTIFACT_DIR / f"{CHECKPOINT}.png"
diff_path = ARTIFACT_DIR / f"{CHECKPOINT}-diff.png"
BASELINE_DIR.mkdir(parents=True, exist_ok=True)
ARTIFACT_DIR.mkdir(parents=True, exist_ok=True)

options = Options()
if os.environ.get("HEADLESS", "1") == "1":
    options.add_argument("--headless=new")
options.add_argument("--window-size=1365,900")
options.add_argument("--force-device-scale-factor=1")
options.add_argument("--disable-dev-shm-usage")

# Use explicit waits for specific interface conditions. Do not combine a
# nonzero implicit wait with these explicit waits.
driver = webdriver.Chrome(options=options)
try:
    driver.set_window_size(1365, 900)
    driver.set_page_load_timeout(45)
    driver.get(TARGET_URL)
    WebDriverWait(driver, 20).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, READY_SELECTOR))
    )

    # Add application-specific actions here, then wait for the resulting state.
    # Example: driver.find_element(By.CSS_SELECTOR, "[data-testid='open-menu']").click()
    # WebDriverWait(driver, 10).until(
    #     EC.visibility_of_element_located((By.CSS_SELECTOR, "[role='menu']"))
    # )

    # Optional targeted stabilization: wait for fonts and images to finish.
    # This does not replace an application-specific ready condition.
    WebDriverWait(driver, 10).until(
        lambda d: d.execute_script("return document.readyState") == "complete"
    )
    driver.execute_script(""
        return Promise.all(Array.from(document.images, image => {
          if (image.complete) return Promise.resolve();
          return new Promise(resolve => {
            image.addEventListener('load', resolve, {once: true});
            image.addEventListener('error', resolve, {once: true});
          });
        }));
    """)
    driver.save_screenshot(str(actual_path))
finally:
    driver.quit()

if not baseline_path.exists():
    if not ACCEPT_BASELINE:
        print(f"No baseline at {baseline_path}. Review {actual_path}, then rerun with ACCEPT_BASELINE=1 to create it.")
        sys.exit(2)
    Image.open(actual_path).save(baseline_path)
    print(f"Created baseline: {baseline_path}. Review and commit this image deliberately.")
    sys.exit(0)

baseline = Image.open(baseline_path).convert("RGBA")
actual = Image.open(actual_path).convert("RGBA")
if actual.size != baseline.size:
    print(f"Image dimensions differ: baseline={baseline.size}, actual={actual.size}")
    sys.exit(1)

difference = ImageChops.difference(baseline, actual)
# Save a contrast-enhanced diff artifact for diagnosis; comparison uses the
# original difference, not the amplified preview.
difference.save(diff_path)
mean_difference = sum(ImageStat.Stat(difference).mean[:3]) / 3
print(f"Mean absolute RGB difference: {mean_difference:.4f}; limit: {MAX_MEAN_DIFFERENCE:.4f}")
print(f"Actual: {actual_path}; diff: {diff_path}")
if mean_difference > MAX_MEAN_DIFFERENCE:
    print("Visual difference exceeds the configured threshold. Review the images; do not auto-accept the baseline.")
    sys.exit(1)
print("Visual comparison passed.")

The code captures the viewport, not an entire long page. Its mean pixel difference is a deliberately simple teaching example, not a perceptual visual-testing algorithm. It can flag harmless antialiasing differences and can also make a meaningful small defect look insignificant across a large image. For a production suite, use region-aware or perceptual comparison and a review interface, or integrate a visual testing service.

Create and approve a baseline

  1. Run the test without ACCEPT_BASELINE. It saves the actual capture and exits with a message that the baseline is missing.
  2. Inspect the capture for correct data, layout, and state. If it is right, run ACCEPT_BASELINE=1 python visual_check.py.
  3. Commit the baseline image with the test code. Keep the checkpoint name and test conditions stable.
  4. On later runs, inspect the actual and diff artifacts when the check fails. Update the baseline only after confirming the change is intentional.

3. Drive the application into repeatable states

Navigation’s readyState is not a reliable proxy for a modern application’s finished rendering. JavaScript may still update content or reveal elements after navigation returns. Selenium recommends synchronization around the condition the next action requires, and warns that mixing implicit and explicit waits can produce unpredictable timing. See [Selenium waiting strategies](https://www.selenium.dev/documentation/webdriver/waits/).

Use explicit, state-specific waits

Wait for a visible element, a spinner to disappear, a result count to stabilize, or an application-owned readiness marker. Prefer a condition that describes the state you need over a fixed sleep. A short fixed delay can be appropriate for a known animation or third-party transition, but it adds runtime and does not guarantee readiness under load.

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

wait = WebDriverWait(driver, 15)
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "[data-testid='results']")))
wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, ".loading-spinner")))
# Trigger the screenshot only after the target view is ready.

For a dynamic application you control, a test-only marker such as data-testid or a documented readiness flag is often more dependable than waiting for network silence. A page can continue making analytics or polling requests after its visible content is stable.

Control volatile content narrowly

Potential sources of noise include timestamps, randomized avatars, rotating promotions, live maps, ads, cursors, and animation. The strongest fix is deterministic test data or a stable test environment. If a region cannot be stabilized, mask only that region and document why; every ignored pixel is an area the check no longer protects.

  • Freeze or disable animation through a test stylesheet when appropriate.
  • Replace time-dependent text with a controlled clock or predictable fixture.
  • Wait for images and web fonts that affect layout when they matter to the checkpoint.
  • Use stable accounts and data, and reset state between tests.
  • Keep browser version, viewport, device scale, locale, and color scheme consistent across baseline and comparison runs.

Some visual products expose capture-time controls such as CSS injection, animation freezing, and ignored selectors. For example, Percy documents those options for its Selenium snapshot flow; confirm the current package and option names in the [Percy Selenium integration repository](https://github.com/percy/percy-selenium-python). Do not assume a vendor option works with a raw WebDriver screenshot.

4. Choose viewport, element, or full-page coverage

A viewport screenshot checks what is visible at a specific scroll position and browser size. An element screenshot narrows the comparison to a component, which can make a focused test easier to interpret. A full-page screenshot checks more content but may take longer and can behave differently around sticky headers, lazy-loaded sections, or elements that move during scrolling.

Capture scope Useful for Trade-off
Viewport Common user view, dialogs, responsive breakpoints Content outside the viewport is not checked
Element Stable component or isolated widget Can miss surrounding layout and overlap defects
Full page Long pages where below-the-fold layout matters More pixels; scrolling and sticky content can create artifacts

WebDriver’s screenshot commands and third-party capture tools do not all offer identical full-page behavior. Verify the selected browser and integration rather than assuming a viewport capture covers the document. Percy documents a full_page option in its integration; the implementation and support are tool-specific. Applitools’ [screenshot guidance](https://applitools.com/blog/full-page-screenshots-in-selenium/) illustrates the longstanding scrolling and stitching caveat, but it should not be read as a universal statement about every current browser or service.

For responsive testing, capture at the viewports that correspond to real supported layouts. A desktop baseline does not establish that mobile rendering is correct. Selenium Grid can distribute browser sessions across machines; Selenium describes Grid as a way to run tests against different browsers and platforms. See the [Selenium project documentation](https://www.selenium.dev/documentation/).

5. Keep visual comparisons useful

Interpret diffs, not just pass or fail

A visual mismatch is a signal for review. Check whether the difference is localized to the intended component, whether the page reached the same state, and whether the test ran with matching browser and viewport conditions. Keep the baseline image, actual image, and diff together as CI artifacts so the failure can be diagnosed without rerunning locally.

Do not set a generous global threshold just to make flaky checks pass. Thresholds can tolerate rendering noise, but broad tolerance can hide subtle regressions. If using raw pixel comparison, record the metric and threshold; use targeted regions or a visual engine when a single whole-image score is too blunt.

Make baseline changes reviewable

Treat baseline changes like code changes. Include them in the same pull request as the intentional UI change, and review the before/after images. Separate expected design updates from unexpected diffs. If a baseline change has no clear product change behind it, investigate browser version, test data, fonts, network state, and timing before accepting it.

6. CI, parallel runs, performance, and reliability

  • Pin the capture environment: use a consistent browser image and viewport. A browser upgrade can alter rendering and should trigger a deliberate baseline review.
  • Isolate test data: parallel workers should not edit the same account or record if that changes the captured page.
  • Use unique artifact names: include test, browser, viewport, and build identifiers so concurrent runs do not overwrite each other.
  • Retry carefully: retries can help identify intermittent infrastructure issues, but a test that passes only on retry still needs investigation.
  • Keep artifacts on failure: preserve baseline, actual, diff, browser logs, and checkpoint context.
  • Limit capture volume: begin with high-value screens and add checkpoints where they protect a user-visible workflow.

Screenshot capture and comparison cost grows with the number and size of images, browser sessions, and repeated runs. Full-page captures and broad cross-browser matrices take more time and storage than a small viewport suite. Measure your own CI duration and storage needs; the dossier does not provide a neutral benchmark or pricing comparison.

Selenium WebDriver BiDi adds a WebSocket connection for browser events such as network activity, console messages, and JavaScript errors. It can improve diagnostics for event-aware tests, but screenshot comparison does not require BiDi. Selenium notes that BiDi support is being implemented; consult the [current BiDi documentation](https://www.selenium.dev/documentation/webdriver/bidi/) before depending on a particular event API.

7. Troubleshooting common failures

Symptom Likely cause Fix
Element missing or screenshot is blank Capture ran before the app rendered, or navigation reached an error page Wait for a visible app-specific marker; inspect URL, page title, and browser logs before capture.
Intermittent timeout Race condition, unstable network, or a wait condition that never becomes true Wait for the actual UI state, use a realistic timeout, and capture diagnostics on timeout. Selenium identifies poor synchronization as a common source of errors in its [troubleshooting guide](https://www.selenium.dev/documentation/webdriver/troubleshooting/).
Images differ on every run Dynamic content, animation, random data, or inconsistent browser/device scale Stabilize data and capture settings first; narrowly mask irreducible volatile regions.
Entire image shifts or has a different size Viewport, browser chrome, device scale factor, or responsive breakpoint changed Set the same window size and scale factor, and check that both runs use the same browser configuration.
Only text edges differ Font availability, browser build, operating system, or antialiasing changed Use a consistent CI image, wait for fonts where relevant, and avoid overbroad thresholds.
Full-page screenshot has repeated or misplaced sticky content Scrolling/stitching interacts with fixed elements or lazy loading Try viewport or element scope, or use a capture integration with documented full-page support; verify the result in the target browser.
Driver session fails to start Browser/driver mismatch, missing browser dependencies, or unsupported headless setup Check installed browser and Selenium versions, CI libraries, and driver logs; reproduce in a second browser if possible.
Baseline is absent in CI Baseline images were not committed or the working directory differs Commit reviewed baseline files and use an explicit project-relative baseline path.
Baseline updates unexpectedly Test automatically overwrites the approved reference Require an explicit baseline-accept action and review the image change in version control.

8. Optional visual testing integrations

If the team needs centralized baseline review or managed visual comparison, use the same evaluation questions for each integration: does it fit your test language and runner; support your needed viewport, element, and full-page scopes; provide controls for animation and volatile regions; make baseline approval reviewable; cover the browsers you run; meet your storage and privacy needs; and fit your CI and cost constraints?

Applitools documents a Java Selenium quickstart that runs a Visual AI test and reviews results; it requires an account and API key. See its [Selenium Java quickstart](https://applitools.com/docs/eyes/sdks/selenium-java/quickstart). Percy documents a Python Selenium integration and capture options in its [repository](https://github.com/percy/percy-selenium-python). Those examples establish that integrations exist, not which service is superior or what current pricing is. Verify current options and terms with each vendor.

9. Or skip the browser setup

If your immediate need is a page screenshot rather than Selenium-driven interaction, ScreenshotNeo can capture a URL with one API request. Selenium remains useful when the test must log in, click through a workflow, or establish a particular application state. ScreenshotNeo is a website screenshot API and MCP server from [ScreenshotNeo](https://screenshotneo.com); its [API documentation](https://screenshotneo.com/docs/) covers request options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: process.env.SCREENSHOTNEO_API_KEY,
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Use Selenium when the capture requires browser interaction; use a direct screenshot API for URL-based capture.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

10. Frequently asked questions

Does Selenium compare screenshots by itself?

Selenium drives the browser and can capture screenshots. You need comparison logic or a visual testing integration to compare captures against approved baselines and review differences.

Should every changed screenshot fail the build?

It should create a reviewable signal. Your team should decide whether the change is intended and update the baseline only after reviewing it.

Do I need WebDriver BiDi for visual regression testing?

No. BiDi can stream browser events useful for diagnostics, while screenshot comparison can use the regular WebDriver workflow.

Can one browser baseline cover every browser?

No. Rendering can vary by browser and environment. Run captures in the browsers and viewports that matter to your users.

Sources