Visual Regression Testing with Selenium
Learn how to build reliable Selenium visual regression checks with screenshots, baselines, diff review, stabilization, and CI troubleshooting.

Visual regression testing with Selenium means capturing a known UI state, comparing that screenshot with an accepted reference image, and reviewing any differences before deciding whether to update the reference. Selenium drives the browser and puts the application into the right state. A separate image-comparison and baseline workflow determines what changed. A difference is evidence for review, not automatic proof of a defect.
On the first run, the screenshots you approve become baselines. On later runs, new screenshots are compared with those stored images. If a change is intentional, accept it and save a new baseline. If it is a regression, reject it and keep the previous baseline. This article shows a complete Selenium workflow, explains how to make captures repeatable, and covers review, CI, troubleshooting, performance, and cost.
How the Selenium visual regression workflow works
A useful test has five distinct stages:

- Arrange: start the browser with a known viewport, browser version, locale, timezone, data set, and authentication state.
- Act: use Selenium WebDriver to navigate and interact with the application.
- Checkpoint: wait until the meaningful UI state is stable, then capture a screenshot.
- Compare: compare the new image with the accepted baseline using an image-diff workflow or visual-testing service.
- Review: inspect the diff. Accept an intentional product change as a new baseline, or reject a defective image and retain the old reference.
Choose checkpoints that represent user-visible states: a logged-in dashboard, a completed checkout step, an error state, or a responsive navigation menu. A screenshot taken halfway through an animation or before data has loaded creates noise instead of useful coverage.
Selenium supplies browser automation and context handling. Its WebDriver documentation covers controlling browser windows and tabs, which is useful when a checkpoint opens a new context. Read the Selenium windows and tabs documentation for the browser-control details.
Build a deterministic Selenium screenshot test
The example below uses Python, Selenium, and Pillow. It captures a checkpoint image and compares it with a baseline using a pixel difference. This is a deliberately small implementation that you can run locally or adapt to a CI suite.
1. Install the dependencies
python -m pip install selenium pillow
Your machine also needs a browser such as Chrome or Firefox. Recent Selenium versions can manage the matching driver through Selenium Manager, or you can provide a driver explicitly in your environment.
2. Capture and compare a checkpoint
from pathlib import Path
from time import sleep
from PIL import Image, ImageChops
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URL = 'https://example.com/dashboard'
BASELINE = Path('visual-baselines/dashboard.png')
CURRENT = Path('visual-current/dashboard.png')
DIFF = Path('visual-diffs/dashboard.png')
def capture_dashboard():
options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,1000')
options.add_argument('--force-device-scale-factor=1')
options.add_argument('--disable-gpu')
driver = webdriver.Chrome(options=options)
try:
driver.get(URL)
wait = WebDriverWait(driver, 20)
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, '[data-testid="dashboard"]')))
# Wait for application-specific loading to finish.
wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, '[data-testid="loading"]')))
driver.execute_script('window.scrollTo(0, 0)')
sleep(0.25) # allow a final paint after the state transition
CURRENT.parent.mkdir(parents=True, exist_ok=True)
driver.save_screenshot(str(CURRENT))
finally:
driver.quit()
def compare_with_baseline():
if not BASELINE.exists():
BASELINE.parent.mkdir(parents=True, exist_ok=True)
BASELINE.write_bytes(CURRENT.read_bytes())
print(f'Created baseline: {BASELINE}')
return
expected = Image.open(BASELINE).convert('RGBA')
actual = Image.open(CURRENT).convert('RGBA')
if expected.size != actual.size:
raise AssertionError(f'Screenshot sizes differ: {expected.size} vs {actual.size}')
diff = ImageChops.difference(expected, actual)
bbox = diff.getbbox()
if bbox:
DIFF.parent.mkdir(parents=True, exist_ok=True)
diff.save(DIFF)
raise AssertionError(f'Visual difference detected; inspect {DIFF}')
if __name__ == '__main__':
capture_dashboard()
compare_with_baseline()
Run it once to create the baseline, review that image, and commit it to the repository or baseline store. Run it again after code changes. A real project usually adds a tolerance for antialiasing and a richer diff report rather than requiring every pixel to match exactly.
3. Capture an element instead of the whole viewport
Full-page screenshots are useful for page-level checks, but a component checkpoint can be less sensitive to unrelated changes. Selenium can capture an element by locating it and using its screenshot method:
card = driver.find_element(By.CSS_SELECTOR, '[data-testid="pricing-card"]')
card.screenshot('visual-current/pricing-card.png')
Use a stable selector such as a test ID. Avoid selectors based on generated class names that change when the build tool emits a new bundle.
Make screenshots repeatable
Most false positives come from nondeterministic rendering rather than a meaningful product change. Stabilize the inputs around each checkpoint.
| Source of variation | What to do |
|---|---|
| Viewport and device scale | Set the same window size and device scale factor on every run. |
| Fonts | Install the same fonts in local and CI images. Wait for document.fonts.ready. |
| Animations | Disable transitions and animations with test-only CSS, or wait until the animation ends. |
| Async data | Use deterministic fixtures, seed test data, and wait for a specific loaded marker. |
| Images | Wait for important images to finish loading. Use fixed test assets where possible. |
| Time and dates | Freeze the application clock or configure a fixed timezone and test date. |
| Randomness | Seed random data and remove rotating content from the checkpoint. |
| Consent and overlays | Set the consent state deliberately and close chat, newsletter, and development overlays. |
| Browser differences | Keep browser, operating-system, and rendering versions consistent for a baseline set. |
One practical stabilization snippet disables motion and waits for fonts:
driver.execute_script("""
const style = document.createElement('style');
style.textContent = `*, *::before, *::after {
animation: none !important;
transition: none !important;
caret-color: transparent !important;
}`;
document.head.appendChild(style);
""")
driver.execute_async_script("""
const done = arguments[arguments.length - 1];
if (document.fonts && document.fonts.ready) {
document.fonts.ready.then(() => done());
} else {
done();
}
""")
Apply this only when disabling motion matches the state you intend to protect. A motion-specific test should capture the application with motion enabled and use a different assertion strategy.
Baseline management and visual review
Keep baselines versioned and name them by the state they represent, for example checkout-payment-chrome-1440.png. Store the browser, viewport, and test-data assumptions alongside the image so a reviewer can reproduce the context.

When a diff appears:
- Open the expected image, actual image, and diff overlay together.
- Identify whether the changed region is intentional, such as a redesigned button, or accidental, such as a missing font or shifted layout.
- Check the test environment before changing the baseline. A different browser build or viewport can create broad differences.
- Accept the new image only when the product change is deliberate and reviewed.
- Reject the image and keep the existing baseline when the difference indicates a defect.
- Commit approved baseline updates with the application change so the reason for the update is traceable.
A passing comparison proves consistency with the selected baseline and conditions. It does not prove that every page, browser, accessibility behavior, or business rule is correct.
Using a visual-testing service with Selenium
You can maintain image comparison and baseline storage in your project, or add a visual-testing service to the Selenium suite. Applitools documents Selenium SDKs for Java, C#, JavaScript, Python, and Ruby and describes a checkpoint and baseline workflow. Its overview defines visual testing as regression testing that checks whether previously correct screens changed unexpectedly. See the Applitools visual UI testing overview and SDK documentation for its stated integration options.
Evaluate tools on the dimensions that affect your team:
- Selenium integration: does the SDK fit the language and test runner you already use?
- Review workflow: can reviewers inspect a clear diff and accept or reject a baseline change?
- Execution scope: do you need one browser and viewport or a defined matrix?
- Storage and operations: will your team maintain image artifacts, retention, permissions, and cleanup?
- Failure reporting: can CI expose the expected, actual, and diff images to the person fixing the test?
The reviewed sources establish the checkpoint and baseline model, but they do not provide an independent comparison of vendor coverage, pricing, or maintenance effort. Treat those as project-specific evaluation questions.
Or skip the browser setup
If your goal is to obtain stable page screenshots for visual checks, documentation, or review artifacts, ScreenshotNeo provides a website screenshot API. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. The API can capture a full page, a CSS-selected element, a chosen device or viewport, dark mode, retina scale, custom CSS and JavaScript, a click before capture, selector or network-idle waits, headers, cookies, user agents, authorization, timezone, geolocation, resource blocking, caching, signed links, asynchronous jobs, webhooks, and bulk requests.
Use the ScreenshotNeo API documentation for the complete parameter list. A minimal cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
CI design for reliable visual checks
- Run functional setup and authentication before the visual checkpoint.
- Use a fixed browser container or runner image so fonts and rendering libraries remain stable.
- Publish expected, actual, and diff artifacts when a check fails.
- Make baseline approval a deliberate review step rather than an automatic retry.
- Separate visual failures from infrastructure failures such as a browser crash or unavailable test environment.
- Run a small smoke set on every change and a wider browser or viewport matrix on a scheduled job when runtime is significant.
Retries can hide flaky infrastructure, but they should not silently approve a different screenshot. Record whether a retry reproduced the same visual difference.
Performance, reliability, and cost considerations
Performance
Browser startup is often the largest cost in a Selenium screenshot test. Reuse a driver for related checkpoints when isolation permits, but reset application state between tests. Keep the checkpoint page focused, avoid unnecessary third-party resources, and wait for a precise readiness condition instead of a long fixed sleep.
Reliability
Pin browser and operating-system versions for a baseline set. Keep test data deterministic and review broad diffs as possible environment drift before changing application code. When a page contains intentionally dynamic content, hide or replace it in the test environment rather than increasing the diff tolerance until defects disappear.
Cost
Self-managed Selenium consumes CI minutes, browser infrastructure, image storage, and engineering time for diff tooling and review. A hosted capture API can reduce browser maintenance, but account for request volume, image retention, and the options your test requires. With ScreenshotNeo, only clean shots are billed, while failed loads, bot checks, blank pages, timeouts, and cache hits are not billed. Select caching and a TTL when repeated captures do not need a fresh browser render.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
TimeoutException waiting for an element |
The selector is wrong, the page is still loading, or the test data did not reach the expected state. | Verify the selector, add an application-specific readiness marker, and capture the page source and console logs on failure. |
| Large diff after a harmless change | Different fonts, browser versions, viewport size, device scale, or operating system. | Pin the environment and compare screenshot dimensions before inspecting individual pixels. |
| Only timestamps or ads differ | Dynamic content is included in the checkpoint. | Freeze time, use fixtures, block the resource, or hide the region with test-only CSS. |
| Screenshot is taken before content appears | A fixed sleep was too short or the wrong element was used as the readiness signal. | Wait for a specific visible element, loading marker to disappear, network-idle condition, or completed image load. |
| Blank or partially rendered image | Navigation failed, authentication expired, a script crashed, or the browser closed too early. | Check the URL, status and console logs, wait for the authenticated marker, and keep the driver alive until the screenshot is written. |
| Element screenshot has the wrong size | Responsive layout changed because the element was outside the intended viewport or fonts had not loaded. | Set the viewport explicitly, scroll the element into view, and wait for fonts before capture. |
| CI fails but local passes | CI uses different fonts, locale, timezone, browser, or data. | Run the same container locally, print environment details, and make locale and timezone explicit. |
| Screenshot API returns a bot-check result | The target site presented a CAPTCHA or bot challenge. | Inspect the response verdict headers and treat it as an unavailable capture rather than approving a blank baseline. |
Checklist before approving a baseline
- The checkpoint represents a meaningful user-visible state.
- Viewport, browser, scale factor, fonts, locale, timezone, and data are known.
- Animations, loading indicators, overlays, and dynamic content are handled intentionally.
- The expected, actual, and diff images were reviewed together.
- The change is tied to an intentional product or test change.
- The baseline name identifies the page state and environment.
- CI will publish enough artifacts to diagnose the next failure.
FAQ
Is Selenium itself a visual regression tool?
No. Selenium automates the browser and can capture screenshots. You still need image comparison and a process for storing, reviewing, accepting, or rejecting baselines.
Should every Selenium test take a screenshot?
No. Capture states that protect important visual behavior. Too many arbitrary checkpoints create maintenance work and make meaningful failures harder to find.
Can a pixel-perfect comparison work across browsers?
It is safest to maintain separate baselines for materially different rendering environments. Browser, operating system, fonts, and scale factor can change pixels without a product defect.
When should I update a baseline?
Update it after a reviewer confirms that the visual change is intentional and the screenshot was taken in the correct, stable state.
What should a failed visual test retain?
Retain the old accepted baseline and publish the new screenshot and diff as artifacts. This preserves the last known-good reference while the change is investigated.


