How to test whether a website screenshot change is caused by a browser update
Compare old and updated browser versions under controlled conditions to find out whether a browser update explains a visual regression.
A changed screenshot does not, by itself, show that a browser update caused the change. To test that hypothesis, capture the same page with the old and updated versions of the same browser while holding the operating system, test setup, page state, viewport, and capture settings fixed. Repeat the captures, compare each version with the known-good baseline, and, if possible, switch back to the old version to see whether the earlier rendering returns.
If the difference consistently appears with the updated browser and disappears after reverting, the evidence supports a browser-version-related visual regression. It does not prove that the browser engine has a defect: the browser may interact with the page, fonts, graphics hardware, or test harness.
1. Preserve the failure and record the setup
Before changing browser versions or refreshing expected images, save the unexpected screenshot and the known-good baseline. Record the URL or test route, the test steps, and any fixture or account state needed to reproduce the page. Do not update the expected screenshot yet; doing so can replace the evidence you need to investigate.
For each run, record:
- Browser product and channel, such as Chromium, Chrome, or Edge, and its exact version.
- Automation framework and version. Record this separately from the browser version because Playwright’s managed browser binaries are tied to Playwright releases.
- Operating system, container image, and, where practical, CI host or hardware.
- Headless or headed mode, viewport dimensions, and device scale factor.
- Test code, browser settings, fonts, page data, capture options, and timing or wait conditions.
Rendering can vary with the host operating system and its version, browser settings, hardware, power source, and headless mode. For consistent screenshots, Playwright recommends running tests in the same environment used to generate the baseline: Playwright visual comparisons.
2. Run a controlled old-versus-new comparison
Compare two versions of the same browser. Change the browser version while keeping the rest of the capture environment and inputs constant. A comparison between different browser families can reveal cross-browser differences, but cannot establish that a recent update caused the change.
Playwright projects can run tests across different configurations and browsers: Playwright projects. For this diagnosis, configure two runs of the same browser family with the old and updated browser binaries. Keep the test code and framework version fixed if possible. If updating Playwright is also necessary to obtain a new browser binary, record that as a second changed variable and avoid attributing the result to the browser version alone.
Runnable Playwright example
The following JavaScript test captures the same route twice using two Chromium projects. Configure each project’s browser executable or container so that chromium-old launches the previous version and chromium-new launches the updated version. The exact browser installation mechanism depends on your CI image. Keep the OS image and Playwright package version identical between jobs.
// tests/visual.spec.js
const { test, expect } = require('@playwright/test');
test('page matches the visual baseline', async ({ page }) => {
await page.goto('https://example.com', { waitUntil: 'networkidle' });
await expect(page).toHaveScreenshot('page.png', {
fullPage: true,
animations: 'disabled',
});
});
// playwright.config.js
const { defineConfig } = require('@playwright/test');
module.exports = defineConfig({
projects: [
{ name: 'chromium-old', use: { browserName: 'chromium' } },
{ name: 'chromium-new', use: { browserName: 'chromium' } },
],
});
Run each project separately in the corresponding pinned environment, for example:
npx playwright test --project=chromium-old
npx playwright test --project=chromium-new
The project names alone do not pin browser versions. Ensure each job actually launches the intended binary; Playwright manages browser binaries alongside its releases, and its browser documentation describes browser installation and headless behavior: Playwright browsers.
3. Repeat captures and compare the right images
One screenshot can differ because the page was not stable. Dynamic data, animations, delayed fonts or images, and timing can all produce a noisy comparison. Playwright’s screenshot assertion waits for two consecutive screenshots to be identical before comparing with the expected image: PageAssertions. Repeat the old and new runs and check whether the difference recurs.
- Compare the old-version screenshot to the known-good baseline.
- Compare the updated-version screenshot to that same baseline.
- Inspect the changed regions. Determine whether they are stable, localized, and plausibly related to rendering, rather than moving content or missing resources.
- Where feasible, switch back to the old browser and capture again. If the old rendering returns while the updated rendering continues to differ, the version-related explanation is stronger.
Keep the comparison axis explicit: same browser family, old versus updated version, matched environment, and repeated stable captures. If you also compare headed and headless mode, a different operating system, or another browser family, treat those as separate experiments.
4. Interpret the result carefully
| Observed result | What it supports | What to check next |
|---|---|---|
| Only the updated version consistently differs; reverting restores the baseline. | A browser-version-related visual regression is plausible. | Reduce the page to a minimal reproduction and review the browser’s release notes before claiming an engine defect. |
| Both versions differ from the baseline. | The browser update alone does not explain the result. | Check the OS image, test runner, page state, fonts, timing, and capture settings. |
| Results vary between repeated captures of one version. | The page or capture is unstable, so the version comparison is inconclusive. | Stabilize data and loading, disable animation where appropriate, and repeat. |
| Chrome differs from Firefox, but old and new Chrome match. | There may be a browser-family difference. | Test two versions of the same browser to investigate an update. |
Report a result as a “browser-version-related visual regression” when it reproduces under controlled conditions. That wording reflects the evidence without claiming that the browser alone, rather than an interaction with the site or environment, is responsible.
5. Common confounders
- Operating system or container changed: Match the OS image and version across runs. System fonts and rendering behavior can affect pixels.
- Browser and Playwright both changed: Record both versions. Playwright’s browser binaries track its releases, so a combined upgrade does not isolate the browser update.
- Headless mode changed: Keep the mode fixed. Chrome and Edge’s newer headless implementation differs from Playwright’s default Chromium headless shell.
- Page content changed: Seed test data or use a stable fixture. A timestamp, rotating banner, personalized response, or changing API value can create a diff.
- Capture happened too early: Wait for the relevant content, fonts, and images to load. Use a stable selector or an appropriate readiness condition rather than assuming navigation completion means the visual page is ready.
- Device scale factor or viewport changed: Match both exactly. A different scale can change text rasterization and layout breakpoints.
- Different browser family used as the control: Use the same family at two versions for the update hypothesis. Cross-browser comparisons answer a different question.
6. Troubleshooting
| Problem | Likely cause | Fix |
|---|---|---|
| The old and new runs launch the same browser build. | The project configuration selects a browser name but does not pin the executable version. | Use separate pinned environments or browser binaries and log the launched browser’s version in each run. |
| A screenshot diff appears on every run, including the old browser. | The baseline environment or page inputs changed. | Restore the baseline OS image, fonts, fixture data, viewport, scale factor, and capture settings before testing the version hypothesis. |
| The diff changes from run to run. | Animation, live content, delayed resources, or race conditions. | Freeze data, disable animations where suitable, wait for stable content, and repeat captures. |
| Headed and headless screenshots disagree. | They use different rendering paths or settings. | Keep capture mode constant for the primary comparison; test the other mode as a separate axis. |
| Updating Playwright fixes or introduces the difference. | The framework upgrade may have changed the managed browser binary or capture behavior. | Compare browser versions with the framework held fixed where possible. Otherwise, report both changed inputs and isolate them in additional runs. |
| The difference appears only in CI. | CI may use a different OS image, fonts, hardware, power profile, or browser mode. | Reproduce in the same CI image and record its configuration; do not compare a local run against CI as if browser version were the only variable. |
7. Performance, reliability, and cost
Each browser-version comparison requires running the page and taking screenshots in both environments. Repeated runs add confidence but consume more CI time and browser resources. Start with a small number of matched captures; increase repetitions when output is unstable or the difference is subtle. Keep the same test route and avoid broad, unrelated suites until the cause is narrowed down.
Reliability depends on reproducibility: preserve the baseline, pin the environment, record exact versions, and retain the generated images and diff. A passing result in one machine is weaker evidence than repeated results in the same controlled environment. Testing another environment can establish scope, but introduces another variable.
The principal cost is the compute and CI time for additional browser runs and maintaining the old browser environment. There is no topic-specific published statistic for how often browser updates cause screenshot changes, so avoid using unrelated visual-testing statistics to estimate the likelihood.
8. Or skip the browser setup
If you need a clean capture of a page without installing and maintaining a browser environment, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF; the example below captures a page as WebP. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
Frequently asked questions
Does a changed screenshot prove a browser bug?
No. A controlled old-versus-new comparison can support a version-related cause, but the change may depend on the site, fonts, hardware, or test setup. A minimal reproduction and release-note review help narrow it further.
Should I update the expected screenshot after a browser upgrade?
First determine whether the new rendering is intended and stable. Preserve the old baseline during investigation; update it only after reviewing and accepting the visual change.
Can I use another browser as the old-version control?
No, not for testing an update. Use two versions of the same browser. Another browser is useful for a separate cross-browser comparison.
Why record the automation framework version separately?
The framework can determine which browser binaries are installed or launched. Recording both versions helps identify which input changed.


