Playwright Screenshot Tests Failing on Windows: Common Fixes
Fix Playwright screenshot mismatches on Windows by aligning environments, reinstalling browser binaries, stabilizing page state, and inspecting the diff and trace.
Start by matching the environment that generated the baseline to the environment running the comparison. Playwright screenshots can differ across operating systems, browser versions, settings, hardware, power source, and headless mode. Then confirm Playwright and its browser binaries are aligned, make the page deterministic before capture, and inspect the expected, actual, and diff images plus the trace before updating a snapshot.
This guide covers Windows local runs and Windows CI. It assumes Playwright Test and TypeScript; the same diagnosis applies to JavaScript tests.
1. Identify what kind of failure you have
First separate a browser launch or test setup problem from a visual comparison failure. A launch error means the browser did not start correctly. A screenshot mismatch means the test ran and Playwright compared an actual image with its baseline. Fixing the wrong layer wastes time.
| What you see | Start here |
|---|---|
| Browser executable missing or launch fails | Check the Playwright package version and reinstall its browser binaries. |
| Test passes on one OS but fails on Windows | Generate and compare baselines in the same OS and browser environment, or maintain separately reviewed baselines per platform. |
| Intermittent mismatch on the same machine | Stabilize asynchronous page state and inspect the trace for timing, network, or volatile-content differences. |
| Large, consistent visual change | Inspect the actual and diff images and verify the application change before updating the baseline. |
| Only a specific region differs | Determine whether that content is intended to vary. If it is outside the test’s purpose, mask or style that region narrowly. |
Playwright’s visual comparison documentation cautions: “Browser rendering can vary based on the host OS, version, settings, hardware, power source (battery vs. power adapter), headless mode, and other factors.” See the [Playwright Visual comparisons documentation](https://playwright.dev/docs/test-snapshots).
2. Align the baseline, browser, and Windows environment
Use the same rendering environment
Create baselines in the same environment where the test will compare them. A Windows-generated baseline is the appropriate reference for a Windows comparison job. If a project intentionally runs visual tests on multiple platforms, keep platform baselines separate and review each one.
When local Windows and CI Windows still differ, compare the details that can affect rendering: Playwright version, browser version, headed versus headless mode, viewport and device scale factor, machine settings, and whether the machine is on battery or external power. A baseline created on a different OS or browser build may encode those differences rather than an application regression.
Reinstall the browser binaries after a Playwright upgrade
Playwright packages expect matching browser binaries. After changing the Playwright package version, install the browsers again from the project directory:
npx playwright --version
npx playwright install
For a specific browser, install only that browser, for example:
npx playwright install chromium
On Windows, the default browser cache is %USERPROFILE%\AppData\Local\ms-playwright. If a browser launch fails, check that the current Windows account has access to its cache and that the installed browser matches the project’s Playwright version. See the [Playwright browsers guide](https://playwright.dev/docs/browsers).
Check the test project and capture settings
Set the browser project, viewport, and screenshot scale explicitly when they affect the intended image. Keep those settings consistent between baseline generation and CI. Here is a small Playwright Test configuration with an explicit Chromium project and viewport:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
use: {
browserName: 'chromium',
viewport: { width: 1280, height: 720 },
deviceScaleFactor: 1,
headless: true,
},
});
Do not copy settings blindly: use the browser and viewport your product needs, then keep them fixed for the comparison job.
3. Make the page deterministic before capture
Wait for the application state you intend to test, not an arbitrary amount of time. Playwright’s web-first assertions retry while waiting for their condition. Also, toHaveScreenshot() waits for two consecutive identical captures before comparing. That helps with some transient changes, but it cannot make an unstable application state correct.
This runnable TypeScript example waits for a meaningful UI state, then compares a screenshot. Replace the URL and selector with those from your application:
import { test, expect } from '@playwright/test';
test('dashboard visual baseline', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 720 });
await page.goto('http://127.0.0.1:3000/dashboard');
// Wait for the UI state that matters to this test.
await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
await expect(page.getByTestId('dashboard-data')).toBeVisible();
await expect(page).toHaveScreenshot('dashboard.png', {
fullPage: true,
animations: 'disabled',
});
});
Run it with Playwright Test:
npx playwright test tests/dashboard.spec.ts
Use these controls selectively:
- Wait for the actual state: use a retrying assertion for a heading, loaded record, or other condition that shows the page is ready. Avoid fixed sleeps as a substitute for understanding readiness.
- Animations: disable animations for the screenshot when motion is not under test. This can prevent a capture from landing at a different animation frame.
- Volatile content: use a narrow mask or screenshot stylesheet for timestamps, rotating promotions, or other content intentionally outside the test. Avoid broad masks that could hide a regression.
- Viewport and scale: set consistent dimensions and device scale when they influence layout or pixel output.
- Network and data: make test data repeatable where possible. If a response or third-party resource changes between runs, inspect network activity and control the relevant test input.
See the [PageAssertions API](https://playwright.dev/docs/api/class-pageassertions) for screenshot assertion options and [Writing tests](https://playwright.dev/docs/writing-tests) for retrying, web-first assertions.
4. Inspect the failure before changing a baseline
- Open the HTML report and locate the failed screenshot assertion.
- Compare the expected image, actual image, and diff. Identify whether the change is global, layout-related, font-related, or limited to a changing region.
- Open the test trace. Review the action timeline, DOM snapshots, logs, network requests, and screenshot state near the assertion.
- Check whether the application changed intentionally and whether the actual image was captured in the canonical environment.
- Only then update the baseline, using the environment that owns that baseline.
For example, update snapshots with:
npx playwright test --update-snapshots
Review the resulting image changes in version control before keeping them. A larger tolerance is a comparison policy choice; it does not explain why pixels changed. Start with the diff and evidence, then set tolerances only when small rendering variation is acceptable for that test. Playwright’s [Trace Viewer guide](https://playwright.dev/docs/trace-viewer) describes the available trace evidence.
5. Diagnose Windows CI failures
The [Playwright CI guide](https://playwright.dev/docs/ci) says Windows and macOS agents need no additional configuration beyond installing Playwright and running tests. Browser installation and comparable baselines still matter.
If the browser will not launch, enable browser debug logging and rerun the failing command. In PowerShell:
$env:DEBUG = 'pw:browser'
npx playwright test
In Command Prompt:
set DEBUG=pw:browser
npx playwright test
In a POSIX shell:
DEBUG=pw:browser npx playwright test
Use the resulting launch logs to investigate the browser setup. If tests launch and fail only at image comparison, focus first on environment and baseline alignment, then on deterministic page state. For CI troubleshooting practices, see [Playwright Best Practices](https://playwright.dev/docs/best-practices).
6. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Executable does not exist or browser launch fails after an upgrade | Browser binaries are missing or do not match the installed Playwright package. | Run npx playwright install from the project, then retry. Check the Windows browser cache path if the error persists. |
| Snapshot differs on Windows but passes on Linux | The baseline and comparison use different rendering environments. | Generate and compare the Windows baseline on the intended Windows browser setup, or maintain reviewed baselines per platform. |
| Screenshot assertion fails intermittently | The UI or a resource is still changing, or the capture encounters volatile content. | Wait for a meaningful application state, inspect trace timing and network activity, and isolate only intentionally variable regions. |
| Many unrelated pixels differ | Browser build, OS, headless mode, viewport, scale, or other host settings changed. | Compare the full rendering configuration with the baseline generation environment before touching thresholds. |
| Only animation frames or transitions differ | The capture occurs at different points in an animation. | Disable animations in the screenshot assertion if motion is not what the test covers. |
| CI retry passes after an initial failure | The test is flaky; the retry does not identify the cause. | Inspect the failed attempt’s trace and eliminate the source of nondeterminism. Playwright marks a test that fails initially and passes on retry as flaky. |
| Updating snapshots makes the test pass, but the change is unexplained | The baseline was replaced without validating whether the visual change was intended. | Restore or review the change, inspect expected/actual/diff and trace, and regenerate only in the canonical environment after confirming the change. |
Playwright’s [Retries guide](https://playwright.dev/docs/test-retries) explains how retries classify flaky tests. Treat a retry as diagnostic evidence, not as a fix.
7. Reliability, runtime, and maintenance
Visual tests are most reliable when their rendering environment and inputs are stable. Pin the project’s Playwright dependency through its normal package management, install the corresponding browser binaries in CI, and generate baselines with the same project settings used for comparison. Keep baseline changes reviewable.
Screenshot assertions may capture more than once while waiting for two consecutive identical images, so unstable pages can take longer and still fail. Prefer waiting for the intended state over adding arbitrary delays. Keep screenshots focused on the behavior under test: full-page captures can cover more layout, while a targeted element capture can reduce irrelevant page variation when that is appropriate.
No Windows-specific performance benchmark is established by the documentation used for this guide. The practical cost is test time and maintenance: repeated captures and large visual diffs take review time. Do not broaden masking or tolerance just to reduce failures; first remove avoidable sources of variation.
8. Or skip the browser setup
For a clean website screenshot without configuring Playwright and a browser on Windows, [ScreenshotNeo](https://screenshotneo.com) offers a one-request screenshot API. Its capture flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents.
See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/). Example using cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots. All features are available on every plan. [Create a free ScreenshotNeo account](https://screenshotneo.com/account/sign-up/) to get started.
FAQ
Does Playwright guarantee identical screenshots across Windows and Linux?
No. Rendering can vary with the host operating system and other environment factors. Compare screenshots in the environment that generated their baselines, or maintain separate reviewed baselines.
Does a passing retry mean the screenshot test is fixed?
No. It indicates the test was flaky. Use the failed attempt’s trace and diff to find and remove the instability.
Should I raise the screenshot threshold until CI passes?
Only if the accepted visual variation is a deliberate policy. First determine what changed by inspecting the images and trace.
Can I use Playwright screenshot assertions without Playwright Test?
toHaveScreenshot() is part of Playwright Test’s assertion API. Use Playwright Test for that matcher and its snapshot workflow.


