How to fix Playwright screenshot tests that fail intermittently
Diagnose flaky Playwright screenshot tests by checking rendering environments, stabilizing capture, inspecting diffs, and tuning thresholds only when justified.
To fix intermittent Playwright screenshot failures, find and remove the source of nondeterminism before relaxing the visual comparison. Keep baseline generation and comparison in the same rendering environment, inspect repeated-run diffs, stabilize only the state or regions that are meant to be stable, and adjust pixel thresholds only after reviewing the differences.
Playwright Test’s toHaveScreenshot() assertion already waits for two consecutive captures to match before comparing the result with the stored baseline. It also handles animations by default. A fixed sleep or a larger diff tolerance is therefore not the right first response to every flaky test.
1. Confirm the screenshot assertion and baseline
Use Playwright Test’s screenshot assertion for visual comparisons. The first run can generate a reference image; inspect that image and commit it only after approving it as the intended appearance. See the visual comparisons guide.
import { test, expect } from '@playwright/test';
test('account page matches its approved appearance', async ({ page }) => {
await page.goto('/account');
await expect(page).toHaveScreenshot('account-page.png');
});
When a comparison fails, preserve and inspect the baseline, actual screenshot, and diff image produced by the test run. Confirm that the baseline belongs to the expected project and rendering environment before updating it. Updating a baseline to make a failure disappear can approve an unintended visual change.
2. Align the local and CI rendering environments
A screenshot can differ even when the page code is unchanged. Playwright warns that rendering depends on factors including host operating system, browser version and settings, hardware, power source, and headless mode. Generate and compare baselines with the same OS, browser project, and relevant rendering setup where possible. See Playwright’s snapshot guidance.
- Check which Playwright project and browser produced the baseline and which one CI is running.
- Compare local and CI operating systems and browser versions.
- Check headless settings and other launch or rendering configuration.
- Keep baseline updates on the same setup used for CI comparisons, or intentionally maintain separate project-specific snapshots.
Snapshot naming can incorporate browser, platform, or project identity, so a baseline for one environment may not be the baseline used in another. Treat “passes locally, fails in CI” as a reason to compare environments before changing the assertion.
3. Stabilize capture without hiding real changes
toHaveScreenshot() waits until two consecutive screenshots are identical before it compares the last capture with the expected image. Its default animation handling disables CSS animations, transitions, and Web Animations; finite animations are fast-forwarded and infinite animations are temporarily canceled. These behaviors are documented in the Page assertions API.
If the page still changes between captures, identify what is changing. JavaScript-driven state, timestamps, rotating content, external content, or data that changes between requests may need deterministic test setup. This is application-specific diagnosis: the screenshot API’s built-in stabilization cannot make changing application data identical.
When a changing region is genuinely outside the visual contract, use a narrowly scoped screenshot stylesheet with stylePath to hide or neutralize it. Do not mask broad page areas that could contain meaningful regressions. The option can be supplied to an individual assertion or configured for a project; see Playwright’s snapshot documentation and assertion options.
import { test, expect } from '@playwright/test';
test('account page ignores only its live clock', async ({ page }) => {
await page.goto('/account');
await expect(page).toHaveScreenshot('account-page.png', {
stylePath: './tests/screenshot-styles.css',
});
});
/* tests/screenshot-styles.css */
.live-clock {
visibility: hidden !important;
}
4. Repeat the test and inspect what moves
Use repeated execution to establish whether the failure is intermittent and which region varies. Playwright’s repeatEach setting is intended to help debug flaky tests. Retries can also reveal whether a later attempt passes, but a retry pass is evidence of flakiness, not proof that the screenshot check is reliable. See test configuration.
import { defineConfig } from '@playwright/test';
export default defineConfig({
repeatEach: 5,
retries: 0,
});
Run the affected test repeatedly in the same environment first, then compare local and CI artifacts if the failure only appears remotely. Keep the actual, expected, and diff files from failed runs so the changing area is visible. Once investigation is complete, remove a temporary high repeatEach value if it is slowing the normal suite unnecessarily.
5. Tune visual-diff tolerance only after reviewing the diff
Playwright offers maxDiffPixels, maxDiffPixelRatio, and a per-pixel color threshold. These options determine how much difference is accepted. Set a value based on a reviewed, understood rendering variation, rather than broadly increasing tolerance to silence an unexplained failure. A permissive comparison can also admit a real visual regression. See the snapshot comparison options.
await expect(page).toHaveScreenshot('account-page.png', {
maxDiffPixels: 20,
});
Use one justified tolerance adjustment at a time and review the resulting comparison. Prefer environment consistency or a targeted treatment for an irrelevant dynamic region when those address the demonstrated cause.
6. Keep flaky status visible in CI
Retries are disabled by default. If you enable them, configure CI to fail on tests Playwright classifies as flaky so intermittent visual failures remain visible. failOnFlakyTests is a configuration option documented in the Playwright test configuration.
import { defineConfig } from '@playwright/test';
export default defineConfig({
retries: process.env.CI ? 2 : 0,
failOnFlakyTests: !!process.env.CI,
});
Retries can help surface intermittent behavior and make CI outcomes informative. They do not remove the underlying nondeterminism; keep investigating any test that passes only on a retry.
Common failure symptoms and fixes
| Symptom | Likely cause to investigate | Next step |
|---|---|---|
| Passes locally, fails in CI | Different OS, browser, headless mode, or rendering configuration | Compare environments and generate/compare baselines on a consistent setup. |
| Different pixels in the same region across repeated runs | That region or its underlying application state changes during capture | Inspect repeated actual and diff images; make test state deterministic or narrowly use stylePath if the region is outside the visual contract. |
| Failure disappears on retry | The test remains intermittent | Keep the flaky result visible in CI and investigate the changing state; do not treat the retry as a fix. |
| Too many small differences are reported | Rendering variation or a comparison threshold that is too strict for a known, accepted difference | Review the diff and environment, then tune a specific diff option only if the difference is understood. |
| A large hidden area makes the test pass | The screenshot stylesheet may be hiding meaningful content | Narrow the selector to only content outside the visual contract and check that important UI remains covered. |
| Baseline update removes the failure | The newly approved reference may simply encode an unreviewed change | Review the baseline and its environment before committing it. |
Performance, reliability, and cost considerations
Repeated runs and retries consume additional browser time, so use repetition to locate intermittent behavior rather than applying a large repeat count to every test indefinitely. Consistent environments reduce uncertainty in both failures and approved baselines. Targeted screenshot styles preserve more visual coverage than hiding large areas, while tolerance increases trade detection strictness for acceptance of more pixel differences. The Playwright sources describe these behaviors and configuration options but cannot identify the cause of a particular application’s failures; use the artifacts and environment comparison to diagnose that cause.
Or skip the browser setup
For capturing a page outside a test runner, ScreenshotNeo returns a screenshot or PDF from one API request. Its screenshot API can help when you need a clean page capture rather than a Playwright visual assertion; it does not replace Playwright’s baseline comparison or test diagnostics. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers say which page verdict was returned and whether it was billed.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Does toHaveScreenshot() already wait for the page to settle?
It waits for two consecutive screenshots to match. That stabilizes capture output, but it does not make application data or external content deterministic.
Should I add a fixed timeout before every screenshot?
Not as the default fix. First inspect the changing region and use the assertion’s built-in stabilization. Add application-specific setup only when you have identified state that needs controlling.
Should I update the baseline when CI reports a difference?
Only after reviewing the actual screenshot and diff, confirming the intended appearance, and checking that the baseline corresponds to the intended rendering environment.
Can retries make a flaky screenshot test reliable?
No. Retries can help diagnose and report intermittency. A test that passes on retry still has a source of nondeterminism to investigate.


