How to Reduce Flaky Playwright Screenshot Tests
Make Playwright screenshot tests more reproducible by controlling the browser environment and page state, then diagnosing diffs before adjusting thresholds.
To reduce flaky Playwright screenshot tests, make the baseline and test run use the same rendering environment, wait for the page’s expected state, and isolate genuinely changing content. Use Playwright’s screenshot assertions, inspect the actual and diff images when a test fails, and change comparison thresholds only after you understand the difference. Retries can identify intermittent failures, but a retry pass does not fix their cause.
1. Keep the rendering environment consistent
A screenshot baseline is tied to the conditions under which it was created. Playwright documents that rendering can vary with the host operating system, browser version and settings, hardware, power source, and headless mode. Keep those factors consistent between baseline creation and comparison wherever possible.
For CI, a practical approach is to generate and compare snapshots in the same pinned CI image. Avoid creating baselines on one operating system and expecting pixel-identical results on another. If your suite intentionally tests multiple rendering configurations, keep separate baselines for the relevant Playwright project, browser, platform, viewport, and device scale factor. Playwright snapshot names include browser and platform information, and different browsers or platforms may need distinct references. See the Playwright visual comparisons guide.
| Record or control | Why it matters |
|---|---|
| Playwright version | Rendering and screenshot behavior can change across releases. |
| Browser project and version | Chromium, Firefox, WebKit, and branded browsers can render differently. |
| Operating system and CI image | Fonts, graphics libraries, and platform rendering affect pixels. |
| Viewport and device scale factor | They change layout, wrapping, and rasterization. |
| Headless or headed mode | Capture conditions can affect the result. |
Keep the project configuration explicit so you can reproduce a failure. For example, set the browser project and viewport in Playwright configuration rather than relying on different local and CI defaults:
// playwright.config.ts
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
use: {
browserName: 'chromium',
headless: true,
viewport: { width: 1280, height: 720 },
deviceScaleFactor: 1,
},
});
This example makes the browser project, viewport, and scale factor visible in the configuration. Pin your Playwright dependency and use the corresponding browser installation in the environment that creates and checks the snapshots. If you add Firefox or WebKit projects, review their snapshots separately.
2. Prefer Playwright’s screenshot assertion
Use await expect(page).toHaveScreenshot() for a page comparison, or use a locator screenshot assertion when only one component matters. The assertion waits for two consecutive screenshots to match before comparing the last capture to the baseline. This helps avoid capturing a page while it is still changing.
import { test, expect } from '@playwright/test';
test('checkout summary matches its reference', async ({ page }) => {
await page.goto('/checkout');
await expect(page.getByRole('heading', { name: 'Order summary' })).toBeVisible();
await expect(page.getByTestId('order-summary')).toHaveScreenshot('order-summary.png');
});
The locator assertion narrows the comparison to the component under test. This can make failures easier to interpret and avoids unrelated page areas becoming part of that particular snapshot. Use a page screenshot when the overall page composition is what you need to protect.
Playwright’s screenshot assertion options include controls for animations, caret visibility, masks, stylesheets, and pixel comparison. Check the current API reference for the options and defaults for your installed version.
3. Wait for the application state, not a guessed delay
A fixed sleep such as waitForTimeout(2000) guesses how long the page needs. It can be unnecessarily slow on a fast run and too short on a slow one. Instead, assert the visible state that must exist before the screenshot.
await page.goto('/dashboard');
await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
await expect(page.getByTestId('account-balance')).toHaveText('$125.00');
await expect(page).toHaveScreenshot('dashboard.png');
Playwright’s assertions retry while checking their condition, which helps avoid races between the test and the application. Prefer a specific assertion for the content or state relevant to the screenshot. If the page depends on an API response or a client-side transition, wait for the resulting UI condition rather than assuming a request finishing means the interface is ready. See Playwright assertions and auto-waiting and actionability.
4. Control animation, caret, and volatile regions
Screenshot assertions disable animations by default and hide the caret by default. Finite animations are fast-forwarded, while infinite animations are canceled for the capture. These defaults help remove common sources of pixel variation. You can make the behavior explicit when it improves readability:
await expect(page).toHaveScreenshot('profile.png', {
animations: 'disabled',
caret: 'hide',
});
If an animated state is the behavior under test, do not disable it blindly. Arrange a deterministic state or test the animation with an assertion suited to that behavior.
For timestamps, rotating avatars, live counters, or other truly volatile content, mask only the unstable region. A locator screenshot assertion can also reduce the capture area. The page assertion supports masks; check the installed version’s API reference for exact option types and mask styling behavior:
await expect(page).toHaveScreenshot('activity.png', {
mask: [page.getByTestId('last-updated-time')],
});
Another option is stylePath, which applies a stylesheet during capture. Playwright documents that this stylesheet can pierce Shadow DOM and inner frames. Keep the rule narrow so it does not hide meaningful interface changes:
await expect(page).toHaveScreenshot('activity.png', {
stylePath: './tests/screenshot-stability.css',
});
/* tests/screenshot-stability.css */
[data-testid="live-clock"],
[data-testid="rotating-ad"] {
visibility: hidden !important;
}
Masking or hiding a region trades coverage for stability: changes inside that region will not be meaningfully checked. Document why each exception exists and keep the masked area as small as possible.
5. Set comparison thresholds deliberately
Playwright’s documented default YIQ color threshold is 0.2. Screenshot assertions also support limits such as maxDiffPixels and maxDiffPixelRatio. These options define how much difference the comparison tolerates; they do not make a changing page deterministic.
await expect(page).toHaveScreenshot('catalog.png', {
maxDiffPixels: 20,
});
Treat a threshold as an explicit tradeoff. A stricter comparison can catch small visual changes but may expose harmless rendering noise. A more permissive comparison can reduce sensitivity and also conceal real regressions. Review representative actual and diff images before choosing a value. Avoid increasing tolerance simply because a failure is inconvenient.
Use either an absolute pixel limit or a ratio when it best describes the intended allowance. The right value depends on the screenshot and project; the documentation does not establish a universal value that is safe for every page.
6. Create and update baselines carefully
Commit reference snapshots so code review can show what visual change is expected. When a test reports a difference, first decide whether the application change is intended. Update snapshots only after that review:
npx playwright test --update-snapshots
Run the command in the same rendering environment used for normal comparison, then review the changed snapshot files. Updating references to make a failing test pass without inspecting the difference can turn a real regression into the new baseline.
7. Use retries as a diagnostic signal
Playwright can retry failed tests when retries are configured. A test that fails and then passes is categorized as flaky. That classification is useful evidence that the result is intermittent, but it does not identify the cause or establish that the test is fixed. Keep retry information visible in CI and investigate the first failure rather than treating a later pass as proof of stability. See Playwright test retries.
// playwright.config.ts
import { defineConfig } from '@playwright/test';
export default defineConfig({
retries: process.env.CI ? 1 : 0,
});
Retries cost additional CI time when failures occur. Use them to collect evidence while working on reproducibility, not as a substitute for controlling the environment and page state.
8. Diagnose a failure from the diff outward
- Confirm the run conditions. Compare Playwright version, browser project and version, operating system or CI image, headless mode, viewport, and device scale factor with the baseline environment.
- Inspect expected, actual, and diff images. Identify whether the difference is a layout shift, missing content, animation frame, font rendering change, or a small volatile region.
- Check the test’s readiness condition. Confirm the test waited for the specific content or state shown in the image.
- Inspect browser logs and network activity. Look for failed requests, console errors, or late data that explains the visual difference.
- Reproduce in the baseline environment. Rerun the specific test with the same project and settings.
- Fix the source of nondeterminism. Narrowly mask truly variable content, wait on the right UI condition, or align the environment.
- Review any threshold change or snapshot update. Confirm it represents an intentional visual policy or application change.
Playwright UI Mode can show expected, actual, and diff screenshots; its browser tools can help inspect console logs and network requests around the capture. See the UI Mode documentation.
9. Runnable minimal example
Here is a small test file that waits for application state, captures a stable component, and keeps its comparison strict unless a reviewed difference justifies another setting:
// tests/profile.spec.ts
import { test, expect } from '@playwright/test';
test('profile card visual reference', async ({ page }) => {
await page.goto('/profile');
const card = page.getByTestId('profile-card');
await expect(card).toBeVisible();
await expect(card.getByRole('heading')).toHaveText('Account profile');
await expect(card).toHaveScreenshot('profile-card.png', {
animations: 'disabled',
caret: 'hide',
});
});
Run the test with npx playwright test tests/profile.spec.ts. To create or intentionally refresh references, use npx playwright test tests/profile.spec.ts --update-snapshots, then review the resulting files. This is a JavaScript/TypeScript Playwright example; screenshot assertions are provided by @playwright/test.
10. Troubleshooting common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Local passes, CI fails | Different OS, browser version, headless mode, fonts, or rendering dependencies. | Compare run conditions and create and verify baselines in the same CI environment. |
| Failure moves between runs | Uncontrolled page state, late content, animation, or volatile data. | Inspect the diff, assert the relevant UI state, and narrowly disable or mask the changing element. |
| Snapshot is blank or partly loaded | The test captured before the app reached its expected state, or a resource failed. | Assert the required content and inspect browser logs and network requests. |
| Many unrelated pixels differ | Environment drift or a broad layout/font change. | Check OS, browser, viewport, scale factor, fonts, and the full-page diff before adjusting tolerance. |
| Test passes only on retry | An intermittent failure remains; retries only surfaced it. | Use the first failed attempt’s artifacts and logs to locate the unstable condition. |
| Updating snapshots removes the failure | The baseline changed, but the visual change may not have been intended. | Review the snapshot change and application diff before committing the new reference. |
| Mask hides a real bug | The masked region includes meaningful UI. | Narrow the mask or assert the hidden content separately with a functional assertion. |
11. Performance, reliability, and cost considerations
Visual assertions need browser rendering and image comparison, so keep their scope purposeful. A locator screenshot can be more focused than a full-page image when the component is the behavior under test. Retries add runtime on failing attempts; a large number of broad screenshots also means more reference files to review and maintain.
Reliability comes primarily from stable inputs: a repeatable browser environment, known application state, and deliberate handling of volatile pixels. Thresholds and retries are controls with tradeoffs, not guarantees. No universal cost or runtime figure applies across CI providers and pages; measure the suite in your own environment.
12. Or skip the browser setup
If you need a clean reference capture without configuring a browser for that capture, ScreenshotNeo is a website screenshot API and MCP server for developers. It does not replace Playwright’s assertion against committed visual baselines, but it can return a screenshot or PDF from one request. See the ScreenshotNeo site and API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', await res.arrayBuffer());
Replace the example URL with the page you need and provide your API key. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Should I use a full-page screenshot for every test?
No. Use a full-page capture when page composition is what you need to protect. For a component-level check, a locator screenshot keeps the comparison focused.
Does a passing retry mean the test is fixed?
No. A fail-then-pass result is classified as flaky and indicates an intermittent failure. Investigate the original failure and remove its cause.
Can I use a higher diff threshold to stop flakes?
A threshold can tolerate pixel differences, but it cannot make page state or environment deterministic. Review diffs first and accept the risk that a permissive threshold may miss a regression.
When should I update a snapshot?
After confirming that the visual change is intentional and reviewing the new reference in the same environment used for comparison.


