ScreenshotNeo

BlogHow-to

How to Retry Failed Visual Regression Tests

Configure Playwright retries for visual regression tests, preserve failure evidence, and keep flaky results visible instead of treating them as fixed.

By the ScreenshotNeo team4 October 20267 min read

Playwright Test does not retry failed tests by default. Add retries to playwright.config.ts or pass --retries=N to the test command. A test that fails first and passes on retry is reported as flaky, not healthy; a test that fails on every attempt remains failed. Retries are useful for collecting evidence and reducing transient noise, but they do not fix unstable screenshots or approve a visual change.

1. Configure retries in Playwright Test

Install and use Playwright Test in your project, then configure a retry count. This TypeScript configuration makes two additional attempts after the initial run:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: 2,
});

Run the suite with npx playwright test. To set retries just for one invocation, use:

npx playwright test --retries=2

The number is additional attempts, not total executions: retries: 2 allows up to three attempts. A retry setting can also be scoped to a project or overridden through project configuration when different browsers or environments need different policies. Keep the policy intentional; a high retry count can lengthen feedback and conceal repeated instability in routine reports.

Check the Playwright version installed in the project before copying newer options from documentation. Defaults and available configuration features can change over time. See the official Playwright retries documentation.

2. Make a visual regression test reproducible

Retries help only when the repeated test is meaningful. Screenshot comparisons are sensitive to the environment as well as to the page. Keep the operating system and browser versions consistent between baseline creation and comparison. Also hold constant the viewport, device scale factor, locale, timezone, fonts, test data, and application state wherever those affect rendering.

A minimal visual assertion uses Playwright’s screenshot matcher:

import { test, expect } from '@playwright/test';

test('home page matches its visual baseline', async ({ page }) => {
  await page.goto('http://127.0.0.1:3000');
  await expect(page).toHaveScreenshot('home.png');
});

Run it with npx playwright test. On a first run, the project may need an approved baseline generated using its normal snapshot workflow. Review baseline changes deliberately: an expected design change belongs in baseline review, while a transient failure belongs in test diagnosis. Do not use retries to silently accept a changed baseline.

For stable captures, wait for the relevant application state before asserting, use deterministic test data, and avoid animations or time-dependent content where possible. If the page includes content that genuinely varies, make the test isolate or mask that content according to the project’s visual testing approach. Keep any such exclusions narrow so meaningful regressions remain visible.

3. Choose retry scheduling and CI policy

Playwright replaces the worker process after a test failure. When retries are enabled, the failed test runs again in a replacement worker, which helps avoid carrying corrupted worker state into the retry. Retry scheduling affects interference and total runtime:

Choice Behavior Trade-off
Immediate Retry when a worker is available; retries can interleave with the remaining tests. Often gives earlier feedback, but concurrent activity may contribute to interference.
Isolated Run retries at the end, one by one in a single worker. Reduces interference between failed and healthy tests, while increasing elapsed run time.

The documented retryStrategy option supports immediate and isolated behavior in versions that provide it. Verify availability and exact syntax for your installed version before adding it; do not assume this option exists in every release.

Playwright recommends one worker in CI when stability and reproducibility are priorities. A suite can use parallel workers or sharding when the CI infrastructure supports it and the resulting environment remains reproducible. More workers can shorten wall-clock time, but resource contention can itself introduce timing or rendering instability.

Keep retry outcomes visible. Playwright can fail a run when any test is marked flaky, using the failOnFlakyTests configuration option or its CLI counterpart where supported. This is useful when retries should collect evidence but a flaky result still requires investigation. For traces, Playwright documents enabling trace collection on the first retry, which preserves diagnostic detail for a failing attempt without collecting a trace for every successful run.

import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: 2,
  use: {
    trace: 'on-first-retry',
  },
});

Consult the versioned documentation for the precise failOnFlakyTests configuration and CLI spelling supported by your installed Playwright release. See Retries and Playwright CI guidance.

4. Read the result correctly

  • Passed on first attempt: the test passed without a retry.
  • Flaky: the initial attempt failed, then a retry passed. Preserve the failure details and investigate the source of instability.
  • Failed: the initial attempt and all configured retries failed.

A retry pass does not prove the test is stable. Track flaky outcomes over time, inspect the first failed attempt, and address the underlying cause. If a visual baseline changed intentionally, route it through the team’s normal review or visual-testing approval process. Runner retries do not approve snapshots in a visual-testing service.

5. Troubleshoot common retry problems

Symptom Likely cause What to do
No retry occurs Retries are disabled, the CLI flag was omitted, or a project setting overrides the expected configuration. Check the effective project configuration and command. Confirm the installed package is Playwright Test and the retry option is supported in that version.
The run takes much longer Each retry repeats navigation, setup, and screenshot comparison; isolated scheduling can defer retries until the end. Use a small retry count based on the team’s tolerance for delayed feedback. Fix repeated slow setup and investigate whether isolated scheduling is necessary.
Tests pass only on retry Race conditions, shared state, slow rendering, worker contamination, external dependencies, or resource contention. Inspect the first-attempt trace and logs. Make setup deterministic, wait for the correct page state, isolate shared resources, and reproduce with the CI browser and OS.
Screenshots differ across machines Different OS or browser versions, fonts, viewport, device scale factor, locale, timezone, or dynamic page content. Align capture environments and inputs with the baseline. Stabilize or narrowly exclude genuinely variable content.
Retries keep failing on the same difference The UI may have changed, the baseline may be stale, or the assertion may target the wrong state. Review the image diff and application change. Update a baseline only through the normal approval process when the visual change is intended.
Flaky tests do not fail CI The project reports flaky status but has no fail-on-flaky policy. Enable the documented fail-on-flaky configuration or CLI option for the installed version if flaky outcomes must block the run.
Trace is missing for the failed attempt Trace collection may not be configured for retries, or artifacts may not be retained by CI. Configure trace collection on the first retry and ensure CI uploads the resulting report and artifacts.

6. Performance, reliability, and cost

Retries trade faster recovery from transient failures for longer runs and slower feedback. With two retries, a repeatedly failing test can execute three times. The actual cost depends on navigation, application setup, capture work, worker scheduling, and CI billing; no retry count is universally correct. Start with a low count, inspect the flaky results, and adjust based on evidence.

For reliability, retain the first failure’s report, trace, and screenshot diff. Keep the browser and OS versions pinned or otherwise controlled. Consider one CI worker when reproducibility is more important than throughput; add parallelism or sharding only when infrastructure and test isolation support it. Fail-on-flaky policy makes recovery visible to CI consumers instead of allowing a retry pass to erase the signal.

Or skip the browser setup

If your goal is to capture a page for a visual workflow without maintaining browser automation, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. For a screenshot check, capture the same target URL and compare the returned image in your existing workflow; this does not replace Playwright Test’s retry or visual-baseline approval policy.

See the ScreenshotNeo API documentation. Example cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
  • Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does a retry passing mean the visual test is fixed?

No. Playwright reports that outcome as flaky. Investigate the initial failure and address its cause.

Do Playwright retries approve a changed visual baseline?

No. Retries rerun a test. Baseline review or approval is a separate visual-testing workflow.

Should every visual test use the same retry count?

Not necessarily. Choose a small, explicit policy based on feedback time and observed instability, then keep flaky results actionable.

Can retries make a nondeterministic test reliable?

They can reveal intermittent behavior and help classify outcomes, but only a deterministic test and controlled capture environment establish repeatability.