ScreenshotNeo

BlogEngineering

How to Detect and Customize Flaky Test Detection

Learn how to detect flaky tests without hiding failures, configure retries in Playwright and pytest, and choose a CI policy that keeps retry-pass results visible.

By the ScreenshotNeo team4 October 202611 min read

A flaky test is one whose result changes across runs in a way that appears non-deterministic. To detect one, preserve the first attempt and every retry as separate outcomes, repeat the test under controlled conditions, and flag tests that fail and then pass. Treat that retry-pass as evidence of inconsistency—not proof of its cause. Then choose explicitly whether flaky results should appear only in reports or fail CI.

Retries can make a pipeline green while concealing a real defect. Keep the original failure visible, capture enough context to investigate it, and give any quarantine or retry policy an owner and review date. The examples below cover Playwright Test, pytest, and Azure Pipelines; their retry and reporting semantics differ.

1. What flaky-test detection should tell you

A useful detection workflow answers three separate questions:

  1. What happened first? Record the initial result, error, logs, and diagnostics.
  2. Did a repeat change the result? A fail-then-pass sequence is evidence of flakiness in systems that classify retries this way. A test that fails every attempt is still a failure.
  3. What should CI do? Decide whether a flaky classification is informational, a warning, or a build failure. Detection and enforcement are separate choices.

A final green status by itself is not enough. If the runner or report discards the initial failure, the signal that a test is unstable disappears. Keep retry outcomes in the report and make the gate policy visible to developers.

2. A practical detection workflow

  1. Keep the first result. Store each attempt’s status and failure context rather than reporting only the final status.
  2. Repeat a narrow scope. Repeat the affected test or group to gather evidence. Prefer a runner’s repeat-for-debugging option when available; repeated passes help demonstrate variability, but they do not establish a root cause.
  3. Compare attempts. Check test order, shared data, parallel workers, environment, resource pressure, timing, and external dependencies. Compare a suite run with an isolated run.
  4. Save reconstruction data. Retain logs, traces, screenshots or video for UI failures, environment details, and the exact test command.
  5. Choose a gate deliberately. Keep flakes visible. Decide whether they block CI, and apply the decision consistently.
  6. Assign remediation. Track the suspected cause, owner, and review date. Remove temporary retries or quarantine after the underlying issue is fixed.

Do not select a universal retry count. A small, explicit retry budget limits added runtime, but the right value depends on suite duration and the impact of a missed failure. Retries are diagnostic evidence and temporary containment, not a substitute for fixing nondeterminism.

3. Playwright Test: retry classification, repeated runs, and CI gates

Playwright Test has retries disabled by default. When retries are enabled, a test that fails initially and passes on retry is reported as flaky; a test that continues to fail remains failed. The repeatEach option runs each test multiple times and is useful for debugging. See the official retry guide, configuration reference, and CLI reference.

Configure retries and fail CI on flaky results

In playwright.config.ts, set a small retry count for CI and opt into failing the run when a test is classified as flaky:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: process.env.CI ? 1 : 0,
  failOnFlakyTests: Boolean(process.env.CI),
  reporter: 'list',
  use: {
    screenshot: 'only-on-failure',
    trace: 'retain-on-failure',
    video: 'retain-on-failure',
  },
});

This example uses one retry in CI as an explicit policy choice, not a universal recommendation. The failure artifacts help explain the attempt; review the artifact-retention settings and available options for the installed Playwright version.

Alternatively, set policy from the command line:

npx playwright test --retries=1 --fail-on-flaky-tests
npx playwright test --repeat-each=10 tests/checkout.spec.ts

The first command retries failures once and fails the run if any test is marked flaky. The second repeats each selected test ten times as an investigation exercise; it is not a sensible default for every full CI suite. CLI options are documented in the Playwright Test CLI reference.

Scope retry behavior

Global configuration is convenient when the same policy applies to all tests. A group can override retries when a known area needs a temporary diagnostic policy:

import { test, expect } from '@playwright/test';

test.describe('checkout flow', () => {
  test.describe.configure({ retries: 1 });

  test('submits a valid order', async ({ page }) => {
    await page.goto('/checkout');
    await page.getByRole('button', { name: 'Place order' }).click();
    await expect(page.getByRole('status')).toHaveText('Order confirmed');
  });
});

Use the project’s normal test server and fixtures for this snippet to run. Group-level settings are useful for scoping investigation, but a permanent exception can conceal a defect. The configuration reference also documents retry strategy options, including isolated retry behavior; availability depends on the Playwright version. failOnFlakyTests is documented since v1.52 and retryStrategy since v1.62. Check the installed version before using versioned properties.

Use auto-retrying assertions to avoid avoidable flakes

Retries help classify unstable outcomes; they do not repair brittle synchronization. For asynchronous UI state, prefer Playwright’s retrying locator assertions over checking immediately with a one-shot assertion:

await expect(page.getByTestId('status')).toHaveText('Submitted');

Playwright’s assertion documentation explains retrying assertions and polling. Wait for a meaningful application state, with a bounded timeout, rather than inserting a fixed sleep.

4. pytest: repeat failures and preserve outcomes

pytest’s core documentation describes causes and investigation techniques for flaky tests, while rerunning, random-order, and replay behavior is provided by plugins rather than one universal built-in retry setting. Choose a plugin compatible with the project and make sure its report retains both the initial failure and subsequent attempt outcomes. See pytest’s flaky-test guidance.

Run a focused repeat investigation

First run the suspect test alone, then repeat it and vary order or parallelism if the project already uses those modes. For example, a shell loop gives a simple repeat experiment without adding a retry plugin:

for run in 1 2 3 4 5; do
  echo "Run $run"
  pytest -q tests/test_checkout.py::test_submit_order || true
done

This shell loop is for diagnosis: || true allows later iterations to run after a failure, so the loop’s final exit status does not represent the test result. Read and retain each run’s output. For machine-readable tracking, use your runner or a compatible plugin to emit per-attempt results rather than relying on the loop’s exit code.

To expose order dependence, run the relevant suite with a random-order plugin installed and configured according to that plugin’s instructions. To investigate parallel-only failures, compare serial execution with the project’s parallel mode. pytest notes that parallel execution can reveal dependence on test ordering, such as state not cleaned up by another test.

Use xfail cautiously

xfail communicates an expected failure; it is not a flaky-test detector. A non-strict xfail can let a failing test stop breaking the build, but pytest warns that using it permanently as manual quarantine is dangerous. If used temporarily, give the marker a specific reason and an owner or issue reference:

import pytest

@pytest.mark.xfail(reason="Temporary quarantine: issue 1234")
def test_checkout_recovers_after_network_reset():
    ...

Consider strict xfail when an unexpected pass should draw attention. Current pytest configuration naming has changed across versions, so consult the documentation for the installed version before setting a project-wide strict option. See pytest’s skip and xfail guide.

5. Azure Pipelines: report flakes separately from build policy

Azure Pipelines documents flaky-test management with automatic detection through reruns or custom detection, branch-level flaky data, reporting choices, and manual marking or unmarking after analysis. Its test summary settings can suppress flaky failures from failing a pipeline, but the report’s behavior depends on the test publishing path. Microsoft notes that the summary behavior applies to the Visual Studio Test task and Publish Test Results task; other scenarios may need a custom script. See Microsoft Learn: Manage flaky tests.

Use Azure’s flaky tag to support investigation and tracking, not as a permanent substitute for fixing the test. Decide whether flakes should affect the build’s pass percentage or be visible for troubleshooting only. Record who marked a test flaky and revisit the status after a fix; marking or unmarking applies to future executions, not retroactively to the current pipeline.

6. Root-cause triage

Signal What to inspect Useful response
Fails only after another test Shared files, database rows, globals, browser context, cleanup, and order Make setup and teardown explicit; isolate data; rerun independently and in varied order.
Fails mainly in parallel Shared mutable resources, worker collisions, resource limits Isolate per-worker state or serialize access where required; check resource pressure.
UI assertion fails intermittently Whether the test observes a transient state or assumes a fixed event order Synchronize on a meaningful application state and use retrying assertions; retain screenshots, video, or traces.
Timeout or slow response Application logs, request and response timing, service dependencies, machine load Find the slow or unresponsive dependency; set appropriate bounded timeouts rather than adding arbitrary delay.
Failure follows environment changes Clock, locale, timezone, network, OS, hardware, test data, and dependency versions Control environment inputs where possible and include them in failure metadata.
Failure appears only in a suite Order dependencies, leaked state, setup/cleanup, shared external resources Run the test alone and under varied ordering; remove hidden coupling.

Google’s testing guidance recommends synchronizing on specific application states and warns against arbitrary delays: a delay can become unreliable again and slow the suite. pytest likewise identifies uncontrolled state and test-order dependence as common sources. See Google Testing Blog: Test Flakiness, Part II and pytest’s flaky-test documentation.

7. Customize detection without hiding defects

Make four choices explicit in configuration and team documentation:

Choice Questions to answer Practical starting point
Detection signal Does the runner flag fail-then-pass, or are tests repeated intentionally to investigate? Use retry classification for routine visibility; repeat narrowly during diagnosis.
Scope Does the policy apply globally, to a group, or to a file? Keep the default broad and scope exceptions to the affected area.
Gate policy Does a flaky result fail the job, warn, or appear only in reports? Choose intentionally; do not let a green final status erase the flaky label.
Retry isolation and runtime Do retries run immediately or in isolation after the suite? How much runtime is acceptable? Check runner-version support and consider whether retry isolation reduces interference enough to justify longer runs.

Track first-attempt failures, retry outcomes, final status, test identifier, commit, worker, environment, duration, and diagnostic artifact references. These fields let a team distinguish a genuine fix from a policy change that merely stopped surfacing the signal.

8. Performance, reliability, and cost

  • Runtime: retries add work primarily when tests fail; repeat-each multiplies work across the selected tests. Keep broad repetition out of normal CI unless its cost and purpose are understood.
  • Signal quality: a retry-pass establishes inconsistent behavior under the observed conditions, not its probability of recurrence or its cause. Reproducing conditions and comparing artifacts is more useful than accumulating retries without context.
  • Reliability: a retry may pass because a transient dependency recovered, because ordering changed, or because the product defect itself is intermittent. Preserve the first error and keep meaningful assertions.
  • Resource use: parallelism can shorten elapsed time while increasing contention for shared resources. A failure that appears only under parallel load needs isolation or resource investigation, not just more retries.
  • Engineering cost: temporary quarantine can protect a release workflow, but invisible or permanent quarantine makes test results less trustworthy. Maintain a visible owner and follow-up.

9. Troubleshooting

Problem Likely cause Fix
CI is green despite a first-attempt failure Retries are enabled and flaky results do not fail the job, or the report shows only the final outcome. Enable a flaky-test gate if desired and retain attempt-level results in the reporter.
A test is labeled failed instead of flaky It failed every allowed attempt, or the runner did not execute a retry. Inspect retry count and attempt logs. Keep a consistently failing test classified as a failure.
Playwright rejects a configuration property The installed version does not support that option. Check the installed Playwright version and its matching configuration reference; the documented availability of newer properties is version-specific.
pytest loop exits successfully after failures The diagnostic example uses || true to continue to later iterations. Read each captured result or use a reporting plugin/script that preserves per-run status and returns a meaningful aggregate exit code.
pytest xfail stops surfacing a known issue Non-strict xfail treats failures as expected and can become permanent quarantine. Add a reason and follow-up owner, evaluate strict behavior, and remove the marker after remediation.
Azure summary does not reflect flaky settings The reporting path may not be one of the supported test publishing tasks. Check whether results use Visual Studio Test or Publish Test Results; Microsoft notes other scenarios may need custom suppression logic.
A fixed sleep seems to solve the failure temporarily The test is racing an asynchronous application state. Wait for a specific state or event with a bounded timeout and investigate the timing dependency.

10. Or skip the browser setup

If your investigation needs screenshots of the application under test, ScreenshotNeo is a website screenshot API and MCP server. It can capture a page with one GET request; for automated test evidence, pass the test environment’s URL and use the response image as an artifact. See the ScreenshotNeo API documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Replace the sample URL with the page you need to inspect. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents screenshot, page-info, and PDF capture tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

11. FAQ

Does passing on retry prove the test is flaky?

It is evidence of inconsistent outcomes under those run conditions. It does not identify the cause or prove that the test itself, rather than the application or environment, is responsible.

Should every flaky test block the build?

That is a team policy decision. Keep the flaky result visible either way, then choose a gate based on the test’s impact and the cost of delaying a change.

Is a test that fails on every retry flaky?

It is consistently failing in that run and should remain a failure. Investigate it as a defect or test/environment problem rather than classifying it as a retry-pass.

How long should a test stay quarantined?

There is no universal duration. Set an owner and review date when quarantine begins, and remove it when the cause is fixed or the test is deliberately replaced.

Primary references