ScreenshotNeo

BlogEngineering

How to Quarantine Flaky Tests in CI

Quarantine confirmed flaky tests without losing failure evidence, ownership, or a clear path to repair and reintroduction.

By the ScreenshotNeo team4 October 20269 min read

Quarantine a confirmed flaky or broken test only as a tracked, temporary measure: preserve the failure evidence, link the test to an issue and an accountable owner, exclude it from the CI blocking path through a framework-supported mechanism, and set a review date and exit plan. Keep investigating. Then fix, remove, replace, or deliberately reintroduce the test after stability checks.

A retry that passes is evidence of inconsistency, not proof that the failure is harmless. A flaky result can come from the test, the environment, or the application. Quarantine does not repair any of those causes.

1. Confirm and characterize the failure

Call a test flaky when it produces inconsistent results under the same intended conditions. Before changing CI behavior, preserve enough information for someone to reproduce and investigate the failure:

  • The failing job or pipeline link and the test’s full identity.
  • The stack trace, assertion output, and failure pattern, including whether a retry passed.
  • The random seed, test order, shard, runtime, dependency versions, and relevant environment details.
  • Any screenshots, logs, traces, videos, or application state that help explain the failure.

GitLab’s issue guidance, for example, calls for the failing pipeline or job, stack trace, failure pattern, and an appropriate failure label. Treat that as a useful example, not a universal issue template. GitLab’s flaky-test guidance

Investigate before muting the test. Look for state leakage between tests, order dependence, shared resources, race conditions, time-sensitive assertions, unstable external dependencies, and application defects. Reproduce with the same seed when possible. A passing rerun can narrow the investigation, but it does not establish which cause is responsible.

2. Decide whether quarantine is justified

Fix the cause directly when it is understood and a prompt fix is practical. Quarantine is appropriate when the test is blocking useful development and a fix cannot be delivered promptly, provided the failure remains visible and someone owns the follow-up.

Before quarantining, classify what is known. The failure may be caused by flaky test code, a product bug, stale expectations, a broken framework or dependency, an unstable environment, or an investigation that is still in progress. A known product regression should remain clearly classified as a product defect; do not describe it as harmless flakiness to make it disappear from the blocking path.

Ask these questions in the issue or review:

  1. What exact test is being excluded, and what evidence shows inconsistent or broken behavior?
  2. What coverage or regression protection is lost while it is quarantined?
  3. Who will investigate, and when will the team review the status?
  4. Will the exit decision be to fix, remove, replace, or reintroduce the test?

3. Choose a quarantine mechanism

Use the test framework and CI system’s supported mechanism. The mechanism should remove the test from the blocking path without erasing it or making it impossible to run deliberately. Keep the reason and issue reference close to the test or in the quarantine record.

For example, GitLab documents an RSpec workflow that attaches quarantine metadata and an issue URL to an example or enclosing context. Its Jest example skips the test with a quarantine comment and provides a way to run quarantined Jest tests. These are GitLab repository conventions, not portable instructions for every RSpec or Jest project. Check your project’s current framework documentation and CI configuration before copying a pattern. GitLab’s quarantine implementation guide

GitLab also documents repository-specific constraints, including restrictions on quarantining certain shared examples through its shown mechanism and prerequisites involving feature-category metadata, a test-failure issue, and merge-request linkage. Do not assume those constraints apply outside that repository.

A useful quarantine record should include:

  • Test identity: file, suite, case, and any relevant platform or shard.
  • Reason and evidence: what failed, how often, under what conditions, and why it is not currently blocking.
  • Issue and owner: one link to the investigation and a named team or person accountable for the next update.
  • Dates: quarantine start, next review, and an explicit expiry or decision date.
  • Coverage plan: what protection remains and how it will be restored or replaced.

4. Set a review window and ownership

Choose the shortest review interval your team can reliably meet. A test blocking a critical release may need an immediate, short-lived exception. A poorly understood failure may need a longer investigation period, but it still needs regular updates and an expiry decision.

GitLab’s handbook distinguishes a fast quarantine from a longer-term quarantine. Its current policy gives the fast path a maximum of three days, expects ownership acknowledgement within 48 hours, and limits long-term quarantine to three months. It also describes a one-week deletion warning and uses more than 100 local runs as one dequarantine stability criterion. These are GitLab policy thresholds, not industry standards or guarantees. Adapt the intervals to your own release risk and team process. GitLab’s quarantine policy

Do not let an issue become a passive parking place. At each review, record what was learned, what action is next, and whether the test still deserves to exist.

5. Investigate and resolve the underlying problem

Use the preserved evidence to work through likely causes:

  • Order dependence or state leakage: run the test alone and in different suite orders; inspect shared fixtures, global state, database records, files, and cleanup.
  • Concurrency and timing: replace arbitrary sleeps with condition-based waits where appropriate; check synchronization, timeouts, and assumptions about scheduling.
  • Environment instability: compare CI and local dependencies, resource limits, clocks, locale, network access, browser versions, and service readiness.
  • External services: identify reliance on live APIs or mutable data; use controlled fixtures or a stable test double where appropriate.
  • Application behavior: determine whether the failure exposes a real regression. Keep a product bug linked and visible if it does.
  • Test quality: verify that the assertion checks the intended behavior and that the setup and teardown leave no shared state behind.

GitLab’s RSpec debugging material describes reproducing from the CI seed, bisecting where useful, inspecting state leakage and order dependence, and rerunning after a fix. Those commands and procedures are specific to its setup; the investigation principles generalize. GitLab’s debugging guidance

When resolved, choose one deliberate outcome:

  • Fix it: repair the test, application, or environment and restore the test to the blocking suite.
  • Remove it: delete it when it is redundant, obsolete, or cannot provide useful coverage.
  • Replace it: move the check to a better level or replace it with more reliable coverage.
  • Reintroduce and monitor it: restore it after the root cause is addressed and your stability criteria are met.

6. Dequarantine and monitor

Before restoring a test to the blocking path, confirm that the root cause has been identified and fixed, or that the test has been replaced or removed with a deliberate coverage decision. Run it repeatedly under relevant conditions, including the previously failing seed, order, shard, or environment when possible.

GitLab’s policy uses more than 100 local passes as one example criterion, followed by a week of monitoring after reintroduction. That run count is not a mathematical guarantee that a test is stable; use a threshold that fits the test’s risk and execution conditions. If the failure returns, capture the new evidence and decide promptly whether to re-quarantine or address a newly understood cause.

7. Common mistakes and troubleshooting

Symptom Likely cause What to do
A test passes on retry, so the issue is closed The retry was treated as proof instead of evidence Keep the failure record and investigate the test, environment, and application. Close only after a resolution or explicit risk decision.
The test still fails the CI job The quarantine marker is unsupported, attached at the wrong scope, or the job uses a different test command Check the framework and runner documentation, confirm the affected job and test identity, and inspect collection or skip output.
The test silently disappears from local runs too The skip rule applies in every environment, or there is no supported opt-in path Provide a deliberate way to run quarantined tests and document it. Verify that the test remains discoverable and maintained.
The same failure returns after dequarantine The root cause was not fully resolved, or the reproduction conditions were not covered Preserve the new seed and environment data, restore the quarantine if necessary, and reopen the investigation with the evidence.
The quarantine has no updates No accountable owner or review date was recorded Assign an owner, set a near review date, and ask for a status and next action at each check-in.
A shared example cannot be quarantined with the chosen annotation The repository’s implementation has a scope or metadata limitation Follow the implementation’s documented constraints. Quarantine at a supported scope or adjust the test structure without losing coverage.
A known product bug is being treated as flaky The issue label or quarantine reason obscures a regression Classify and link the product defect explicitly, keep the risk visible, and set an owner and resolution plan.

8. Reliability, performance, and cost considerations

Quarantining reduces the chance that a known intermittent test blocks unrelated work, but it also reduces the suite’s immediate protection. The practical cost is the risk of a regression going undetected until other coverage or production feedback catches it. Record which behavior is no longer gating and make that loss visible in the issue or release process.

Retries can improve signal collection, but they consume CI time and can hide failure frequency if dashboards show only the final result. Preserve first-attempt failures and report retry outcomes separately where your CI tooling supports it. Avoid increasing retries as a substitute for identifying the cause.

For a large suite, track quarantine count, age, owner, failure reason, and time to resolution. A growing backlog can indicate that the review process is not working. Prefer your existing CI reports and issue workflow; no additional service is required to apply the quarantine process described here.

Or skip the browser setup

For a separate need—capturing web pages as screenshots or PDFs during a workflow—ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns a PNG, JPEG, WebP, or PDF. It does not manage test quarantine or replace CI test reporting.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for parameters and response details. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Does a retry that passes mean a test is healthy?

No. It shows that the result is inconsistent under the conditions observed. The cause may still be in the test, environment, or application.

Should a quarantined test run locally?

That depends on your framework and repository convention. Keep a deliberate way to run it for diagnosis; GitLab says its quarantined tests run locally by default.

How long should a test stay quarantined?

Set a review and exit date based on urgency and risk. GitLab’s three-day and three-month limits are examples of its own policy, not universal standards.

Can I quarantine a test without an issue?

Keep a durable record with an owner and follow-up date. An issue is a straightforward way to preserve that context and track the resolution.

Sources