ScreenshotNeo

BlogHow-to

How to Find and Fix Flaky Cypress Tests Using Code Smells

Find the code smells behind intermittent Cypress failures, replace timing guesses with reliable synchronization, and verify fixes under varied conditions.

By the ScreenshotNeo team4 October 20269 min read

A flaky Cypress test passes sometimes and fails at other times without a meaningful change to the code under test. The fix is to reproduce the failure, identify the source of nondeterminism, and make the test establish its own state and wait for the condition it actually needs. Common code smells include relying on another test’s leftovers, brittle selectors, arbitrary time delays, and branching on a DOM that is still changing.

Retries can reveal that a failure is intermittent, but they do not remove its cause. Cypress has two different retry mechanisms: query and assertion retries that wait for application state, and optional test retries that rerun a failed test. Use the first for ordinary synchronization; use the second as diagnostic evidence while you investigate.

1. Reproduce the failure before changing the test

  1. Preserve the failure context. Keep the assertion, Cypress command log, browser, test data, Cypress version, operating environment, and whether the failure occurred in cypress open or cypress run.
  2. Run the suspect test by itself. If it fails alone, focus first on its setup, asynchronous dependencies, selector, and expected condition.
  3. Run it with its spec and suite. If it fails only after another test, investigate state leakage, order dependence, and shared server-side data.
  4. Repeat it and vary load. Cypress recommends excessive repetition and simulating different network and CPU conditions. Its example uses 100 executions; that is an example, not a universal threshold or a statistically meaningful sample size.
  5. Sort the symptom into a hypothesis. An element timeout may mean the expected state was never reached, the selector is wrong, or an asynchronous dependency is unresolved. A failure only after a preceding test suggests leftover state. A failure under CI load may point to timing or resource assumptions. Treat these as leads to investigate, not certain diagnoses.

Animations, API calls, server or database availability, resource availability, and network issues can all contribute to race-related failures. Record what changes between passing and failing runs before editing; otherwise, a code change can hide the symptom without explaining it.

2. Fix the code smells that create nondeterminism

Smell: one test relies on another test’s state

A test may pass in the full suite because an earlier test logged in, created a record, or navigated to a page. Run it alone, reorder tests, or retry it and that hidden precondition disappears. Cypress enables end-to-end test isolation by default, but browser isolation does not automatically reset server-side records or shared external systems. Each test should arrange the state and data it needs.

// cypress/e2e/orders.cy.js
describe('orders', () => {
  beforeEach(() => {
    // Establish deterministic server-side data for every test.
    cy.request('POST', '/test-support/reset-and-seed', {
      orders: [{ id: 'order-123', status: 'ready' }],
    });
    cy.visit('/orders/order-123');
  });

  it('shows the seeded order', () => {
    cy.get('[data-cy=order-status]').should('have.text', 'Ready');
  });
});

The reset endpoint above is application-specific: implement it only in a test environment and adapt the request to your app’s test-support mechanism. Programmatic login or setup can make tests faster and more isolated, but retain a separate user-flow test for the login experience itself.

Smell: selectors depend on styling or incidental markup

Long CSS paths, presentation classes, and implementation-specific IDs can break after a style or markup refactor even when the user-facing behavior is unchanged. Prefer purposeful testing attributes such as data-cy, made specific enough to identify the intended control.

// Fragile: tied to layout and styling
cy.get('.panel > div:nth-child(2) .green-button').click();

// Stable contract intended for tests
cy.get('[data-cy=save-profile]').click();
cy.get('[data-cy=save-confirmation]').should('be.visible');

Cypress recommends data-* attributes to separate selectors from CSS and JavaScript changes. Use the project’s chosen attribute consistently; avoid adding generic selectors that match multiple controls.

Smell: a fixed delay guesses when work will finish

cy.wait(5000) encodes a timing guess. It can still be too short on a slow run and wastes time on a fast one. Prefer an assertion about the state required by the test: Cypress retries linked queries and assertions until they pass or time out. Commands that are not queries run once.

// Avoid guessing how long rendering takes
cy.wait(5000);
cy.get('[data-cy=results]').should('be.visible');

// Synchronize on a known request, then verify the visible outcome
cy.intercept('GET', '/api/search*').as('search');
cy.get('[data-cy=search]').type('cypress');
cy.wait('@search');
cy.get('[data-cy=results]').should('contain', 'Cypress');

A wait for a specific intercepted request can be a useful synchronization boundary. Follow it with an assertion about the UI if the test’s goal is to verify what the user sees. A request completing does not prove that rendering or a later client-side update has finished.

Smell: branching on a DOM that has not settled

Checking whether a transient element exists and taking different test paths is unsafe while a client application may still render or update asynchronously. If the DOM is not known to be settled, the element’s temporary absence cannot tell you which state the application will ultimately reach.

// Risky when the app may still be rendering:
cy.get('body').then(($body) => {
  if ($body.find('[data-cy=welcome-modal]').length) {
    cy.get('[data-cy=welcome-modal-close]').click();
  }
});

Make the state deterministic instead. For example, set an experiment using a URL parameter, seed the user state, or read a stable source of truth such as a cookie, local storage, explicit test data, or server state before choosing a path. Conditional testing is appropriate only when the state being inspected is known to be settled.

Smell: required cleanup happens only after the test

If a runner refreshes or a run stops before an after or afterEach hook completes, cleanup-only logic can leave stale server data for later tests. Establish required preconditions before each test so it can recover from leftovers. First determine whether the state is browser state that Cypress’s automatic test isolation already clears or server-side state that needs explicit reset.

beforeEach(() => {
  cy.request('POST', '/test-support/reset');
  cy.request('POST', '/test-support/seed', { user: 'test-user' });
});

Keep destructive reset endpoints unavailable outside the test environment, and scope test data so parallel runs do not overwrite each other.

3. Understand Cypress’s two retry mechanisms

Mechanism What repeats Best use What it does not prove
Query and assertion retry-ability Linked queries and assertions retry until success or timeout Wait for the application condition the test requires It cannot make a wrong selector or impossible condition correct
Test retries The failed test runs again when retries are configured Expose and report intermittent failures A later pass does not show that the underlying cause is fixed

Test retries are disabled by default. If enabled, the configured count is the number of additional attempts. The beforeEach and afterEach hooks run again for each attempt, so setup must be safe to repeat. Preserve the failure history and investigate tests that fail once and then pass; do not treat a larger retry count as a repair.

// cypress.config.js
const { defineConfig } = require('cypress');

module.exports = defineConfig({
  retries: {
    runMode: 2,
    openMode: 0,
  },
  e2e: {
    testIsolation: true,
  },
});

Adapt the retry counts to your diagnostic and CI policy. Cypress documentation also describes experimental retry strategies for flake detection, such as retaining a failing result after a later pass or requiring a threshold of passing attempts. These settings are experimental and can change; check the documentation and configuration supported by the Cypress version in use before adopting them.

4. Validate the proposed fix

  1. Run the test alone and confirm the assertion describes the required user-visible state.
  2. Repeat the test under varied network and CPU conditions to expose timing assumptions.
  3. Run it in its normal spec and suite, then run neighboring tests to check for leaked state.
  4. Run in the environment where the failure occurred, such as CI or cypress run, and compare with the local browser run.
  5. Record the Cypress version, browser, operating environment, and run mode with the result.

A useful fix makes the test establish deterministic preconditions, select the intended element reliably, and synchronize on the state the test needs. Repetition is evidence about consistency, not a guarantee against every future environment.

5. Troubleshoot common intermittent failures

Symptom Likely cause to check Practical next step
Element command times out Expected state was not reached, selector no longer matches, or an async dependency is unresolved Inspect the command log and DOM; verify a stable selector and assert the condition that should make the element appear
Passes in suite but fails alone Test relied on earlier navigation, login, or data Make setup explicit in the test or beforeEach; seed required state
Passes alone but fails after another test Shared server-side data or external state leaked across tests Reset or namespace data before each test; check parallel-run collisions
Fails only in CI or under load Timing, resource, network, server, or database assumptions Reproduce with throttled CPU/network and synchronize on a request or application assertion
Fixed wait still occasionally fails The delay is shorter than some runs need, or the awaited work is not the work represented by the delay Remove the arbitrary delay; wait for the specific request and assert the resulting state
Test passes on retry Nondeterminism remains, even though a later attempt passed Keep the failure visible in triage; compare attempts and investigate timing and state
Test changes behavior based on element presence DOM may be transient while the app is rendering Use deterministic test data or a stable source of truth to select the expected path
Failure appears after a browser or runner refresh Cleanup may not have completed Put required reset and setup before the test; distinguish browser state from server state

6. Keep the fix maintainable

Compare candidate fixes by whether they establish deterministic setup, survive styling changes, identify the actual synchronization condition, preserve diagnostic visibility, and remain easy to maintain. A broad timeout increase may hide a slow condition without clarifying it. A data attribute adds a small test-facing contract but usually survives visual redesigns. Programmatic setup reduces dependence on earlier tests, while a dedicated user-flow test should still cover the real login or workflow when that behavior matters.

7. Capture browser evidence when a failure is hard to inspect

When an intermittent failure is difficult to compare across runs, a screenshot can preserve what the page looked like at the failure point. Keep the test’s assertions and logs as the source of diagnosis; a screenshot is supporting evidence, not proof of why the failure occurred.

// Capture the current page at a useful checkpoint
cy.screenshot('search-results-after-response');

For repeatable external-page captures or CI evidence, you can also use a screenshot API. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its captures remove known consent banners, newsletter popups, and chat widgets before the shot; failed loads, bot checks, blank pages, and cache hits are not billed. Its MCP server provides screenshot tools for AI agents. See the ScreenshotNeo website and API documentation.

Or skip the browser setup

One GET request returns a screenshot. Create a free API key, then replace YOUR_API_KEY and the target URL in any example below. The response is an image file; check the response headers for the page verdict and billing status.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free and get your API key.

FAQ

Why does a Cypress test pass locally but fail in CI?

CI can change network, CPU, and resource conditions. Treat the environment difference as a clue: reproduce under varied load, inspect the failed condition, and replace timing assumptions with explicit synchronization.

Should I increase the test timeout?

Only when the expected operation legitimately needs more time and the timeout reflects that requirement. First verify the selector, preconditions, and synchronization point; a larger timeout alone can make a broken assumption slower to report.

Are Cypress retries always bad?

No. They can surface intermittent failures and help teams collect diagnostic evidence. Keep retry results visible and fix the nondeterminism that caused the first failure.

Can a screenshot prove what caused a flaky test?

No. It records visible page state at a point in time. Use it alongside the Cypress command log, assertion, and run context to investigate the cause.

Official Cypress references