ScreenshotNeo

BlogEngineering

How to Write End-to-End Tests Without Slowing Development

Keep end-to-end tests focused on critical journeys, then measure and remove the setup, waits, and CI bottlenecks that slow feedback.

By the ScreenshotNeo team4 October 202611 min read

Short answer: Keep end-to-end (E2E) tests for critical journeys and system behaviors that smaller tests cannot establish, make each test independent, and measure where time is going before changing the suite. The biggest gains usually come from reducing unnecessary UI setup and waits, stabilizing shared state, and using parallel CI only when the tests and machines can support it.

E2E tests exercise a system through user-visible behavior, often across a browser, application server, and dependent services. They provide confidence at important boundaries, but they are slower and more operationally demanding than unit, component, or integration tests. A useful suite is deliberately small enough to run and diagnose quickly.

1. Put each behavior at the smallest useful test level

Use the E2E layer for behavior that depends on the integrated system: a user completing a high-value journey, components working together across a meaningful boundary, or a system property such as resource allocation, concurrency, or API compatibility that smaller tests cannot reliably prove. Test routine logic and component behavior at lower levels when those tests provide sufficient confidence.

For each important use case, consider one E2E test for the successful path and tests for important classes of failure. Keep the total E2E count low enough that failures remain actionable. This is a selection principle, not a universal percentage quota; the right mix depends on the system and the risks it needs to cover.

Test level Use it for Typical feedback characteristics
Unit Business rules, parsing, calculations, and isolated edge cases Usually the quickest and easiest to pinpoint
Component A UI component’s behavior with controlled dependencies Useful for interaction and rendering without full-system setup
Integration or API Contracts and behavior across a service or data boundary More realistic than a unit test with less browser overhead
End-to-end Critical user journeys and system behavior that smaller tests cannot establish Higher setup, runtime, and maintenance cost

Choose the level by asking: “What uncertainty does this test remove?” If a unit or integration test can prove the behavior with the same confidence, moving it to E2E adds run time and maintenance without adding much signal.

2. Establish a baseline before optimizing

Record representative local and CI runs before making broad changes. Identify the slowest individual tests and spec files, then determine whether the time is spent in browser startup, repeated authentication, UI-driven setup, real network calls, application readiness, fixed waits, or a resource-constrained runner.

  1. Capture total suite wall time and per-test or per-spec duration.
  2. Compare a representative local run with CI to identify environment-specific delays.
  3. Inspect the longest contributors first; optimizing already-short tests may not change total time.
  4. Check machine CPU, memory, and contention before increasing workers.
  5. Change one major source of waste at a time and compare the next run with the baseline.

Cypress publishes these vendor reference ranges in its performance guidance: under 3 seconds per individual test using stubs and programmatic setup is “Excellent”; 3–10 seconds for an E2E test against a real server is “Acceptable”; 10–30 seconds merits investigation; over 30 seconds is “Poor.” For spec files, it describes under one minute as excellent for memory and parallelization, and more than five minutes as poor. These are guidance ranges, not independent benchmarks or guarantees for every application and CI environment. Cypress also cautions that specs under 10 seconds may not benefit from splitting because browser launch and video overhead can outweigh the gain. See Cypress: Optimizing test performance.

3. Make tests independent before running them concurrently

A test should set up the state it needs and should not depend on another test’s order or side effects. Give tests their own data and, where appropriate, their own cookies and storage state. Clean up or namespace created records so retries and concurrent workers do not collide.

  • Use unique test identities or data per run and worker when the system allows it.
  • Do not rely on a preceding test to create a user, populate a cart, or establish a session.
  • Reset or isolate mutable state, including database records and browser storage.
  • Make setup repeatable so an individual test can run alone during debugging.
  • Use fixtures or setup helpers to share mechanics while keeping each test’s required state explicit.

Isolation improves reproducibility and prevents failures from cascading. It is also a prerequisite for safe parallel execution. Playwright’s best practices and parallelism guide both emphasize independent tests and state.

4. Reduce setup without weakening what the test proves

Repeated UI-driven setup is a common source of wasted time. If every test signs in through the full login screen, consider authenticating once and reusing a saved session or establishing state programmatically. Keep at least the appropriate E2E coverage for the login journey itself; using programmatic setup in other tests does not prove that the login UI works.

Likewise, replace slow external dependencies with controlled stubs when the test’s purpose is application behavior under a known response. Keep tests that validate the real integration where the network or third-party contract is itself the behavior under test. The boundary should be explicit: a stubbed test checks your handling of a response; a real dependency test checks the integration.

For browser checks, prefer assertions tied to user-visible outcomes and semantic locators. A test for a purchase journey should establish that the user reached the expected confirmation state, rather than asserting an internal CSS class that can change without changing behavior. Playwright recommends user-facing locators and web-first assertions in its best practices.

5. Wait for conditions, not guessed time

Fixed sleeps such as “wait five seconds” add their full delay even when the page is ready sooner, and can still fail when a slower run takes longer. Wait for the condition the test actually needs: an element becoming visible, a response completing, or a user-visible state changing. Use the framework’s condition-based wait and assertion features.

There are cases where a short delay is intentional, such as verifying behavior after a debounce interval, but make the timing part of the behavior being tested and keep it as small as the requirement allows. Do not use a generic sleep as a substitute for understanding why the application is not ready.

6. Use parallel CI and affected-test selection with care

Once tests are independent, parallel workers can reduce wall time. Begin with a conservative worker count, then watch both duration and machine saturation. More workers can make a run slower when processes compete for CPU, memory, browser resources, or a shared service. If tests are distributed across CI jobs, split long spec files at sensible feature boundaries and compare actual timings after the change.

For a Playwright suite, a minimal CI command can run the suite with a bounded number of workers:

npx playwright test --workers=4

Choose the worker count for the capacity of the runner and the test environment; four is only an example. Playwright also documents sharding, for example:

npx playwright test --shard=1/4
npx playwright test --shard=2/4
npx playwright test --shard=3/4
npx playwright test --shard=4/4

Run each shard as a separate CI job when appropriate. Ensure that the jobs do not write conflicting shared state. Playwright’s CI guide shows CI execution and sharding patterns.

Playwright’s --only-changed option can serve as a preliminary pull-request pass for likely affected tests. Treat it as prioritization for early feedback, not as a replacement for the broader checks your CI policy requires. See the Playwright CI documentation.

7. Keep failure evidence useful and selective

A fast suite is not useful if failures cannot be diagnosed. Preserve enough information to reproduce and explain failures: clear test names, relevant logs, screenshots, traces, or state snapshots when appropriate. Capture expensive diagnostics selectively. Playwright documents tracing on the first retry in CI and warns that tracing every test has a performance cost; see its Trace Viewer guide.

Retries can prevent a transient failure from blocking a run, but a passing retry does not make the underlying test reliable. Keep retry counts low, record which tests are flaky, and investigate timing assumptions, shared state, environment load, and external dependencies. Cypress likewise advises using retries cautiously and addressing root causes in its test retries guidance.

8. Troubleshooting slow or unreliable E2E suites

Symptom Likely cause Practical fix
One test takes much longer than the rest Repeated UI setup, fixed sleeps, slow real network calls, or a long application path Inspect its timeline; remove unnecessary setup, wait on actual conditions, and stub dependencies only when the test is not meant to validate them.
A whole spec is slow Too many cases in one file, repeated expensive setup, or test data accumulation Use duration data to find repeated work; split very long specs along feature boundaries and check whether setup can be shared safely.
Tests pass alone but fail in a suite Order dependence, leaked state, shared accounts, or data collisions Run tests in different orders and in parallel; assign isolated data and remove dependencies on earlier tests.
Tests fail only in CI Different resource limits, network conditions, browser configuration, or environment setup Compare local and CI evidence, inspect runner saturation, and make readiness waits condition-based. Avoid simply increasing timeouts without finding the slow condition.
More workers increase total time CPU or memory contention, shared service limits, or test-state conflicts Reduce workers, isolate external state, and increase capacity only when measurements show the runner is the constraint.
Retries hide recurring failures Race conditions, unstable dependencies, or non-independent tests Track retry outcomes and fix the underlying timing, isolation, or dependency issue; retain retries only as a low-count containment measure.
Splitting specs does not help Specs are already short, or browser startup and reporting overhead dominate Compare wall time and resource use before and after; avoid splitting files that are already below the project’s useful granularity.
Tests pass, but a visual check misses a broken page The test verified a locator or response without checking the rendered output that matters Add an assertion for the user-visible state. If reviewing a page artifact is part of the workflow, capture it deliberately and keep that capture separate from functional assertions.

9. A practical rollout checklist

  1. Choose critical journeys: name the user outcomes and system risks that need full-stack confidence.
  2. Move routine checks down: cover logic and component behavior at smaller test levels where they are sufficient.
  3. Measure: record slow tests and specs in representative local and CI runs.
  4. Isolate: make each test own the state and data it needs.
  5. Remove measured waste: replace guessed sleeps, repeated UI setup, and irrelevant real network calls where appropriate.
  6. Parallelize gradually: add bounded workers or CI shards and check for contention and state collisions.
  7. Keep evidence actionable: capture enough failure detail to debug, without paying diagnostic overhead on every passing test.
  8. Review flake data: treat retries as a signal to investigate, not as proof of health.

10. Capture page evidence without maintaining browser capture code

When a workflow needs screenshots of pages or rendered states, a browser automation script can capture them as part of a test or review process. Choose a screenshot method based on the actual need: whether it must share test state, capture a particular point in the journey, or produce a reusable artifact. Do not add screenshot capture to every test if it slows the suite and the evidence is not useful.

DIY: capture a page with Playwright

This runnable Node.js example opens a page and saves a full-page screenshot. Install Playwright and its browser first:

npm install -D playwright
npx playwright install chromium
// save as capture.mjs
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded', timeout: 30000 });
  await page.locator('body').screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}
node capture.mjs

For tests, capture only on failure or for selected cases if the artifact helps diagnosis. A full-page screenshot may not include content that appears only after scrolling; trigger the relevant interaction or scroll/load behavior before capture when that is part of the scenario. Avoid brittle selectors, and wait for the actual page state needed for the screenshot rather than adding a fixed delay.

Or skip the browser setup

ScreenshotNeo provides a screenshot API and MCP server. Its one-call API returns an image or PDF, and the screenshot options can be used separately from a Playwright test when a standalone capture is enough. The API also supports full-page capture, a CSS selector, device presets, viewport and scale, wait conditions, custom headers and cookies, and more; see the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo, then sign up for 1,000 free screenshots a month, no card required.

Performance, reliability, and cost trade-offs

  • Performance: Removing repeated setup and unnecessary waits reduces per-test time. Parallelism can reduce wall time, but only if workers have enough resources and do not contend for shared state.
  • Reliability: Independent tests, controlled data, and condition-based waits make failures more reproducible. A retry can help diagnose or contain intermittent issues, but it cannot establish that the test is healthy.
  • Maintenance cost: Every E2E path depends on more of the system and can be affected by UI changes or slow services. Keep coverage purposeful and use smaller tests for behaviors they can prove.
  • CI cost: More workers and shards can trade lower elapsed time for more concurrent compute and operational complexity. Measure against the team’s actual CI capacity rather than assuming that more concurrency is always faster.
  • Capture cost: ScreenshotNeo states that only clean shots are billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Its plans include 1,000 monthly shots free, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. All features are on every plan.

FAQ

How many E2E tests should a project have?

There is no universal count or percentage. Cover the critical journeys and system behaviors that smaller tests cannot establish, including important error classes, and keep the suite maintainable.

Should every pull request run the entire suite?

Use affected-test selection as an early feedback pass when it fits your workflow, while retaining the broader CI coverage needed to catch interactions and unrelated regressions.

Are retries a good way to make the suite reliable?

Low-count retries can contain an intermittent failure or collect diagnostic evidence. Track retry outcomes and fix the underlying cause rather than treating a retry pass as proof of reliability.

Should screenshots be captured on every test?

Capture them when they help review or diagnose a relevant state. For routine passing tests, selective failure artifacts can avoid unnecessary storage and capture overhead.

Sources