How to Run End-to-End Tests Faster in CI
Speed up CI feedback by measuring the bottleneck, cutting avoidable setup, and scaling test work only when your suite and runners can support it.
To run end-to-end tests faster in CI, first measure where the time goes, then remove avoidable setup and add concurrency only while the runner has capacity. Start with a comparable baseline, inspect slow specs and job imbalance, install only the browsers you test, wait for the application to be ready instead of sleeping, and then evaluate local workers or multiple CI jobs. Keep tests isolated and watch failures and retries as closely as duration.
There is no universal worker count or machine count. A change that helps a suite with independent, CPU-light tests can slow one that already saturates memory, depends on shared data, or spends most of its time starting runners and browsers.
1. Find the bottleneck before changing concurrency
Record the total wall-clock time from CI start to usable test results. Break it down where possible: dependency setup, browser download, application startup, test execution, artifact upload, and reporting. Note the slowest specs, the duration of each job or shard, machine CPU and memory use, and failure and retry rates.
Compare runs with the same test selection and a comparable runner configuration. A faster test command does not shorten the developer feedback loop if setup or report generation dominates the job.
- Long setup, short tests: look at browser installation, dependency caching, app startup, and job initialization.
- A few very slow specs: profile those specs and their fixtures, network calls, waits, and cleanup before adding machines.
- One job finishes much later than the rest: inspect spec or shard assignment and balance. Adding more jobs may not help if one shard still owns the long-running tests.
- Duration grows as workers increase: check CPU and memory saturation, browser crashes, database contention, and rate limits.
- Tests finish quickly but failures surface late: use failure and spec history to prioritize diagnostics; do not mistake early failure reporting for a shorter successful run.
Cypress documents run duration, slow tests, resource contention, and machine/spec allocation as useful diagnostic signals in its test performance guide. Its Machine view can reveal an idle machine waiting while another runs long specs. Cypress also describes a failure-prioritization scenario where a failure found around 20 minutes into a run may surface in 2–3 minutes when earlier failing specs run first and remaining work is cancelled. That is a Cypress example, not a general speed guarantee.
2. Remove avoidable setup time
Install only the browser engines in your test matrix
Browser downloads add time and consume disk space. If a job tests only Chromium, install Chromium in that job rather than every supported browser. Keep the full browser matrix in the jobs that actually need it; narrowing one job must not accidentally remove required coverage.
# Playwright: install Chromium and its system dependencies when Chromium is the target
npx playwright install chromium --with-deps
Playwright recommends installing only the browsers needed for the project in CI. See its best practices and check the command against the Playwright version pinned in your lockfile.
Wait for readiness, not a guessed delay
A fixed sleep is often both wasteful and unreliable: a short delay lets tests race the application, while a long one adds idle time to every run. Gate the test command on an actual readiness response or endpoint. Cypress documents start and wait-on options in its GitHub Action for starting an app and waiting for it to respond.
# Cypress GitHub Action example; pin action versions according to your repository policy
- name: Cypress run
uses: cypress-io/github-action@v6
with:
start: npm run start
wait-on: 'http://localhost:3000'
browser: chrome
This example is specific to the Cypress GitHub Action. For other runners, use that CI provider’s readiness mechanism or a project-owned wait command that checks the real health endpoint. Do not replace a readiness check with sleep 30.
3. Use local workers only while the machine has headroom
Parallel workers can shorten independent test work on one runner, but each browser and test process consumes resources. Increase worker count gradually and compare wall time, resource usage, and failures. If duration stops falling or rises, return to the last stable setting and investigate contention.
Playwright worker configuration
import { defineConfig } from '@playwright/test';
export default defineConfig({
// Conservative baseline. Try a small increase only after measuring runner headroom.
workers: process.env.CI ? 1 : undefined,
});
Playwright’s general parallelism guide shows configurable worker counts, including an example of using two workers in CI. Its CI guide recommends one worker as a stability and reproducibility baseline. These are choices to measure, not a promise that one or two workers is fastest for every suite.
Playwright workers are separate processes and each starts a browser. Tests that run in parallel must not depend on shared module globals or mutable shared state. See Playwright parallelism and its CI guide.
Cypress concurrency is a different mechanism
Cypress can distribute specs across machines in a recorded run using Cypress Cloud. Its parallelization command is not a generic Cypress or framework-agnostic switch:
# Cypress only: requires a recorded run and a secret configured in CI
npx cypress run --record --key="$CYPRESS_RECORD_KEY" --parallel
Store the record key in your CI secret facility; do not commit a real key in workflow source. Cypress Cloud balances specs across available machines for recorded parallel runs. See the Cypress CI overview and performance guide.
4. Scale across CI jobs with sharding
Use sharding when a single machine has reached its useful worker limit and the suite can be divided into independent pieces. Multiple jobs add runner startup, checkout, setup, browser, and artifact overhead. They help when the test work saved exceeds that overhead and the shards finish at roughly similar times.
Playwright: split the suite into shards
For four CI jobs, each job runs a different shard index with the same total:
# Job 1
npx playwright test --shard=1/4
# Job 2
npx playwright test --shard=2/4
# Job 3
npx playwright test --shard=3/4
# Job 4
npx playwright test --shard=4/4
Configure the CI provider to start these as separate jobs, retain each job’s report or blob output, and merge results using Playwright’s report guidance. The exact artifact and matrix syntax depends on the CI provider; follow Playwright’s CI sharding examples. Keep shard indexes and the total consistent across the jobs in a run.
Cypress: use Cloud’s spec allocation
For Cypress, recorded parallel runs let Cypress Cloud distribute specs among the available CI machines. A CI matrix or equivalent starts the machines, and each runs the Cypress parallel command with the same project and record credentials. Follow the Cypress GitHub Actions guide for the provider-specific setup. Cypress notes that rolling runner-image updates can temporarily give parallel Linux jobs different browser versions; its guide discusses using a consistent Cypress browser Docker image for that case.
| Approach | Where work runs | Coordination and reporting | Best fit |
|---|---|---|---|
| Local workers | Several workers on one machine | Runner and test framework coordinate local work | Independent tests and unused CPU/memory on a runner |
| Playwright shards | Separate CI jobs or machines | Shard index/total are configured in jobs; reports can be merged | Playwright suites that outgrow one runner |
| Cypress Cloud parallel run | Separate CI machines | Recorded run and Cloud spec distribution | Cypress teams using recorded runs and Cloud |
No approach guarantees lower elapsed time. Measure runner startup and browser setup along with test duration, and compare the slowest shard to the median. One overloaded shard can set the completion time for the entire suite.
5. Make parallel tests safe and repeatable
Before raising concurrency, verify that tests can run in any order and at the same time. Parallel runs often expose hidden dependencies that were masked by serial execution.
- Give each test or worker unique accounts, records, filenames, and other mutable data where practical.
- Reset state at the test boundary; do not assume a previous test created a shared fixture.
- Serialize access to a genuinely shared external resource when it cannot be partitioned. Playwright documents named locks for resources shared by tests.
- Check external API quotas, database connection limits, and application capacity before multiplying concurrent traffic.
- Keep browser and runner versions consistent across parallel jobs where the CI image can change independently.
Playwright’s guidance is to isolate tests so they can run independently. Isolation improves reproducibility as well as enabling safe parallel work. If a test only passes after another test has run, repair the dependency before scaling.
6. Use retries to measure flakiness, not to hide it
Retries rerun work and can make a run longer, especially when applied broadly. Cypress warns that if 10% of tests have some flakiness, a high global retry count can add many minutes to a CI run. That 10% is an illustrative condition in its warning, not a claim about typical suites.
Track which tests retry, how often they fail on the first attempt, and whether failures cluster by browser, shard, or resource pressure. Fix the cause—such as shared state, a readiness race, or an unstable dependency—then keep any retry policy explicit and bounded. A retry may help a transient failure policy, but it is not a speed optimization.
7. Measure the result and stop at the useful limit
After each change, compare the same suite on comparable runners. Keep a short record of:
- Total elapsed CI time and time to first actionable failure.
- Slowest specs and the spread between jobs or shards.
- CPU, memory, browser crashes, and application or database saturation.
- Failure, retry, and cancellation rates.
- Setup and artifact overhead added by extra jobs.
Keep a change when it lowers the feedback time without an unacceptable increase in instability or cost. Stop adding workers or machines when contention and overhead erase the gain. Cypress’s published performance guide contains product-specific illustrations, including a Kitchen Sink run changing from 1:51 serial to 59 seconds with a second machine. Treat those as Cypress examples rather than an estimate of your project’s outcome.
Common problems and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| More workers make the suite slower | CPU or memory contention; each worker launches a browser | Reduce workers, inspect machine utilization, and profile the slow specs. |
| One shard runs much longer | Uneven spec duration or poor assignment | Inspect per-spec timing and shard allocation; rebalance using runner-supported history or grouping. |
| Tests fail intermittently after parallelizing | Shared accounts, records, files, or external resources | Allocate unique state per worker or serialize access to the shared resource. |
| Tests fail at startup, or a fixed delay is needed | Application readiness is not checked | Wait for a health endpoint or response before starting tests. |
| CI time is dominated by browser setup | Unneeded browser engines are downloaded in every job | Install only the browser engines required by that job’s coverage. |
| Cypress parallel option does not distribute specs | The command is being used without a recorded run and Cypress Cloud configuration | Configure the Cypress project and CI secrets, then follow the Cypress recorded-run setup. |
| Parallel Cypress jobs behave differently | Jobs may have different browser versions during runner image rollout | Use a consistent browser image as described in Cypress’s Linux CI guidance. |
| Retries increase duration without stabilizing the suite | Retries repeat work while the underlying flaky cause remains | Use retry records to find and fix the failure source; keep retries bounded. |
| Playwright report is incomplete after sharding | Shard artifacts were not all retained or merged | Upload each job’s report output and follow Playwright’s report merge guidance. |
Or skip the browser setup
If CI needs screenshots of pages for visual checks, review artifacts, or debugging, ScreenshotNeo is a website screenshot API and MCP server: send one GET request with a URL to receive a PNG, JPEG, WebP, or PDF. It can complement an E2E suite when the task is capturing a page artifact rather than exercising an interactive user journey.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Only clean shots are billed, and response headers report the page verdict and billing status.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Should CI always use one Playwright worker?
One is Playwright’s conservative CI recommendation for stability and reproducibility. Measure a small increase on your runner before deciding whether additional workers help your suite.
How many CI machines should I use?
There is no generally supported number. Add machines while balanced test work saves more time than runner startup, browser setup, and reporting overhead.
Does Cypress parallel mode work without Cypress Cloud?
The documented Cypress parallel mode across machines uses a recorded run and Cypress Cloud’s spec distribution. Playwright sharding is a separate mechanism and does not require Cypress Cloud.
Should I retry failed tests until CI is green?
No. Retries repeat work and can hide nondeterminism. Use retry data to identify flaky tests and fix their causes; if retries are part of a policy, keep them bounded.


