How to Reduce Automated Test Execution Time
Speed up automated test feedback by measuring bottlenecks, parallelizing safely, sharding CI work, and reducing flaky reruns without trusting incomplete coverage.
To reduce automated test execution time, first measure where elapsed time goes. Then parallelize independent tests, shard work across CI machines when one runner is the limit, shorten initial feedback with carefully scoped changed-test runs, and fix flaky tests that trigger reruns. Keep a full-suite run as a correctness check whenever you use a selection heuristic. More workers do not guarantee a faster or more reliable suite.
1. Measure the baseline before changing the suite
Record total wall-clock time in both local and CI environments. Where your framework provides it, capture durations by test or file. Separate time spent doing test work from setup and teardown, waiting on services, environment startup, and scheduling. This helps distinguish a slow test from a slow runner or an expensive shared fixture.
- Run the same suite on a representative machine or CI runner and save the elapsed time.
- Collect per-test or per-file durations if available; note the slowest tests and unusually expensive setup.
- Check whether long waits are intentional synchronization or fixed sleeps that could be replaced with condition-based waits.
- Repeat the measurement after each change under comparable conditions, including the same test selection and runner size.
Do not treat a single fast run as proof of improvement. Resource contention, caches, and variable external services can change elapsed time. Compare repeated runs and watch failure and retry rates alongside duration.
2. Parallelize tests that are safe to run concurrently
Parallel execution reduces elapsed time when tests have independent work and the machine has spare CPU, memory, and service capacity. It can also expose hidden dependencies: tests may share a database, filesystem paths, accounts, global state, or cleanup assumptions. Isolate data and make teardown reliable before raising concurrency.
pytest with pytest-xdist
Install pytest-xdist in the project environment, then run the suite with automatic worker selection:
python -m pip install pytest-xdist
pytest -n auto
pytest-xdist distributes tests across worker processes. Its documented auto setting selects workers based on the number of physical CPU cores. Treat that as a starting point, not a universal optimum. For example, compare a bounded worker count with automatic sizing:
pytest -n 4
Choose a count based on measured wall-clock time and stability. Too many workers can increase memory pressure, overload a database or browser service, and add scheduling overhead. Consult the pytest-xdist distribution guide.
Playwright Test workers
Run Playwright Test locally with an explicit worker count:
npx playwright test --workers=4
Playwright can run test files in parallel, and their order is not guaranteed. In CI, Playwright recommends one worker by default to favor stability and reproducibility; capable self-hosted systems may choose parallel workers after measuring their capacity. For multiple machines, use sharding as described below. See the official Playwright parallelism guide and CI guidance.
Check isolation before increasing worker count
- Give each test unique records, accounts, directories, and other mutable resources.
- Avoid relying on another test to create state or perform cleanup.
- Reset shared state explicitly, including process-level or global state where relevant.
- Make tests independent of execution order and safe to retry.
- When parallel-only failures appear, investigate shared state and timing assumptions instead of masking them with repeated retries.
3. Shard the suite across CI jobs
Sharding distributes test work across separate CI jobs or machines. It can reduce elapsed time when a single runner is the bottleneck and the suite can be split into independent groups. It may increase total compute use and setup overhead, and it adds work to aggregate reports and diagnose failures.
Playwright Test supports shards. A typical two-shard arrangement uses one command per job:
# CI job 1
npx playwright test --shard=1/2
# CI job 2
npx playwright test --shard=2/2
Configure those commands as separate CI jobs or machines in your provider. Keep the shard count aligned with available capacity; twice as many shards do not necessarily halve elapsed time. Uneven test durations can leave one shard running long after the others finish, so use duration data to assess balance. Playwright documents this approach in its CI guide.
4. Use changed-test runs for an early signal
A targeted run can shorten the first feedback loop while a developer is working. Its coverage is incomplete by design, and dependency analysis may miss tests affected indirectly by a change. Playwright describes this approach as heuristic and advises following it with the full suite.
- Use the project’s supported selection mechanism to run tests directly related to the change.
- Label that result as preliminary; do not treat it as equivalent to a full-suite pass.
- Run the complete suite in CI or before merging, especially for shared code, configuration, fixtures, and cross-cutting changes.
Do not invent a dependency mapping if the test framework or project does not maintain one. A fast, incomplete run is useful for feedback, but it does not establish that unaffected tests remain unaffected.
5. Fix flaky tests to eliminate repeated work
Flaky failures cost time through reruns and investigation, and they weaken confidence in the suite. pytest’s documentation identifies shared system state, missing cleanup, and order dependence among causes of flakiness; parallel execution can make these problems visible.
- Shared state: use isolated test data and resources, or serialize tests that truly cannot be isolated.
- Missing cleanup: use reliable teardown so failed tests do not leave state for later tests.
- Order dependence: ensure each test establishes its own prerequisites.
- Timing assumptions: wait for observable conditions rather than assuming an operation finishes within a fixed delay.
- External dependencies: identify services and resource limits that can make results vary, then make their use explicit and diagnosable.
Retries can help reveal whether a failure is intermittent, but they do not repair its cause. Track retries as failures to investigate, not as a substitute for a trustworthy result. See pytest’s flaky test documentation.
6. Choose an approach using elapsed time, coverage, and cost
| Approach | When it can help | Main trade-off |
|---|---|---|
| More workers on one runner | Independent tests and spare machine capacity | Resource contention, shared-state failures, and overhead |
| CI sharding | One runner is the limit and separate machines are available | More compute and setup; reports and ownership can be harder to manage |
| Changed-test selection | Fast preliminary feedback during development | May omit relevant tests; follow with the full suite |
| Flake reduction | Reruns and investigation consume time | Requires finding and correcting the source of nondeterminism |
Compare changes using repeated wall-clock measurements, failure and retry rates, coverage expectations, runner capacity, and the complexity of maintaining the workflow. There is no general speedup percentage or optimal worker count that applies to every suite.
7. Troubleshoot common slowdowns and failures
| Symptom | Likely cause | What to do |
|---|---|---|
| More workers make the run slower | CPU, memory, database, or service contention; scheduling overhead | Reduce the worker count and compare repeated runs. Check the bottleneck before adding concurrency. |
| Tests fail only in parallel | Shared files, records, accounts, global state, or order assumptions | Isolate mutable resources, make setup self-contained, and ensure cleanup runs after failures. |
| CI is much slower than local | Different machine capacity, service latency, startup, or resource contention | Measure on the actual CI runner and separate test time from environment and service setup. |
| One shard determines the total duration | Uneven distribution of slow tests | Use duration data to balance work and verify that every shard is included in the workflow. |
| Changed-test run passes but full suite fails | The selection omitted a test affected indirectly by the change | Treat the targeted run as preliminary and use the full-suite result as the correctness check. |
| Retries hide intermittent failures | A flaky test or unstable external dependency | Record the retry and investigate state isolation, cleanup, ordering, and timing assumptions. |
Or skip the browser setup
If your automated workflow also needs website screenshots, ScreenshotNeo offers a screenshot API and MCP server. A single GET request can return an image or PDF, so you do not need to set up a browser for that capture. Its cookie and consent banner handling, newsletter popup removal, and chat widget removal run before the shot; each step can be turned off. Bot checks, blank pages, timeouts, and failed loads are never billed, and cache hits cost nothing. The response identifies the page verdict and billing status in headers. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
Read the ScreenshotNeo API documentation. Here is the one-call cURL example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page screenshots with lazy images loaded, element capture by CSS selector, dark mode, device presets and custom viewports, retina scale, PDF settings, HTML/CSS capture, custom CSS and JavaScript, click and hide selectors, wait conditions, request blocking, headers and cookies, user agent and authorization, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed image links, async jobs with signed webhooks, bulk capture, a usage API, and an OpenAPI spec. The parameter names used by other screenshot APIs also work. All features are on every plan.
Plans are Free with 1,000 screenshots per month and no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. Visit ScreenshotNeo or sign up free for 1,000 screenshots a month with no card.
FAQ
Should I always run the full suite before merging?
Use a full-suite run as the correctness check when an earlier changed-test run is heuristic or incomplete. Your project’s CI policy should define which full checks gate a merge.
Does automatic worker selection guarantee the fastest run?
No. pytest-xdist’s automatic setting uses physical CPU core count as a basis, while the best setting depends on the workload and the resources tests share.
Can retries make a flaky suite reliable?
Retries may reveal intermittency, but they do not remove shared-state, cleanup, ordering, or timing problems. Investigate and fix the cause.


