ScreenshotNeo

BlogGuides

How to Speed Up Regression Testing: A 3-Part Guide

Speed up regression testing by selecting relevant tests safely, balancing parallel workers, and reducing slow setup and flaky reruns.

By the ScreenshotNeo team4 October 20267 min read

To speed up regression testing, reduce unnecessary work in three places: select tests relevant to a change when you can do so safely, run independent tests in parallel, and lower the cost of slow setup and test levels. Also address flaky tests, which cause reruns and investigation. These methods do not guarantee a fixed speedup: results depend on your suite, infrastructure, test dependencies, and how safely you can select tests.

Use a broad test run as a fallback when selection is uncertain, and keep broader validation in your release process. Measure your own suite before and after each change so faster feedback does not quietly reduce coverage.

1. Run the tests that matter for the change

Change-aware test selection runs a subset associated with a code change, giving developers earlier feedback on likely regressions. It is incremental validation, not proof that every unselected behavior still works.

Start with the foundations: map changed code to tests, identify stale or ineffective tests, improve test infrastructure, and make the existing suite run efficiently. AWS recommends addressing these basics before adopting advanced test-selection methods. [AWS Well-Architected DevOps Guidance](https://docs.aws.amazon.com/wellarchitected/latest/devops-guidance/qa.ft.4-balance-developer-feedback-and-test-coverage-using-advanced-test-selection.html)

Set a safe selection policy

  1. Run the selected tests for quick feedback. Make the selection method and its boundaries visible to developers.
  2. Fall back to a broader run when the change cannot be analyzed. A missing or uncertain mapping should not silently become an empty or incomplete test run.
  3. Keep full or broad runs at an appropriate cadence. Run them on a schedule, before release, or at another point where your team needs wider validation.
  4. Review selection misses. When a defect escapes, check whether the changed code was unmapped, a test was missing, or the fallback policy failed.

Microsoft’s Azure DevOps Test Impact Analysis (TIA) illustrates both the value and boundaries of a specific implementation. It selects impacted tests, previously failing tests, and newly added tests, and can fall back to all tests if it cannot reason about a commit. Its documented scope is managed code and a single-machine topology; HTML or CSS changes are examples that can cause a full-suite fallback. TIA also supports configured periodic full runs. These constraints describe Microsoft’s feature, not every test-selection system. [Microsoft Test Impact Analysis](https://learn.microsoft.com/en-us/azure/devops/pipelines/test/test-impact-analysis?view=azure-devops)

When test selection is a poor fit

Be cautious when tests have weak or missing links to the code they exercise, changes span many components, or selection coverage is hard to audit. A broad run may be the safer choice when the selector cannot explain what it omitted. Track selection coverage and fallback frequency alongside elapsed time.

2. Split independent tests across workers

Parallel testing divides work across agents or machines. The goal is to make workers finish at similar times: one slow shard can determine the total duration even when the other workers are idle. Azure Pipelines documents suite slicing for parallel execution; Cypress Cloud documents parallelization and load balancing. Their product documentation describes those products’ capabilities, not a universal performance guarantee. [Azure parallel testing](https://learn.microsoft.com/en-us/azure/devops/pipelines/test/parallel-testing-any-test-runner?view=azure-devops) · [Cypress Cloud Smart Orchestration](https://docs.cypress.io/cloud/features/smart-orchestration/overview/)

Make concurrency safe before increasing it

  • Give each test or worker isolated data where possible.
  • Make setup and cleanup dependable, including after a test fails.
  • Check for order dependence, shared files, global state, and reused accounts or records.
  • Increase worker count gradually and watch for resource contention, longer setup, and newly exposed race conditions.

Parallel execution can expose hidden dependencies. pytest’s guidance describes how uncontrolled state and test ordering cause flaky results, and notes that parallel runs can reveal missing cleanup or tests that modify shared state. Isolate fixtures and data before scaling concurrency. [pytest flaky-test guidance](https://pytest.org/en/latest/explanation/flaky.html)

Balance shards using observed durations

Equal numbers of tests do not necessarily make equal work: one browser test may take much longer than several unit tests. Use per-test or per-file timing data to distribute likely durations, then review the slowest worker after each run. If the CI system supports dynamic assignment or load balancing, compare it against static shards using the same suite and environment.

3. Reduce test cost and flaky reruns

Profile first. Separate time spent in compilation, environment setup, authentication, test execution, network waits, and cleanup. Optimize the largest avoidable cost before changing several parts at once.

Cypress recommends choosing an appropriate test level, caching authentication, stubbing network requests, and setting state programmatically instead of navigating through slow UI setup. Its guide also discusses parallelization, tags for CI tiers, spec prioritization, and cancellation after enough failures. Treat these as Cypress product guidance; they are not independent benchmarks or requirements for every test stack. [Cypress performance guide](https://docs.cypress.io/app/guides/test-performance)

Practical ways to lower per-test time

  • Use the least expensive test level that answers the question. Keep broad end-to-end coverage for behavior that needs it; test logic that does not need a browser at a cheaper level.
  • Avoid repeating expensive setup. Reuse or cache authentication state when that is safe, and set up test state through supported APIs or fixtures where appropriate.
  • Stub external services when the test is not validating the integration. Keep separate tests that exercise the real integration so stubbing does not remove needed coverage.
  • Use CI tiers deliberately. Run a fast, relevant tier for frequent feedback and broader suites at suitable checkpoints.
  • Prioritize useful failures. Run high-signal tests early if your runner supports it; configure cancellation carefully so the team still gets the results it needs.

Fix flakiness instead of relying on retries

pytest defines flaky tests as tests that fail intermittently or sporadically. Causes can include uncontrolled system state, ordering dependencies, and incomplete cleanup. A flaky failure can consume time in reruns and investigation and make developers less confident in results. [pytest flaky-test guidance](https://pytest.org/en/latest/explanation/flaky.html)

  1. Record the failing test, run, worker, and relevant state so the failure can be reproduced.
  2. Repeat the test in isolation and in different orders to check for shared state or order dependence.
  3. Fix timing assumptions, cleanup, test data isolation, or external dependency behavior where possible.
  4. Use retries only as a temporary containment or carefully scoped policy. Cypress notes that retry execution cost compounds when retries are configured carelessly. [Cypress performance guide](https://docs.cypress.io/app/guides/test-performance)

Measure whether the changes helped

Compare the same suite on the same or a representative environment before and after each change. The reviewed documentation does not provide a neutral, comparable benchmark across CI vendors, so use your own baseline rather than assuming a published speedup applies to your project.

Measure What it tells you
Time to useful feedback How soon a likely regression reaches the developer.
Selection coverage and fallback rate Which tests are omitted, how relevance is inferred, and when broader validation runs.
Worker finish times Whether shard imbalance leaves workers idle behind one long-running shard.
Flake and rerun rate How much time is spent repeating or investigating unreliable failures.
Infrastructure and maintenance cost Whether extra workers or selection tooling cost more to operate than the feedback time they save.

Or skip the browser setup

If regression work includes capturing pages for visual checks or documentation, ScreenshotNeo can return a screenshot or PDF with one GET request. The do-it-yourself browser workflow above still applies to running and optimizing your test suite; this API is for capturing a web page as an input to your workflow.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python and Node.js calls are available below. See the ScreenshotNeo API documentation for request options.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners are accepted and removed before capture; 60+ known consent platforms, newsletter popups, and chat widgets can be removed, and each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does running fewer tests mean regression risk is lower?

No. Selection is useful only to the extent that its mapping is reliable. Keep a safe fallback and broader validation so omitted tests still run on an appropriate cadence.

Should every test be parallelized?

No. Tests that share mutable state or depend on order need isolation or restructuring before they can run safely together.

Are retries a fix for flaky tests?

Retries can contain some failures temporarily, but they do not explain the cause. Investigate reproducibility, shared state, cleanup, and timing before relying on retries.

How can a team tell whether its strategy is working?

Compare feedback time, coverage and fallback behavior, worker balance, flaky reruns, and infrastructure effort against a baseline from the same suite and environment.