ScreenshotNeo

BlogEngineering

How to Optimize Tests for Continuous Integration

Speed up CI feedback by measuring bottlenecks, ordering tests by signal, caching carefully, fixing flakes, and parallelizing isolated work.

By the ScreenshotNeo team4 October 20267 min read

To optimize tests for continuous integration (CI), measure where pipeline time goes, run quick and relevant checks early, remove redundant work, cache repeatable dependency downloads, fix flaky tests, and parallelize only independent work. Re-measure after each change. Keep blocking checks reliable: a faster pipeline that misses important failures is not an improvement.

This guide uses GitHub Actions YAML for concrete examples, but the sequence applies to other CI systems. Adapt the syntax to your runner and test framework.

1. Measure the pipeline before changing it

Record the baseline duration across several representative runs. Break it down into queue time, environment setup, dependency installation, build, test execution, artifact handling, and teardown. Then inspect durations at the job, suite, file, and individual-test levels where your tools allow it.

Look for repeated costs, not just the longest individual test. A slow setup step repeated across many jobs can matter more than one slow test. Separate runner wait time from actual work: buying parallel capacity will not fix slow test code, and changing test code will not fix a busy runner queue.

What you observe Likely area to investigate
Long time before a job starts Runner availability, queueing, concurrency limits
Similar setup time in many jobs Repeated environment or dependency setup
One suite dominates execution Slow tests, expensive fixtures, unbalanced shards
Large variation between runs Flakiness, contention, variable external services
More workers but little wall-time gain Resource contention, serial work, poor shard balance

GitLab recommends collecting test-duration information and investigating slow-test patterns; splitting a test file by itself does not make its tests faster. See its unhealthy tests guidance.

2. Run useful checks early

Order stages by how quickly and reliably they can produce an actionable failure. A common sequence is formatting and static checks, focused unit tests, broader unit tests, integration tests, and end-to-end tests. The right order depends on your project: prioritize checks that are fast, relevant to the change, and likely to catch a real problem.

Run tests related to changed code on a pull request only when the mapping from changed files to affected tests is dependable. If that mapping is incomplete, retain a broader blocking suite or run the broader coverage in another required stage. Avoid silently dropping checks because they are slow.

GitLab’s strategy recommends progressive execution and fast feedback, and its efficiency guidance discusses running quick-failing work earlier and avoiding jobs that do not apply to a change. These are useful principles; exact pipeline features vary by CI platform. See the testing strategy and pipeline efficiency guidance.

3. Example: staged checks in GitHub Actions

This illustrative workflow runs a quick check before the broader test suite, and caches npm’s download cache using the lockfile as the dependency key. Replace the commands and Node version with your project’s requirements.

name: CI
on:
  pull_request:
  push:
    branches: [main]
jobs:
  quick:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npm run lint
      - run: npm run test:unit:changed
  full:
    needs: quick
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npm run test:unit
      - run: npm run test:integration

The command test:unit:changed is project-specific; do not copy it without implementing a trustworthy test-selection strategy. If you do not have one, run the full unit suite in the quick job or make it the required suite. In this example, the broad job depends on the quick job, so it starts only after the quick checks pass. You might instead run independent suites concurrently when that reduces feedback time without exhausting shared resources.

4. Remove redundant work and improve slow tests

  • Assign an owner and purpose to each suite. Remove duplicate checks when they provide no distinct coverage.
  • Inspect expensive fixtures and repeated setup. Reuse immutable setup where safe; avoid shared mutable state that makes tests order-dependent.
  • Replace broad or slow checks with focused assertions only when the coverage remains meaningful.
  • Investigate unnecessary network calls, service startup, oversized build images, and slow test discovery.
  • Do not treat splitting a file as a speed fix. Splitting can help distribute independent tests, but the underlying test work remains.

5. Cache dependencies carefully

Caches help when they avoid repeated downloads or reusable build work. Cache keys should reflect the dependency state and relevant environment. For package managers, the lockfile is often a useful key input. A stale or overly broad cache can create confusing failures; a cache that is expensive to restore and save can cost more time than it saves.

Track cache hit rates and compare total job duration, including restore and save steps. Prefer caching downloads when package installation still needs to validate and link dependencies. Avoid treating generated outputs as reusable unless their inputs, tool versions, and invalidation rules are understood.

6. Fix flaky tests before adding retries

A flaky test produces different outcomes without a relevant code change. Reproduce it alone and under the same concurrency as CI, then inspect timing assumptions, test ordering, shared state, service readiness, and resource pressure. Wait for a meaningful condition such as an element becoming available or a service becoming healthy; arbitrary sleeps make feedback slower and can still fail under different load.

Google’s testing guidance specifically cautions against arbitrary delays and recommends diagnosing the source of flakiness: Test Flakiness. Retries can help identify intermittent failures or reduce disruption temporarily, but they can also hide defects. If you quarantine a test, assign an owner and review date, and keep its lost coverage visible.

7. Parallelize only independent work

Parallel workers reduce elapsed time when there is enough independent work and the runner has spare CPU, memory, and service capacity. They also increase resource use and can expose tests that share writable files, databases, accounts, ports, or other state.

  1. Make tests isolated: use separate temporary directories, unique records, and independent fixtures.
  2. Divide work into balanced shards using historical durations when available.
  3. Run the shards under representative CI load and compare wall time, runner minutes, memory, and stragglers.
  4. Adjust shard sizes or worker counts based on the slowest shard and resource contention.

The gtest-parallel project warns that concurrent tests must not write to shared resources. Its implementation applies to Google Test; the isolation principle applies more broadly.

8. Troubleshooting common CI test slowdowns

Symptom Possible cause What to try
Jobs spend a long time waiting to start Runner queue or concurrency cap Measure queue time separately; review runner availability and job fan-out.
Dependency cache has little effect Cache misses, restore overhead, or downloads are not the bottleneck Check hit rate and timed steps; verify keys include the lockfile.
Tests fail only in parallel Shared state, fixed ports, shared files, or resource contention Isolate writable resources and use unique test data; reduce workers to locate contention.
Retries pass after an initial failure Timing race, unstable dependency, or resource pressure Reproduce and fix the underlying issue; track retry frequency rather than accepting it as normal.
A changed-file test job misses failures Incomplete mapping between source files and tests Expand the dependency mapping or restore broader required coverage.
Sharding does not reduce total time Uneven shards, serial setup, or saturated workers Compare per-shard durations and setup costs; rebalance and validate runner capacity.
Tests pass locally but fail in CI Environment difference, ordering dependence, or load-sensitive timing Compare tool versions and configuration; reproduce with CI-like concurrency and services.

9. Evaluate the trade-offs and re-measure

After each meaningful adjustment, compare representative runs against the baseline. Consider feedback time from commit to useful result, coverage and the risk of missed failures, repeatability, runner and service cost, and maintenance burden. Do not promise a fixed percentage speedup: results depend on the repository, tests, and CI environment.

Keep blocking checks stable and preserve broad coverage at an appropriate stage. A practical optimization is one that reduces time developers wait while maintaining a dependable signal about code quality.

Or skip the browser setup

If browser screenshots are part of your CI checks, you can capture a page with one API request instead of managing browser setup. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. It removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Sign up for free.

FAQ

Should every test run on every pull request?

Run the checks needed to catch likely regressions before merge. Use selective execution only when its coverage mapping is reliable, and preserve broader checks in a suitable required stage.

Should I buy faster CI runners?

First establish whether queue time, CPU, memory, or test execution is the bottleneck. More capacity helps some constraints, but will not correct inefficient tests or shared-state failures.

Is caching always faster?

No. Measure restore and save time alongside the work avoided, and ensure cache invalidation matches dependency changes.