ScreenshotNeo

BlogEngineering

How to Manage Tests in a Continuous Integration Pipeline

Choose which tests run at each CI stage, keep feedback useful and fast, and handle flaky failures without losing confidence in your pipeline.

By the ScreenshotNeo team4 October 202613 min read

A reliable continuous integration (CI) pipeline runs the smallest checks that can catch a change early, then adds broader tests where they buy meaningful confidence. Put fast, relevant, stable checks on every pull or merge request; add integration and system tests for interactions; reserve end-to-end (E2E) tests for critical user journeys and deployment boundaries. Measure slow stages, parallelize only when it reduces elapsed time at an acceptable runner cost, and treat a retry that passes as evidence to investigate—not proof that a failure was harmless.

There is no universal pipeline layout, acceptable runtime, shard count, coverage target, or flake-rate threshold. Choose stages and merge rules based on your architecture, risk, test reliability, and available infrastructure. GitLab’s published testing strategy is a useful concrete example, but its tiers are GitLab’s own design. GitLab Testing Strategy

1. Decide what each pipeline stage is for

First identify what a change could break and the cheapest reliable test level that can detect it. Keep most tests at lower levels, where they are usually quicker and less expensive to maintain. Add broader tests to verify interactions and critical user behavior that lower-level tests cannot establish.

Test level What it can establish Typical pipeline placement Common tradeoff
Static checks and build Formatting, lint rules, types, compilation, packaging Early on every change Fast feedback, but not proof of runtime behavior
Unit A small function, class, or module behaves correctly in isolation Every pull or merge request Usually quick; mocks can hide integration problems
Integration Components work together across a boundary, such as a database or service Every request when practical; otherwise a later request tier or protected branch More setup and runtime than unit tests
System or feature A feature behaves across much of the application, sometimes through its UI Changes that affect user-visible workflows or important subsystems Broader confidence, with more environment and maintenance needs
End-to-end A critical user journey works across the deployed stack A focused smoke set at deployment boundaries; broader suite selectively or on a schedule Often the most expensive and failure-prone level to operate

This is a guide, not a quota. GitLab’s testing-level documentation says its own suite is weighted toward unit tests and that E2E tests are the smallest, most expensive layer. Its dated estimate for GitLab Community plus Enterprise Edition lists 218,459 unit tests (75.66%), 57,127 integration tests (19.79%), 12,444 system or feature tests (4.31%), and 704 E2E tests (0.24%). These are GitLab’s inventory figures as of 2025-02-03, not an industry benchmark or target for another project. GitLab testing levels and dated inventory

2. Map tests to change risk and feedback time

For each check, decide when it runs and what it blocks. A useful review asks: How quickly will the author see a result? What failure risk does this suite detect? How reliable is the signal? What runtime and infrastructure does it consume? Who owns failures? Does it block merging, deployment, or release?

  1. On every pull or merge request: run formatting, linting, type checks, build validation, and relevant fast unit tests. Keep these checks deterministic and actionable because they are the first gate.
  2. As request risk or test scope grows: add integration tests for changed boundaries, such as persistence, queues, APIs, or service adapters. In a large repository, use validated change-aware selection, but retain a broader path for changes whose impact cannot be predicted reliably.
  3. For important user-visible behavior: include system tests or a small E2E set when the feature crosses boundaries that lower-level tests cannot cover. Keep the set focused on journeys whose failure matters.
  4. At deployment boundaries: run smoke checks against the deployed environment. A failed staging or canary smoke test can stop promotion; a production post-deploy check may instead alert or trigger a rollback, depending on the team’s release process.
  5. On a schedule or higher-risk change: run broader suites, compatibility matrices, longer E2E scenarios, or expensive environment tests that are too slow for every change.

GitLab’s example places unit tests in merge-request pipelines, expands integration and system coverage in later tiers, and uses E2E checks for selected higher tiers or scheduled pipelines; its deployment example uses smoke checks. Adapt that structure to local risks rather than copying its tiers as a standard. GitLab’s placement example

3. Build a pipeline with clear gates

Here is a runnable starting point for GitLab CI. It uses a Python project with pytest, Ruff, and a build command. Replace the image, dependency installation, commands, test paths, and deploy placeholder with the commands your repository actually uses. The example runs fast checks first, then unit and integration suites, followed by a focused smoke job on the default branch. Jobs in a GitLab stage can run concurrently; stages proceed in order after earlier stages succeed. GitLab pipeline behavior

stages:
  - verify
  - unit
  - integration
  - smoke

variables:
  PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"

.python_setup:
  image: python:3.12
  before_script:
    - python -m pip install --upgrade pip
    - pip install -r requirements.txt
    - pip install pytest ruff
  cache:
    key: "pip-$CI_COMMIT_REF_SLUG"
    paths:
      - .cache/pip/

lint:
  extends: .python_setup
  stage: verify
  script:
    - ruff check .
    - ruff format --check .

build:
  extends: .python_setup
  stage: verify
  script:
    - python -m compileall -q src
    - python -m build
  artifacts:
    paths:
      - dist/
    expire_in: 1 day

unit:
  extends: .python_setup
  stage: unit
  script:
    - pytest -q tests/unit --junitxml=unit-results.xml
  artifacts:
    when: always
    reports:
      junit: unit-results.xml
    expire_in: 7 days

integration:
  extends: .python_setup
  stage: integration
  services:
    - name: postgres:16
      alias: db
  variables:
    POSTGRES_DB: app_test
    POSTGRES_USER: app
    POSTGRES_PASSWORD: test-password
    DATABASE_URL: "postgresql://app:test-password@db:5432/app_test"
  script:
    - pytest -q tests/integration --junitxml=integration-results.xml
  artifacts:
    when: always
    reports:
      junit: integration-results.xml
    expire_in: 7 days

smoke:
  image: python:3.12
  stage: smoke
  rules:
    - if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
  script:
    - echo "Replace with a smoke check against the deployed staging URL"

For this example, install the build package in the setup before invoking python -m build. Add a test database health check or wait step if your runner can start the database service more slowly than the test process. Store real credentials in protected CI variables or a secrets manager, never in the YAML file. The sample password is only a disposable test value.

The important design choices are the stage boundaries, explicit commands, report collection even on failure, and a deployment-boundary smoke check. Adapt syntax when using another CI platform. GitHub Actions, for example, defines jobs that can run concurrently or be sequenced with dependencies. GitHub Actions workflow and job model

4. Keep pull-request feedback useful

  • Run a check early if it is both fast and relevant. A lint error should not wait behind a long browser suite.
  • Do not let selection silently reduce protection. If path filters omit shared libraries, configuration, generated code, or cross-cutting changes, document a fallback that runs the broader suite.
  • Make merge-blocking checks stable. A noisy required check teaches developers to retry or ignore failures. Repair unstable tests or infrastructure before making the signal a hard gate.
  • Keep output diagnosable. Publish JUnit or equivalent results, preserve logs and relevant artifacts, and make failing test names visible in the CI interface.
  • Separate policy from mechanics. A job may run for information without blocking a merge. Define that policy explicitly and assign an owner to decide when it becomes a gate.
  • Cover untrusted contributions safely. Forked pull requests and external contributions may not have access to protected secrets. Design jobs so tests that do not need secrets can still run, and do not expose deployment credentials to untrusted code.

5. Speed up a slow pipeline methodically

  1. Measure the critical path. Record elapsed time by job and stage, queue time, setup time, test runtime, and the slowest suites. The longest job or dependency chain usually determines when the pipeline completes.
  2. Remove avoidable setup work. Pin dependencies, use dependency caches where valid, avoid reinstalling tools in every step of one job, and keep test fixtures and service startup focused. A cache is an optimization, not a correctness dependency: a cache miss must still produce a correct run.
  3. Run independent work concurrently. Separate jobs that need not wait for one another. GitLab stages run jobs within a stage in parallel, while its needs relationships can allow jobs to start as soon as their actual prerequisites finish. GitLab stages and needs
  4. Shard only a measured bottleneck. Confirm that the runner can distribute tests evenly, that each shard has isolated state, and that all shard reports are collected. Compare elapsed-time improvement with additional runner use and queue pressure.
  5. Reduce duplicated coverage thoughtfully. If several expensive tests prove the same behavior, keep the one that offers the clearest signal at the lowest sustainable level. Do not remove a broader test that protects a distinct integration boundary just to lower runtime.
  6. Re-measure after each change. A faster pipeline that misses failures, loses reports, or overloads runners is not an improvement.

GitLab supports splitting jobs with the parallel keyword. Its documented RSpec example is platform-specific; shard configuration and test distribution are runner and framework dependent. GitLab parallel keyword

rspec:
  stage: unit
  parallel: 4
  script:
    - bundle exec rspec --format progress

This only starts four copies of the job; it does not by itself partition the suite. Use the test runner’s supported partitioning mechanism and a stable shard index/total, or all workers may repeat the same tests. Verify the correct environment variables and syntax for your framework and runner. Preserve individual results and combine them or upload each shard’s report so a passing aggregate cannot conceal missing shards.

Parallelism can lower wall-clock time while increasing concurrent runner consumption, database load, API traffic, and contention. Shared mutable test data, rate limits, and resource constraints can also make parallel runs flaky. Increase concurrency gradually and observe both total elapsed time and infrastructure use.

6. Handle flaky tests without weakening confidence

A flaky test sometimes fails and later passes when retried. GitLab’s handbook warns that this behavior undermines trust in test results and wastes investigation time. Possible causes include a brittle test, unstable infrastructure, or an unstable application. GitLab Handbook: flaky tests

  1. Keep the first failure visible. Capture the failing test, attempt number, commit, runner, environment, logs, and relevant artifacts. A retry can help distinguish intermittent behavior, but it must not erase the initial result.
  2. Reproduce with controlled variation. Retry locally or in CI when useful, vary execution order or concurrency where possible, and compare environment and dependency versions.
  3. Classify the cause. Check whether the test relies on timing, shared state, an unstable external dependency, nondeterministic ordering, or a product defect.
  4. Assign an owner and a repair date. Track the test and the evidence. A quarantine is a temporary managed state, not a permanent way to remove coverage.
  5. Requalify before restoring the gate. Confirm stability under the conditions that originally exposed the issue, then return the test to its intended stage and monitor it.

Do not set a retry count as a substitute for fixing flakes. Retries may be useful to collect evidence or reduce disruption while a tracked repair is underway, but teams should know when a retry occurred and what the original result was. GitLab documents quarantine as a way to manage flaky tests until proven stable; keep a visible return path. GitLab pipeline triage guidance

7. Manage test reports, artifacts, and environments

  • Publish machine-readable test results. Use your platform’s test report format so failures appear in the run summary and can be compared over time.
  • Save artifacts on failure. Preserve logs, screenshots, browser traces, crash dumps, and generated reports for a bounded retention period appropriate to your investigation needs.
  • Make environments repeatable. Pin runtime and service versions where practical; define how databases and queues are initialized; avoid depending on mutable shared state.
  • Give each parallel worker isolated resources. Use unique schemas, temporary databases, namespaces, or test identifiers to prevent workers from changing one another’s data.
  • Protect sensitive data. Redact secrets and personal data from logs and artifacts; restrict access to deployment credentials and production-derived data.
  • Keep a scheduled broad run. If request pipelines use targeted test selection or omit expensive cases, retain a broader scheduled run to detect gaps, and route its failures to an owner.

8. Monitor suite health and make changes explicit

Review trends that help explain whether the suite remains useful: pipeline duration and queue time, failures by suite, repeated retries, quarantined tests and age, test-selection misses, report completeness, and runner consumption. A coverage percentage can show which code was executed, but does not establish that assertions would detect a defect.

When changing test stages, selection rules, retry behavior, blocking policy, or parallelism, document the reason, expected effect, owner, and review date. GitLab’s strategy emphasizes clear ownership, progressive testing, and stability; those are useful principles to evaluate locally. It does not establish a universal numeric target for CI duration, flake rate, retries, or coverage.

9. Troubleshooting common CI test failures

Symptom Likely cause What to do
Pipeline fails before tests start Dependency installation, build, configuration, or service startup failed Read the earliest failing log line, distinguish setup from test failure, pin or repair the dependency, and add an explicit service readiness check where needed.
Tests pass locally but fail in CI Runtime/version mismatch, missing environment variable, timezone/locale difference, or hidden local state Compare runtime, dependencies, environment, and service versions. Run the same command in a clean environment and make required configuration explicit.
Tests fail only when parallelized Shared database records, fixed ports, shared files, order dependence, or external rate limits Isolate worker state and resources, remove ordering assumptions, and reduce concurrency until the conflict is understood.
Retry passes after a failure Flaky test, infrastructure instability, race, or nondeterministic product behavior Preserve the first failure, gather logs, classify the cause, assign an owner, and track repair. Do not label the change safe based only on the retry.
Pipeline is slow despite many workers Uneven shards, setup or queue bottlenecks, serial dependencies, or resource contention Measure per-shard duration and critical path, improve distribution or dependencies, and check whether added runner load offsets elapsed-time gains.
Test report is missing although tests ran Wrong report path/format, report not generated after failure, or artifact not uploaded Use the CI platform’s expected report syntax, write the report to the configured path, and upload it even when the job fails.
Change-aware test selection missed a regression Dependency graph or path rules did not capture an indirect effect Repair the mapping, add tests for the dependency boundary, and provide a broad scheduled or higher-risk run as a backstop.
Deployment smoke test fails while earlier suites pass Deployment configuration, environment integration, migration, or live dependency issue Keep promotion blocked when the smoke check protects a critical boundary; inspect deployed version and environment-specific logs before rerunning.

10. Visual checks for web interfaces in CI

For a web application, browser-based checks can complement behavioral tests when layout, rendering, or a critical page state is part of the risk. Keep such checks scoped: a visual artifact helps a reviewer inspect a page, but it does not replace assertions about application behavior or prove that all user journeys work. Capture stable preview URLs, control dynamic content where possible, and retain the artifact with the relevant build or test result.

Or skip the browser setup

If a CI workflow needs a rendered page artifact, a screenshot API can return one from a URL without your job managing a browser installation. ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo website and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace the example URL with a preview URL your CI environment can reach, store the API key as a CI secret, and use the returned file as an artifact. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots per month are free with no card, with paid plans starting at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

11. Keep cost and reliability in view

CI cost includes more than test runtime: runner minutes or compute, queue delays, service containers, storage for artifacts, and engineering time spent diagnosing unreliable failures all matter. Parallel jobs can reduce elapsed time but consume more resources concurrently. Caches can save setup time but should never be required for correctness. A longer scheduled suite may be cheaper in developer interruption than putting every expensive scenario on every change, provided its risk is understood and failures are acted on.

Set local service-level expectations based on how quickly contributors need feedback and what infrastructure is available; the cited guidance does not establish universal numeric thresholds. Review the pipeline when repository architecture or change risk changes, not only when a duration graph rises.

FAQ

Should every test block a merge?

No. Block on checks that are relevant, stable, and important enough to protect the merge. Keep informative or unstable checks visible, assign an owner, and decide deliberately when they are ready to become gates.

Is code coverage a good target for CI?

Coverage is one diagnostic signal. It does not tell you whether a test would fail when behavior is wrong, so evaluate assertions, failure history, and risk as well.

Should E2E tests run on every pull request?

Run a focused set on every request if its runtime and reliability make the feedback worthwhile. Otherwise, run the critical subset at a later gate and broader scenarios on selected changes or a schedule, with the tradeoff documented.

Can a passing retry be treated as a passing test?

Record both outcomes. A passing retry identifies intermittent behavior; it does not explain or resolve the first failure.

Sources and further reading