Continuous Integration Requirements for Automated Testing
A practical CI checklist for builds, layered tests, security checks, merge gates, reporting, and reliable feedback—without treating one pipeline as universal.
A CI pipeline should automatically build and test changes as they enter the shared repository workflow, then show useful results to the people reviewing them. A practical baseline is a reproducible build, fast unit checks, integration tests for important boundaries, broader tests for critical behavior, appropriate security scans, and clear rules for which failures block a merge or release.
There is no universal test matrix, required operating-system list, coverage percentage, or rule that every test must run on every pull request. Choose checks based on the application’s risks, architecture, supported environments, test cost, and team policy. GitHub describes CI as frequently integrating changes and using automated build and test checks; GitLab’s published pipeline tiers are an example of one team’s policy, not an industry standard. GitHub’s CI overview · GitLab’s testing strategy
1. What CI should require
Use this as a starting checklist, then document the project-specific choices.
- Build: Compile or package the code using the same declared dependencies and build steps used by the project.
- Fast automated checks: Run formatting or lint checks and focused unit tests early, so common failures are quick to find.
- Interaction checks: Run integration tests for important connections between modules, services, databases, queues, or APIs.
- Critical behavior: Cover key feature or system behavior and the end-to-end journeys where a regression would matter most.
- Risk-based security checks: Consider source, dependency, secret, infrastructure, and container scans as relevant to the application. Add runtime or API security testing when it fits the exposure and architecture.
- Reviewable results: Surface job status and useful test reports in the pull or merge request. Keep enough output to diagnose failures.
- Explicit gates: State which checks block a merge, deployment, or release, and who owns a failing or flaky check.
- Maintained suites: Review runtime, redundant coverage, and reliability. Repair or remove tests that cannot provide dependable signal for the gate they serve.
Coverage reports can help identify untested areas, but the reviewed guidance does not establish a universal minimum coverage percentage. Set a threshold only when the team can explain what it protects and how exceptions are handled.
2. Choose test layers and when to run them
Different test layers provide different kinds of confidence. A reasonable pipeline starts with cheap, focused checks and expands to slower or broader checks where their added confidence justifies the time.
| Layer | What it checks | Typical placement | Gate decision |
|---|---|---|---|
| Static checks and build | Formatting, lint rules, type checks where used, and whether the project builds | On each relevant change | Usually a merge gate when stable and required by the project |
| Unit | Small components in isolation | Early on each change | Commonly blocking because feedback is fast; choose based on project policy |
| Integration | Interactions across module or service boundaries | On each change when practical; otherwise in a later required stage | Block when the covered boundary is important and results are dependable |
| Feature or system | Important behavior across larger parts of the application | On changes that affect the feature, or in a broader pipeline stage | Set a gate according to risk and runtime |
| End-to-end | Critical journeys through a running system | Prioritized smoke coverage on changes; broader suites later or on a schedule when appropriate | Reserve blocking status for suites with useful, reliable signal |
| Security and quality scans | Potential weaknesses in code, dependencies, secrets, infrastructure, images, or runtime behavior | Choose triggers and scope based on risk, platform, and feedback cost | Define policy for findings, exceptions, and remediation |
Do not copy a test-tier schedule as a universal rule. GitLab’s documented strategy, for example, runs unit checks across its merge-request tiers and expands to broader integration, feature, and end-to-end coverage at later tiers. Its particular staging and production smoke-test policies describe GitLab’s own practice. Another project may need a different schedule.
Separate independent jobs so they can run concurrently; make dependent jobs wait for the outputs they need. If the full suite is too slow for every change, keep a small, risk-focused set in the required path and run broader suites in a later pipeline stage or scheduled workflow. Make that tradeoff visible to maintainers rather than letting important tests silently disappear.
3. Triggers, runners, and reproducibility
Automate checks at the points where they can inform a decision. Common triggers include pushes and pull or merge requests; scheduled and manually or externally triggered workflows can cover recurring checks or operational needs. Keep workflow definitions in version control so changes to CI are reviewable alongside code changes.
Choose a runner environment that matches the project’s needs. Hosted runners reduce the work of maintaining runner machines; self-hosted runners can suit special hardware, network access, or environment requirements, but the team must operate them. Test multiple operating systems or runtime versions when the product promises that support or the risk warrants it—not simply because a matrix is available.
- Pin or otherwise deliberately manage tool and dependency versions so a passing run can be reproduced.
- Declare required services and setup steps rather than relying on hidden state on a runner.
- Keep credentials in the CI platform’s secret mechanism; avoid printing them in logs or embedding them in repository files.
- Use job dependencies for artifacts and setup that downstream checks actually need.
- Review concurrency, timeouts, and cancellation behavior so obsolete runs do not consume capacity unnecessarily.
For platform-specific workflow syntax and event behavior, use the GitHub Actions workflow documentation. GitHub documents both hosted and self-hosted runners and matrix jobs; select them to meet an explicit support or infrastructure need.
4. Example: a GitHub Actions baseline
This illustrative workflow checks a Node.js project on pushes and pull requests. Replace the Node version, package manager commands, and test scripts with those declared by your project. It is a starting point, not a required universal matrix.
name: CI
on:
push:
branches: [main]
pull_request:
permissions:
contents: read
jobs:
checks:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
cache: npm
- run: npm ci
- run: npm run lint --if-present
- run: npm run build
- run: npm test -- --ci
Use action versions and runtime versions consistent with your organization’s maintenance and security policy. Add separate jobs for integration, end-to-end, or security checks when they need distinct services, credentials, environments, or reporting. Configure branch protection or the repository’s equivalent to require the intended status checks; a workflow that runs but is not part of the merge policy is not a merge gate.
Test commands are project-specific. For example, some test runners do not accept --ci; use the command supported by the project. If the project needs a database or other service, declare it in the job and wait for readiness before running integration tests.
5. Reporting, blocking rules, and ownership
Make the pipeline answer three questions for each change: what ran, what failed, and what decision follows. Show pass or fail status in the review workflow and publish test reports or other supported artifacts when they help reviewers diagnose failures. GitHub documents test results in pull requests; GitLab documents unit test reports, coverage, and other report types. Report availability and configuration depend on the platform and project setup.
For every required check, define:
- Scope: What behavior or risk does it cover?
- Gate: Does failure block merge, deployment, or release?
- Owner: Who fixes the test, environment, or code when it fails?
- Exception: How is a temporary bypass approved, recorded, and removed?
- Reliability: How are intermittent failures investigated and tracked?
GitLab states in its own testing strategy: “If a test can’t reliably block a merge, deployment, or release, it shouldn’t exist. Fix it or delete it.” Treat that as GitLab’s stated principle, not a universal standard. The practical lesson is to give each blocking test an owner and a maintenance path. If a suite is flaky, diagnose the source—test race, shared state, external dependency, timing assumption, or unstable environment—before weakening the gate.
6. Security checks: select for risk
Security scanning is not one switch that covers every failure mode. Depending on the application, CI may scan source code, infrastructure definitions, secrets, dependencies, and container images. Runtime testing, API security testing, or fuzzing can find behavior-dependent issues that repository scans will not.
Choose checks based on what the application exposes, its languages and deployment model, organizational policy, and platform support. Define how findings are triaged, which severity or policy violations block a release, and how accepted exceptions expire or are reviewed. Do not assume a scanner runs on every branch or merge request just because the platform offers it: configuration and product tiers can affect behavior. GitLab’s documentation, for example, distinguishes branch pipeline behavior from merge-request scanning setup. Verify the settings for the actual repository.
7. Visual regression checks for web applications
For a web application where layout regressions matter, browser screenshots can complement functional tests. They do not replace assertions about behavior or accessibility. Use a small set of stable, high-value pages or components, control viewport and rendering conditions, and review changes that may be intentional. Screenshot capture in CI is useful only when the baseline, comparison policy, and update process are clear.
A direct browser setup can use Playwright in a Node.js project. Install Playwright and its browser in the project’s normal dependency setup, then run a script such as this against a locally running application:
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
await page.goto('http://127.0.0.1:3000/', { waitUntil: 'networkidle' });
await page.screenshot({ path: 'artifacts/home.png', fullPage: true });
} finally {
await browser.close();
}
In CI, start the application before this script and wait for its readiness endpoint. Upload the generated artifact using the CI platform’s artifact mechanism if reviewers need to inspect it. For repeatable comparisons, keep browser version, fonts, viewport, device scale factor, test data, and animations controlled. Avoid making a remote third-party page the only required visual test: its content and availability can change independently of your code.
Or skip the browser setup
For a screenshot capture step, ScreenshotNeo can return an image from one API request. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
Store the API key as a CI secret, never in workflow source or logs. Replace the example URL with a page you are authorized to capture. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. It also supports options such as full-page capture, selectors, viewport presets, custom waits, and image formats. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. Learn about ScreenshotNeo, then sign up for 1,000 free screenshots a month with no card.
8. Performance, reliability, and cost
CI cost includes runner time, concurrency limits, infrastructure operation, and the engineering time spent diagnosing failures. The source guidance reviewed here does not provide a universal provider price or a benchmark for pipeline duration, so measure your own workflow before setting budgets.
- Keep early feedback focused: Put build, lint, and fast tests before expensive suites where dependencies allow.
- Parallelize independent work: Split jobs when this improves feedback enough to justify runner and setup overhead.
- Use broader checks deliberately: Move expensive suites to later required stages or scheduled workflows only when the risk and merge policy support that choice.
- Watch queue time and runtime separately: A short job waiting for a runner is a capacity issue; a slow job executing is a suite or environment issue.
- Reduce nondeterminism: Control external dependencies and test data, and make service readiness explicit.
- Review runner choice: Compare operating-system and hardware needs, repository integration, reports, parallelism, secrets handling, and the work of maintaining runners.
Do not improve a duration metric by silently dropping important coverage. Record which assurance moved, why, and where it will run instead. Revisit the decision when the system or its risk changes.
9. Troubleshooting common CI failures
| Symptom | Likely cause | Practical fix |
|---|---|---|
| Works locally, fails on runner | Different runtime, dependency resolution, operating system, environment variable, or undeclared local state | Use the declared lockfile and runtime, print non-secret environment diagnostics, and make setup steps explicit. |
| Install step fails intermittently | Registry or network instability, rate limits, or mutable dependencies | Check the package manager error and service status, use lockfiles, and configure retries only for transient failures. |
| Integration test cannot connect to service | Service was not started, is not ready, or uses the wrong host/port in the runner network | Declare the service, add a readiness check, and use the network address appropriate to the CI environment. |
| Test passes alone but fails in suite | Shared state, ordering assumptions, resource collisions, or leaked processes | Isolate test data and resources, clean up in teardown, and make parallel tests use unique identifiers. |
| End-to-end test times out | Application did not start, readiness was assumed, the page waits on persistent network traffic, or the timeout is unsuitable | Check startup logs, wait on a concrete readiness condition, and choose a wait strategy that matches the app. |
| Workflow runs but does not block merging | The status check is not required by branch protection or equivalent repository policy | Configure the intended job status as a required check and verify the rule on a test change. |
| Security report is missing | Scanner is not enabled for that event, configuration is incomplete, or the feature/report is unavailable in the project tier | Check the platform’s current configuration and event-specific documentation; do not infer execution from product availability. |
| Visual screenshots differ unexpectedly | Browser, fonts, viewport, scale, animation, dynamic content, or remote page changed | Pin and control rendering inputs, disable or stabilize dynamic content, and review the diff before changing a baseline. |
| Pipeline is slow or queued | Long serial dependencies, oversized suites, limited runner capacity, or costly environment setup | Inspect per-job queue and runtime, parallelize independent jobs, cache only safe reusable dependencies, and reassess suite placement. |
10. A rollout checklist
- List the build, test, and security checks the project already relies on.
- Identify critical behavior and boundaries that lack automated coverage.
- Put reproducible build and fast checks on the normal change path.
- Add broader checks according to risk, dependencies, and feedback cost.
- Publish useful results where reviewers can find them.
- Configure explicit merge and release gates, with owners and exception handling.
- Track flaky tests, runtime, and queue time; fix underlying causes.
- Review the pipeline when architecture, exposure, supported environments, or release policy changes.
FAQ
Should CI run unit and integration tests on every pull request?
Run unit tests on each change when they are fast and dependable. Run integration tests on each change when the feedback and risk justify their cost; otherwise place them in a clearly defined later required stage or scheduled check.
Does every project need end-to-end tests?
No universal rule requires them. They can protect critical user journeys, but their value depends on system risk, reliability, and maintenance cost.
What test coverage percentage should CI enforce?
The cited guidance does not define a universal minimum. Use coverage as a signal and set a project-specific threshold only when it supports a clear quality policy.
Should every CI failure block a merge?
Define the gate intentionally. Required checks should provide dependable signal; flaky checks need an owner and a repair plan so the team does not normalize ignored failures.


