How to Run High-Performance Tests in a CI/CD Pipeline
Build repeatable CI performance checks with representative workloads, useful thresholds, and retained results. Includes a runnable k6 GitHub Actions workflow.
Run performance tests in CI/CD by defining a service goal, exercising a representative workload against a controlled environment, and making the job fail when pre-agreed latency, error-rate, throughput, or correctness criteria are missed. Keep fast checks close to code changes and schedule broader tests where they can run long enough to provide useful evidence. A passing run applies only to the workload and environment it tested.
This guide uses k6 thresholds and GitHub Actions as a concrete example. The same workflow applies with other CI systems: configure a bounded test, retain its output, and use the test process exit status as the gate.
1. Decide what the pipeline should protect
Start with a user-visible objective or service-level objective (SLO), such as keeping successful requests within an acceptable response time while maintaining a sufficiently low error rate. Avoid copying a sample threshold as if it were a universal standard. Set limits using your own service requirements and baseline measurements.
Choose the signal that answers the question you care about. In k6, the common built-in metrics include:
http_req_duration: request duration, typically assessed using percentiles such as p95.http_req_failed: the rate of failed HTTP requests.http_reqs: the number or rate of generated requests.- Checks: assertions that responses are correct, not merely fast.
Latency, error rate, throughput, and correctness are separate signals. For example, a test that generates many requests but accepts incorrect response bodies can report a misleading success. Define functional checks alongside performance thresholds.
2. Model a representative workload
Build the test around an important API path or user journey. Decide how much traffic to generate, how quickly to ramp it, how long to sustain it, and whether the test should include think time or multiple request types. A brief smoke test catches basic regressions; a staged test can show how the service behaves as load rises.
Document the differences between the test and production: data size, cache state, service dependencies, machine sizes, network path, and traffic distribution can all affect results. Keep the test bounded and use a dedicated staging or pre-release environment when a run could affect real users or shared systems.
3. Write a k6 test with checks and thresholds
Save the following as perf.js. It runs a short staged workload, checks status and response shape, and applies illustrative thresholds. Replace the sample URL, response field, load profile, and threshold values with ones that fit the service.
import http from 'k6/http';
import { check, sleep } from 'k6';
const baseUrl = __ENV.BASE_URL;
if (!baseUrl) {
throw new Error('Set BASE_URL to the API or staging origin');
}
export const options = {
stages: [
{ duration: '30s', target: 5 },
{ duration: '1m', target: 5 },
{ duration: '30s', target: 0 },
],
thresholds: {
// Examples only: replace with service-specific objectives.
http_req_failed: ['rate<0.01'],
http_req_duration: ['p(95)<200'],
checks: ['rate>0.99'],
},
};
export default function () {
const response = http.get(`${baseUrl}/health`, {
tags: { name: 'health' },
});
check(response, {
'status is 200': (r) => r.status === 200,
'body contains healthy status': (r) => r.body.includes('healthy'),
});
sleep(1);
}
The stages example ramps to five virtual users, holds for a minute, then ramps down. This is only a demonstration of a bounded profile, not a claim about the traffic your service should receive. Add the important API paths and realistic request data before relying on results.
k6 thresholds turn metrics into pass/fail criteria. When a threshold is breached, the k6 CLI exits with a non-zero status, allowing CI to fail the job. A threshold should be tight enough to catch a meaningful regression but grounded in observed service behavior.
4. Run it locally before adding the CI gate
Install k6 using the official installation instructions, then run the script with a safe target:
BASE_URL=https://staging.example.com k6 run perf.js
Inspect the summary before making it a merge or deployment gate. Confirm that the test exercised the intended path, checks match real responses, thresholds reflect the service objective, and the target can safely receive the configured load.
5. Add a bounded GitHub Actions workflow
Save this as .github/workflows/performance.yml. Grafana Labs documents official k6 GitHub Actions; this example uses the k6 action to run the script and uploads the text summary even when the test fails. Review the current k6 CI documentation and pin dependencies according to your repository’s supply-chain practices before using a workflow in production.
name: Performance check
on:
workflow_dispatch:
pull_request:
paths:
- 'perf.js'
- '.github/workflows/performance.yml'
schedule:
- cron: '17 3 * * *'
jobs:
k6:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run bounded k6 test
uses: grafana/k6-action@v0.3.1
with:
filename: perf.js
env:
BASE_URL: ${{ vars.STAGING_BASE_URL }}
- name: Upload k6 summary
if: always()
uses: actions/upload-artifact@v4
with:
name: k6-summary
path: summary.json
if-no-files-found: ignore
Set STAGING_BASE_URL as a repository variable or adapt the environment configuration to your deployment setup. Do not put credentials in the script or commit secrets. If the tested endpoints require authentication, pass credentials through your CI secret store and use them in the script without printing them.
The sample runs on manual dispatch, on changes to the test or workflow, and on a schedule. A team can add a pull request gate after confirming the test is stable and its runtime is acceptable. Longer scenarios are often more useful as scheduled or pre-release jobs. Grafana’s automation guide notes that load tests commonly take 3 to 15 minutes or more; actual time depends on the scenario and environment. See Grafana’s automated performance testing guide.
6. Choose where each test belongs
| Pipeline stage | Suitable test | Purpose |
|---|---|---|
| Pull request or frequent integration | Small smoke or focused test | Fast feedback on critical paths and obvious regressions. |
| Nightly or scheduled | Longer staged workload | Observe behavior over a longer period and exercise more paths. |
| Pre-release | Broader, production-like scenario | Build confidence under planned load shapes before release. |
Keep high-volume tests away from uncontrolled production traffic. If production testing is necessary, coordinate it, bound the load, and define an abort condition. A quick CI test improves feedback time; it does not replace a broader investigation when production risk warrants one.
7. Keep results useful for diagnosis
- Retain the test summary and relevant time-series output as CI artifacts.
- Record the commit, target environment, workload parameters, and tool version with each run.
- Compare against a baseline where your workflow supports a meaningful comparison.
- Route failures to the team responsible for the service and include enough output to locate the failed threshold or check.
- Version-control scripts and service objectives where practical; revisit them when traffic patterns or SLOs change.
When a run unexpectedly fails, inspect the result before changing the threshold. A failure can reflect an application regression, unstable test data, environment drift, a changed dependency, or a workload that no longer represents use. Record the cause and any threshold adjustment so the gate remains informative.
8. Optional: automate visual checks for web pages
If a web journey also needs a visual rendering check, a screenshot can provide evidence of what the page looked like during a run. Keep that separate from load-test metrics: a screenshot does not measure throughput or establish behavior at higher concurrency. For browser-based capture, use a controlled target and avoid launching many browser instances in parallel unless the runner can support them.
Or skip the browser setup
For a rendered page snapshot, ScreenshotNeo provides a website screenshot API and MCP server. Add a separate snapshot step where visual evidence helps diagnose a page-rendering issue; it does not replace k6 performance thresholds. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie and consent banners are accepted and removed before capture; newsletter popups and chat widgets are removed too.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses report the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots per month, no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| CI exits non-zero after a run | A threshold or check failed; this is the intended gate behavior. | Read the k6 summary, identify the failed metric or assertion, and compare with the baseline before changing limits. |
| Requests fail before reaching the service | The target URL or network access is wrong, or the environment is unavailable. | Check BASE_URL, runner network access, DNS, and target readiness. Avoid silently allowing a failed setup to pass. |
| Checks fail despite low latency | The response status or body differs from the script’s assumptions. | Inspect a safe response sample, correct the assertion, or fix the application behavior. Do not treat speed as correctness. |
| Results vary substantially between runs | Shared runners, noisy neighbors, unstable test data, cold caches, or environment drift may be affecting the measurement. | Stabilize the environment and workload, record run context, and use a baseline or repeated runs before enforcing a narrow threshold. |
| Pull requests take too long | The scenario is too large for frequent feedback. | Keep a focused smoke test in the pull request workflow and move broader load profiles to scheduled or pre-release runs. |
| Test appears successful but user behavior is broken | The script does not cover the affected journey or assert the relevant result. | Add a representative request path and correctness checks; review coverage when product behavior changes. |
| Test traffic affects other users or shared systems | The workload exceeds what the target environment can safely handle. | Stop the run, lower or bound the load, use an isolated environment, and coordinate any production test. |
Performance, reliability, and cost considerations
- Feedback cost: longer tests delay a pipeline and consume runner time. Put only the test duration needed for the decision at that stage.
- Measurement reliability: a CI runner and staging environment may differ from production. Record those differences and interpret results in that context.
- Load safety: define virtual-user and duration bounds; never let a test scale without an explicit limit.
- Gate reliability: flaky tests teach developers to ignore failures. Investigate instability and separate environment failures from product regressions where possible.
- Service cost: load can consume compute and third-party API quotas. Use test accounts or controlled data and account for dependency limits.
- Tool cost: this example uses k6 CLI in CI; execution time, runner resources, and any hosted testing service should be considered separately. The cited materials do not establish a universal cost or benchmark.
Frequently asked questions
Should every pull request run a load test?
Not necessarily. Run a bounded test on every change when its duration and stability make the feedback worthwhile; schedule heavier tests or use a pre-release stage when they need more time or a more representative environment.
Can a green pipeline prove production performance?
No. It shows that the defined workload met its criteria in the tested environment. Production traffic, dependencies, and infrastructure can differ.
What should block a deployment?
Block on failures tied to a service objective and a representative, trusted test. The exact signals and limits are service-specific; include correctness so a fast but incorrect response cannot pass.
Does ScreenshotNeo replace a performance test?
No. It captures rendered pages; use a load-testing tool such as k6 to measure request behavior under load.


