ScreenshotNeo

BlogEngineering

How to Use Control Charts for Performance Testing

Learn to choose a performance metric, establish control limits from historical test results, select the right chart, and investigate signals without confusing stability with success.

By the ScreenshotNeo team4 October 202611 min read

A control chart shows whether repeated performance-test measurements remain consistent with an established process or signal a change worth investigating. Choose a meaningful metric, collect comparable observations in time order, establish a representative baseline and control limits, then plot new results against those limits. A signal is a reason to investigate, not proof of a cause. Statistical stability also does not mean performance meets a service objective or engineering target.

This guide covers metric choice, baseline construction, chart selection, a runnable Python example, signal interpretation, operational pitfalls, and an option for capturing web pages as part of a test workflow.

1. Decide what performance question to answer

Start with a concrete question such as “Has the median API response time shifted?” or “Is throughput varying more than it did last month?” Pick one primary measure for each chart. Do not combine unlike units such as latency and CPU utilization on one ordinary univariate chart; use separate charts or an appropriate multivariate method.

Possible measures include response or read/write time, CPU time per operation, throughput, and latency. NIST’s NML performance-testing documentation defines examples including maximum and average read/write time, average CPU time, throughput, and message latency. These are examples from that test context, not a universal metric list. Maximum-time measurements can also be affected by clock resolution, so know the limits of your measurement system. See the [NIST NML performance measures](https://www.nist.gov/el/intelligent-systems-division-73500/nist-rml-application-programming-interface/nmlperf).

Define the observation before collecting data. It might be one test run’s p95 latency, a mean over a fixed subgroup of requests, or completed requests per second during a fixed interval. Keep that definition consistent. If one point represents a subgroup summary, record the subgroup size and how it was formed.

2. Make test observations comparable

A control chart is useful only when its ordered observations represent the process you intend to monitor. Keep the test procedure stable and record context that could affect results, such as:

  • Application version, build, and configuration.
  • Load profile, request mix, concurrency, and test duration.
  • Hardware, cloud region, instance type, and deployment topology.
  • Cache state, dataset size, and background jobs.
  • Instrumentation version, clock source, and collection method.

Do not reorder points by value: preserve their time or run order. If conditions intentionally change, annotate the chart and decide whether the new results belong to the same process. A throughput point from a two-minute test under one workload should not be treated as directly comparable to a ten-minute test under a different workload.

3. Establish a baseline in two phases

NIST describes a two-phase approach to control-chart use:

  1. Phase I — establish and review limits. Use historical observations to calculate initial limits. Investigate points outside the limits and other evidence of assignable causes. Correct data problems or document justified exclusions; do not remove inconvenient values without an explanation.
  2. Phase II — monitor the process. Carry the reviewed limits forward and add new comparable observations in order. Investigate signals against the established baseline.

If the system changes materially—for example, a new architecture or measurement method—build and document a new baseline when justified. Do not silently recalculate limits after an unfavorable run: doing so can hide a real regression. NIST’s [SPC phases guidance](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc41.htm) describes historical data, investigation, and later monitoring.

Control limits estimate process behavior. They are not specification limits, an SLO, or an acceptance threshold. A stable service can consistently miss its latency target; a service that usually meets the target can still be unstable. Plot or evaluate the engineering target separately from the control limits.

4. Choose a chart that matches the data

Data and question Chart family to consider What it helps monitor
Continuous measurements collected in rational subgroups X-bar chart, commonly paired with an R or S chart Subgroup mean and within-subgroup variation
Continuous individual observations without subgroups Individuals/moving-range or other moving charts Individual level and short-term variation
Small shifts in process mean matter CUSUM or EWMA More sensitivity to relatively small location shifts
Proportions or counts P/NP charts for binomial data; C/U charts for count data Defect proportions or event counts under the relevant sampling setup

These are selection cues, not automatic prescriptions. NIST’s [control-chart documentation](https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc3.htm) explains chart families and assumptions. Standard continuous-data charts may rely on approximate normality; performance measurements can be skewed or contain discrete values. Check the assumptions, collection design, and distribution before applying a formula. The NIST guide notes that CUSUM and EWMA charts were developed to detect small location shifts.

For many teams, an individuals/moving-range chart is a practical starting point when each test run produces just one value and there is no meaningful subgroup. If repeated requests within a run are naturally grouped, a subgroup chart may better represent the process. Avoid treating many requests from one run as independent observations if shared conditions make them correlated.

5. A runnable Python example for individual measurements

The following example reads one numeric value per line, calculates illustrative three-sigma individuals-chart limits from the first baseline observations using the average moving range, and prints points outside the limits. It assumes a stable baseline, approximately normal individual observations, and consecutive comparable measurements. It is a teaching implementation, not a substitute for validating the chart choice and assumptions for production use.

#!/usr/bin/env python3
"""Simple individuals chart limits from a baseline and signal scan."""
import argparse
import csv
import math
import sys

# For moving ranges of two, d2 = 1.128; individual limits use
# mean +/- 3 * (average moving range / d2).
D2_RANGE_2 = 1.128


def load_values(path):
    values = []
    with open(path, newline="", encoding="utf-8") as handle:
        for row_number, row in enumerate(csv.reader(handle), start=1):
            if not row or not row[0].strip():
                continue
            try:
                value = float(row[0])
            except ValueError as exc:
                raise ValueError(f"row {row_number}: expected a numeric first column") from exc
            if not math.isfinite(value):
                raise ValueError(f"row {row_number}: value must be finite")
            values.append(value)
    return values


def limits(values):
    if len(values) < 2:
        raise ValueError("need at least two baseline observations")
    mean = sum(values) / len(values)
    moving_ranges = [abs(b - a) for a, b in zip(values, values[1:])]
    average_mr = sum(moving_ranges) / len(moving_ranges)
    sigma_estimate = average_mr / D2_RANGE_2
    return mean, mean - 3 * sigma_estimate, mean + 3 * sigma_estimate


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("csv_file", help="CSV with one measurement in its first column")
    parser.add_argument("--baseline-count", type=int, required=True,
                        help="number of initial rows used to establish limits")
    args = parser.parse_args()
    values = load_values(args.csv_file)
    if args.baseline_count > len(values):
        parser.error("baseline-count exceeds the number of observations")
    center, lcl, ucl = limits(values[:args.baseline_count])
    print(f"center={center:.6g} LCL={lcl:.6g} UCL={ucl:.6g}")
    for index, value in enumerate(values[args.baseline_count:], start=args.baseline_count + 1):
        if value < lcl or value > ucl:
            print(f"signal candidate at observation {index}: {value:.6g}")


if __name__ == "__main__":
    try:
        main()
    except (OSError, ValueError) as error:
        print(f"error: {error}", file=sys.stderr)
        sys.exit(2)

Save as control_chart.py, place one observation per CSV row in measurements.csv, then run python3 control_chart.py measurements.csv --baseline-count 20. The first 20 values establish limits; later values are checked. For a production chart, also chart the moving ranges, review Phase I stability, consider run rules appropriate to your use case, and use a statistical package or validated implementation when the chart design calls for it.

6. Read signals without overclaiming

A point above the upper control limit or below the lower limit is a signal candidate. A nonrandom pattern can also indicate a process change even when every point remains within limits. NIST describes a stable process as having points within limits and a random pattern. When a signal appears, check the deployment, workload, infrastructure, data, instrumentation, and test procedure around that time. The chart points to when behavior changed; it does not identify why.

Signal rules involve a false-alarm tradeoff. NIST gives an illustrative Shewhart X-bar example: under a normal unchanged process, the chance of a point beyond three-sigma limits is 0.0027, corresponding to an average run length of about 371 points before a false alarm. This is conditional on that example’s assumptions; it is not a universal rate for every performance chart. Additional run rules alter both detection and false-alarm behavior. See the [NIST discussion of X-bar average run length](https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc32.htm).

7. Put the chart into a repeatable workflow

  1. Write down the operational question and primary metric.
  2. Specify exactly what each plotted point represents and preserve observation order.
  3. Make the test repeatable; record workload, environment, version, and measurement context.
  4. Collect a historical baseline and review it in Phase I for data issues, signals, and assignable causes.
  5. Select a chart family that fits the data type, subgroup design, and shift size of concern.
  6. Fix reviewed limits for Phase II monitoring and plot each new comparable result.
  7. Investigate signals, record findings and corrective actions, and compare separately against the engineering target.
  8. Re-establish limits only when a material process change justifies a new baseline; record the reason and effective date.

8. Performance, reliability, and cost considerations

Control-chart calculations are usually small compared with running the performance test; the operational cost is dominated by test execution, infrastructure, and the effort to preserve comparable conditions. Pick a sampling cadence that can detect the changes you care about without making test runs contend with production workloads or with each other.

  • Measurement noise: verify clocks, units, missing values, outliers caused by collection faults, and instrumentation overhead. Do not automatically discard a slow observation; establish whether it reflects the service or the measurement system.
  • Dependence: neighboring observations may share cache, host, or workload conditions. Strong autocorrelation affects how traditional limits and false-alarm expectations should be interpreted.
  • Changing process: deployments and planned workload changes should be annotated. Separate regimes when they represent materially different processes.
  • Multiple metrics and alerts: monitoring many charts or adding many run rules increases the chance of at least one alert. Define ownership and an investigation path.
  • Cost control: schedule tests at a useful cadence, cap workload duration, and retain raw observations plus context so investigations do not require rerunning expensive tests.

9. Capture a page in a performance-test workflow

Some test workflows need a screenshot artifact alongside timing results—for example, to preserve the rendered page state associated with a run. Keep capture work separate from the timing measurement if browser rendering would affect the metric being measured. Use the same target, viewport, and capture settings when comparing artifacts over time.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Cookie banners are accepted and removed before the shot, along with known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing status. An MCP server lets AI agents use screenshot, page-info, and PDF-capture tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for free.

10. Troubleshooting common chart problems

Symptom Likely cause What to do
Limits are extremely wide or narrow Baseline mixes different workloads, deployments, or environments, or the measurement system changed. Check time order and context, identify distinct regimes, validate instrumentation, then establish documented limits from comparable data.
Frequent alerts with no apparent service change Autocorrelation, non-normal data, an unsuitable chart, or too many added run rules. Check assumptions and dependence; match chart to data and sampling design; review false-alarm expectations before adding rules.
A stable chart still misses the SLO Control limits describe process behavior, not acceptability. Compare measurements to the SLO separately and improve the process; do not move limits to make the target appear met.
There is no signal despite an obvious gradual regression A Shewhart chart may be less sensitive to small shifts, or sampling is too sparse. Consider EWMA or CUSUM, and choose a justified cadence that captures relevant changes.
One test run produces a misleading point Conditions differ, the run is incomplete, or a measurement fault occurred. Inspect run metadata and raw data. Correct only confirmed data errors and preserve an audit note; otherwise investigate it as a real observation.
Python script reports invalid numeric input Header text or a nonnumeric value is in the first column. Remove the header or adapt the reader to skip a named header; ensure the first column contains finite numeric values.
Limits cannot be calculated Fewer than two baseline values are available, or the baseline has no variation. Collect sufficient comparable observations and inspect whether rounded values or measurement resolution conceal variation. A zero moving range should be treated as a measurement/design issue, not as proof of perfect performance.

11. FAQ

How many baseline observations do I need?

There is no universal count that fits every metric and chart. Use enough representative historical data to estimate process behavior and examine Phase I stability; the example’s 20-point default in its invocation is only an illustration, not a recommended minimum.

Should I chart averages or percentiles?

Chart the statistic that answers the operational question and define it consistently. Tail latency often matters, but percentile estimates can have different sampling properties from individual observations; choose a chart and sampling design suited to that statistic.

Can a control chart prove a release caused a regression?

No. A signal identifies a change in observed behavior, not its cause. Compare timing with deployment and environment records, then use controlled investigation to establish causality.

Do I need a paid tool?

No. The method can be implemented with code and a plotting library or a statistical package. The important work is sound measurement, a defensible baseline, chart assumptions, and a response process.

References