ScreenshotNeo

BlogGuides

How to Measure Test Automation Maturity

Measure test automation maturity with evidence, balanced metrics, and a repeatable improvement cycle—not a single coverage percentage.

By the ScreenshotNeo team4 October 202610 min read

Measure test automation maturity as an evidence-backed profile of practices, outcomes, and sustainability. A test count, automation percentage, code coverage figure, or pass rate cannot represent maturity on its own. Define the system and decision in scope, inspect artifacts and operational data, track a small balanced set of measures, and use the findings to choose the next improvement.

This guide gives you a repeatable assessment method, practical metrics, a scoring rubric you can adapt, and a sample implementation for collecting and reporting test-run measures. The rubric is a local decision aid, not a universal or scientifically validated maturity scale.

1. Define the assessment scope and decision

Write down what you are assessing before choosing metrics. The scope might be one team, a product, a portfolio, or the organization’s test automation practice. Record the time period, the people who will use the result, and the decision it should inform.

Assessment detail Example
Scope Checkout service and its delivery team
Period The previous three release cycles
Audience Engineering and product leads
Decision Choose whether to invest first in test data, CI feedback speed, or critical-path coverage

Useful assessment questions include: Are critical customer journeys exercised reliably? Can a developer get an actionable result quickly? Do tests detect important defects before release? Is maintenance work sustainable? Avoid comparing teams without considering differences in product risk, architecture, test mix, and release context. ISO/IEC 33063 describes process assessment using indicators and objective evidence, with subsets selected for the assessment context; it is a software testing process assessment model, not a universal automation scorecard: ISO/IEC 33063:2015.

2. Turn goals into measures with Goal-Question-Metric

For each goal, ask a question whose answer would change a decision, then select a measure that can answer it. For example:

Goal Question Measure
Improve confidence in critical journeys Which agreed high-risk journeys are exercised by reliable automation? Risk-weighted journey coverage, with the inventory and denominator recorded
Shorten developer feedback How long from a code change until a useful result is available? Change-to-actionable-result time and suite execution time trend
Reduce wasted triage How often do test or environment failures appear as product regressions? Flaky-test rate and false-positive failure rate
Reduce escaped defects Which serious post-release defects had a test opportunity that was missed or ineffective? Severity-aware escaped-defect trend with reviewed links to test gaps
Keep automation sustainable Is maintaining the suite displacing useful work? Repair and update effort, duplicate or obsolete cases, and maintenance backlog

Define each measure precisely: formula, numerator, denominator, exclusions, collection window, system of record, and owner. A4Q’s Selenium Tester Syllabus version 3.0 discusses Goal-Question-Metric and examples such as execution time and reliability: A4Q Selenium Tester Syllabus.

3. Assess practices using evidence

Look at the conditions that make automation useful and sustainable, not only at the test results. A literature review synthesized 26 practices across 13 areas from 81 primary studies. Areas include strategy, resourcing, professional competence, tool selection, environments, testability, test data, scripts, test oracles, execution-result analysis, and technology adoption. Treat that list as a source of questions, not as a requirement that every team adopt every practice in the same way. The review also reports that only six practices had formal empirical evaluations of positive effects on maturity improvement, so present recommended practices as evidence-informed guidance with limits, not proven universal prescriptions: Wang et al., “Improving test automation maturity: A multivocal literature review”.

Inspect relevant artifacts, then cross-check interview statements against them:

  • Strategy, risk records, and definitions of critical paths
  • Test code, review practices, test oracles, and ownership
  • Environment setup, test data preparation, and repeatability
  • CI configuration, reports, failure classification, and triage records
  • Defect records, maintenance work, skills plans, and tool decisions

Ask the people who build, maintain, and use the tests how failures are diagnosed, what work gets delayed by suite maintenance, and whether results lead to timely decisions. Record the evidence and the rationale for each rating so another assessor can challenge or repeat the assessment.

A practical local rubric

Level Meaning Evidence example
Ad hoc Practice is inconsistent or depends on individuals Critical paths are not agreed, and failures are handled differently across the team
Repeatable Practice exists for common work but has gaps CI runs a maintained suite, but test data or triage ownership is inconsistent
Managed with evidence Ownership and outcomes are measured against explicit goals Risk inventory, failure categories, and reliability trends are reviewed
Regularly improved Teams use trends and learning to adjust practice Recurring escape and maintenance causes have owners and are reassessed

Score relevant practices individually and report the profile, evidence, and gaps. Avoid averaging every rating into one precise-looking number: a strong CI pipeline cannot compensate for no test ownership or poor coverage of the system’s highest risks.

4. Track a balanced set of outcome and suite-health measures

Choose a compact dashboard that supports decisions. Use trends over time and, where data supports it, meaningful percentiles. Keep definitions stable between reporting periods.

Measure Useful definition What it can tell you Common trap
Risk-weighted automation coverage Agreed critical requirements, operational paths, or journeys exercised by automation divided by the agreed inventory; state weighting and denominator Whether automation reaches important behavior Counting tests without knowing which risks they cover
Code coverage Coverage reported by the relevant instrumentation for a stated build and test set Which code paths may be untested Treating a high percentage as proof that assertions detect faults
Reliability Flaky-test and false-positive rates, with product, test, and environment failures distinguished How much confidence to place in a result Counting reruns as clean passes without tracking the initial failure
Feedback speed Suite duration and time from change to actionable result Whether tests support a useful development loop Optimizing runtime while moving important checks too late
Escaped defects Post-release defects linked to severity and reviewed for missed or inadequate test opportunities Whether important risks are being missed Blaming automation for every production defect
Maintenance and sustainability Effort to repair or update tests, obsolete or duplicate cases, and maintenance backlog Whether the suite remains affordable to change Counting maintenance only as failure rather than necessary upkeep
Test effectiveness Defects detected, risks validated, and evidence that results trigger timely decisions Whether automation changes outcomes Rewarding activity or test volume without useful signal

Microsoft recommends tracking pass rate, defect escape rate, flakiness, execution-time trends, and code coverage, while treating coverage as a signal rather than a target: Build confidence in Azure workloads with effective testing practices. UK Home Office guidance also names defect density, execution time, unreliable-test percentage, defect leakage across levels, and automation coverage: Test pyramid.

5. Decide how much automation coverage is enough

There is no useful universal percentage. First agree on a risk or requirement inventory: critical user journeys, key business rules, failure modes, and operational paths. Then report what proportion is covered, how strong and reliable the checks are, and what remains manual or untested. Code coverage can help locate untested paths, but it does not show whether tests would detect incorrect behavior.

Use coverage as a diagnostic signal alongside reliability, defect escape, execution time, and maintenance effort. Revisit the denominator when the product or risk profile changes. Do not compare a team’s percentage to another team’s without aligning scope and definitions.

6. Compare test layers and team practices in context

Compare options on risk coverage, signal quality, feedback cost, defect outcomes, and operational fit: skills, environment stability, data availability, integrations, and ownership. The UK Home Office test-pyramid guidance recommends emphasizing lower-level tests where practical and reserving end-to-end automation for critical, high-risk flows because end-to-end tests are more complex, fragile, and time-consuming. This is a heuristic, not a required test ratio for every product. Measure whether the chosen mix gives your team useful feedback at sustainable cost.

7. Example: calculate measures from test-run records

This runnable Python example reads a CSV of test executions and reports pass rate, an initial flaky-failure signal, and mean run duration. It is a starting point for suite-health reporting, not a full maturity assessment. The input needs columns test_id, outcome (pass or fail), duration_seconds, and attempt (1 for the first run, 2 for a retry). Add a stable run_id and your own failure classification in production so retries and environment failures are interpreted correctly.

import csv
from collections import defaultdict
from statistics import mean

by_test = defaultdict(list)
run_durations = []

with open("test-runs.csv", newline="", encoding="utf-8") as file:
    for row in csv.DictReader(file):
        outcome = row["outcome"].strip().lower()
        duration = float(row["duration_seconds"])
        attempt = int(row["attempt"])
        if outcome not in {"pass", "fail"} or duration < 0 or attempt < 1:
            raise ValueError(f"Invalid test-run row: {row}")
        by_test[row["test_id"]].append((attempt, outcome))
        run_durations.append(duration)

first_attempts = [
    outcomes[0][1]
    for outcomes in by_test.values()
    if outcomes
]
first_failures = sum(result == "fail" for result in first_attempts)
flaky_candidates = sum(
    any(attempt == 1 and result == "fail" for attempt, result in outcomes)
    and any(attempt > 1 and result == "pass" for attempt, result in outcomes)
    for outcomes in by_test.values()
)

if not first_attempts:
    raise SystemExit("No test executions found")

print(f"Tests with a first attempt: {len(first_attempts)}")
print(f"First-attempt pass rate: {(len(first_attempts) - first_failures) / len(first_attempts):.1%}")
print(f"Fail-then-pass flaky candidates: {flaky_candidates}")
print(f"Mean execution seconds: {mean(run_durations):.2f}")

Save this as measure.py, create a UTF-8 file named test-runs.csv with a header row and records, then run python measure.py. The example assumes each test’s first row is its first attempt; sort or group by run ID and attempt in real data. A retry pass is a candidate flaky signal, not proof of flakiness: classify product regressions, infrastructure failures, and test defects separately. Report trends and sample sizes, and do not use pass rate as a standalone maturity score.

8. Turn assessment results into an improvement cycle

  1. Choose a few gaps with the highest risk or recurring cost.
  2. Assign an owner, a concrete action, and an observable outcome.
  3. Record the baseline and collection window before changing the system.
  4. Make the change, then observe the same measures over a comparable period.
  5. Reassess the evidence and update priorities as risks and architecture change.

Examples include stabilizing a flaky critical journey, improving repeatability of test data, adding a check for a recurring escaped defect, or shortening CI feedback by putting checks at an appropriate layer. Microsoft recommends regular review and maintenance of flaky, duplicate, and obsolete tests. A 2020 practitioner survey reported responses from 151 practitioners across more than 101 organizations and 25 countries; 85% said their teams had sufficient automation expertise, while 47% reported a lack of guidelines for designing and executing automated tests. These are findings from that survey, not current universal benchmarks: Software Test Automation Maturity: A Survey of the State of the Practice.

Or skip the browser setup

If part of your maturity assessment needs consistent website screenshots—for example, checking a rendered page or capturing evidence for a review—ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted like a visitor, and known consent banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers say which page verdict applied and whether it was billed.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up for 1,000 free screenshots a month—no card required.

Troubleshooting an assessment

Problem Likely cause Fix
Coverage rises but important defects still escape The denominator counts tests or lines rather than agreed risks, or checks do not assert meaningful outcomes Review escaped defects against missed test opportunities; define critical journeys and strengthen relevant assertions
Pass rate is high but developers ignore failures Retries hide flaky first failures, or product, test, and environment failures are mixed together Preserve first-attempt results, classify failure causes, and trend false positives separately
Suite runtime improves while confidence drops Important checks may have been removed or moved too late Review risk coverage and defect outcomes with runtime; validate the new test-layer allocation
Teams cannot agree on their maturity rating Scope, evidence, or rubric language is ambiguous Record concrete evidence beside each rating, clarify the assessment scope, and have a second reviewer challenge assumptions
Dashboard metrics change meaning over time Definitions, exclusions, or data sources changed without being recorded Version metric definitions and annotate collection or denominator changes
One score hides severe gaps An average combines unrelated dimensions and masks weak practices Report a profile by practice area and show risk, evidence, and actions for each material gap

Frequently asked questions

Is there a standard test automation maturity score?

ISO/IEC 33063 is a process assessment model for software testing, not a universal test automation score. Define a rubric that fits the assessment scope and retain its evidence and rationale.

Should we measure manual effort saved?

It can answer a useful organizational question, but define what counts as displaced or avoided manual work and compare it with automation design, maintenance, and triage effort.

How often should we reassess?

Reassess on a cadence that gives changes time to affect the measures, and after meaningful changes to product risk, architecture, or delivery practice. Keep the same definitions when comparing periods.

Can one team use this rubric to rank another?

Use it to structure a discussion, not to rank teams without accounting for scope, risk, architecture, and release conditions. A profile with evidence is more useful than a leaderboard.