ScreenshotNeo

BlogEngineering

How to Use Test Analytics to Improve QA

Turn test results into better QA decisions: track meaningful trends, investigate flaky tests, prioritize risk, and close the loop on improvements.

By the ScreenshotNeo team4 October 202612 min read

To use test analytics to improve QA, treat results as a feedback loop: collect comparable test runs, look for meaningful trends and repeated failures, investigate them in the context of product risk, make a targeted change, then check later runs to see whether it helped. Analytics are useful when they change a testing or release decision; a pass-rate chart or coverage percentage alone does not improve quality.

This guide covers how to set up that loop, which metrics to start with, how to investigate signals, and how to report findings to developers, operations, and stakeholders. It also explains how screenshot artifacts can help diagnose failures in browser-based tests.

1. Start with a decision, not a dashboard

Before choosing charts or adding metrics, write down the question the team needs to answer and the decision it may lead to. Examples:

  • Did a recent code change cause a sustained decline in passing tests?
  • Which intermittent failures consume the most triage time?
  • Do critical user journeys have meaningful tests at the appropriate layers?
  • Is the suite becoming too slow to give developers useful feedback?
  • Are defects escaping testing into production, and what test change could address them?

Each question implies a different analysis. A failure investigation needs test-level history and run context. A release-readiness decision needs an agreed scope, current results, known defects, and an explicit account of remaining risk. A coverage review needs a map from critical paths to tests, not just a repository-wide percentage.

Microsoft’s testing guidance recommends measuring quality, tracking defects and coverage, and feeding findings back into development. Use that as a continuous improvement loop rather than a scorecard: observe, investigate, act, and review the effect. Microsoft Azure Well-Architected testing guidance.

2. Collect results that can be compared

Analytics depend on stable, traceable test results. For each test execution, retain enough context to compare one run with another and to reproduce or assign a failure.

Field Why it matters
Stable test identity Allows history to follow the same test even if its display name changes.
Outcome and failure details Distinguishes assertion failures, setup errors, timeouts, and other causes.
Timestamp and duration Supports trend analysis and helps identify slowdowns or time-related patterns.
Build, commit, or release Connects a result to the code and delivery event that produced it.
Environment and dependencies Helps reveal whether failures follow a browser, operating system, service, or configuration.
Artifacts and links Logs, traces, screenshots, and work items provide the context needed for follow-up.

Publish results consistently from CI and release runs. If teams rename tests, change result formats, or omit failed runs, trends can become misleading. Keep historical data tied to a stable identifier where possible, and record the scope of each run: which tests ran, where they ran, and whether the run was complete.

Azure Pipelines Test Analytics is one pipeline-specific example. Its documentation says its insights are based on test results published for a build or release pipeline and describes pass-rate summaries, failing tests, trends, grouping, and test-level drill-down. It documents a 14-day default view; that is a product default, not a universal analysis window. Check current product scope and configuration in the Azure Pipelines Test Analytics documentation.

3. Choose a small set of useful metrics

Begin with a few measures that connect directly to the decisions above. Define each one before comparing teams or setting a release gate. State the numerator, denominator, test scope, and time window, and clarify whether it measures individual test executions or whole runs.

Metric What it can signal How to use it carefully
Test pass rate A sustained decline may indicate a regression, environment problem, or unstable suite. Compare equivalent scopes and inspect which tests changed outcome. A run-level pass rate and a test-execution pass rate are not interchangeable.
Flakiness rate Intermittent failures can waste triage time and reduce trust in results. Define what counts as intermittent and over what observation window. Separate test behavior from infrastructure and dependency failures where possible.
Execution-time trend A slower suite can delay feedback and increase delivery friction. Compare similar environments and suite scopes; investigate the tests or setup responsible for the increase.
Defect escape rate An increase in defects first found in production may point to gaps in test design, scope, or environment coverage. Define what counts as an escaped defect and use a consistent reporting period and severity scope.
Code coverage Low coverage in a critical area can reveal paths that need more attention. Use coverage as a signal for investigation, not a quality target by itself. Exercise critical behavior and assertions, not just lines.

Microsoft’s Azure guidance discusses these quality measures but does not prescribe one universal formula or target threshold. Set definitions that match the test system and product risk. Resist adding metrics unless someone can explain which decision a change in that metric should trigger. A large dashboard of numbers without owners or actions can obscure the quality signals that matter. Microsoft Azure Well-Architected testing guidance.

A single run can tell you that a test failed; it usually cannot tell you whether the failure is new, intermittent, or part of a longer trend. Compare results across enough runs to see a pattern, while keeping the window short enough that changes remain traceable.

  1. Choose a time range that includes the relevant release cadence and enough comparable executions.
  2. Compare like with like: the same suite, environment, and result definition wherever practical.
  3. Mark meaningful events such as code changes, infrastructure changes, test migrations, or dependency updates.
  4. Inspect daily or per-build trends for a shift, then drill into tests and failure details to explain it.
  5. Revisit the same measure after a change to see whether the signal moved as expected.

Do not assume a 14-day view or any other preset is right for every team. A daily deployment pipeline may need a different window than a weekly release cycle. For low-frequency suites, the same calendar period may contain too few runs to support a reliable conclusion.

5. Investigate failures before labeling them

When a pass rate falls or a failure repeats, identify the tests and contexts behind the change. Grouping can expose a shared cause; for example, many failures in one browser or service may point to environment health rather than independent product regressions.

When failures appear after a change

  • Compare failing tests with the last known good build and the commits or files changed since then.
  • Inspect the assertion, stack trace, logs, and test setup. Confirm that the test reached the intended behavior.
  • Check whether the failure reproduces in the same environment and whether related tests changed too.
  • Assign the issue to an owner and record the evidence, impact, and next step.

When a test passes and fails intermittently

Microsoft defines a flaky test as one that inconsistently passes or fails without code changes. Compare repeated executions of the same test and inspect timing, shared test data, concurrency, isolation, infrastructure, and dependencies. A rerun can help diagnose nondeterminism, but a later pass does not prove the original failure was harmless.

Google’s John Micco reported in 2016 that 1.5% of test runs in Google’s corpus produced a flaky result, almost 16% of tests had some flakiness, and about 84% of observed pass-to-fail transitions in its post-submit testing involved a flaky test. These are historical figures from Google’s own test infrastructure, not industry benchmarks or expected rates for another team. Micco also described reruns and quarantine as mitigation techniques while warning that quarantine can hide real bugs such as race conditions. Any quarantined test should have an owner, a remediation issue, and a review condition. John Micco, “Flaky Tests at Google and How We Mitigate Them”.

For browser-based tests, keep visual evidence

A screenshot can show the rendered state at failure time, while logs and traces help explain how the test reached it. Capture evidence at a useful point: for example, when an assertion fails or after a defined navigation step. Avoid treating an image as a substitute for the failure message, DOM state, or trace; it is one piece of context. Store artifacts with access controls and a retention period that fit the data they may contain.

6. Prioritize coverage by product risk

Coverage is most useful when it helps locate a meaningful gap. Map critical user journeys and failure consequences to tests, then decide whether a gap warrants a new test and at which layer. A focused test for a payment or account recovery flow may reduce more risk than increasing coverage across low-impact code.

  • List business-critical workflows and the most consequential failure modes.
  • Map each workflow to existing unit, integration, and end-to-end coverage.
  • Check whether tests validate important outcomes and failure paths, not just execution.
  • Review escaped defects and add regression coverage at the layer that would have caught them most reliably.
  • Account for maintenance cost: a test should provide a clear signal worth keeping.

Microsoft explicitly advises treating coverage as a signal rather than a target. High coverage does not guarantee that tests assert the right behavior, and low coverage does not have the same risk in every area. Microsoft Azure Well-Architected testing guidance.

7. Turn findings into changes and review the result

Analytics only improves QA when findings lead to owned work and the team checks whether that work helped. Match the intervention to the signal.

Finding Possible action Follow-up signal
Failures cluster after a code change Fix the regression or correct the test if its expectation is wrong. Results on later builds and the affected behavior.
Intermittent failures share test data or timing Isolate data, remove ordering assumptions, or make synchronization explicit. Intermittent outcomes for those tests over subsequent runs.
Critical path has a meaningful gap Add focused coverage at an appropriate test layer. Whether the test detects relevant regressions without adding excessive feedback time.
Suite duration keeps rising Profile slow setup and tests; move suitable long, lower-frequency checks to scheduled runs while keeping fast critical checks in the change workflow. Time to useful feedback and the defects each suite still catches.
Escaped defect reveals a test gap Reproduce the defect, add regression coverage, and run it in the environment where the issue occurred. Regression test stability and whether the same class of defect is caught earlier.
Duplicate, obsolete, or low-signal tests Review, repair, consolidate, or remove tests with an owner and reason. Signal quality, suite duration, and retained risk coverage.

Microsoft recommends scheduled maintenance for flaky, duplicate, or obsolete tests and describes nightly full-suite runs in pre-production as one way to find flaky tests and regressions. Apply scheduling to suit the product’s risk and feedback needs; a scheduled run should complement required checks, not make an important release decision invisible until later. Microsoft Azure Well-Architected testing guidance.

8. Report findings for the people making decisions

Use the same underlying evidence to create views suited to each audience. Microsoft’s guidance gives examples of developers using flakiness and coverage, operations using pass rate and execution time, and business stakeholders using defect escape trends.

  • Developer view: actionable failure queue, recent changes, test owner, failure context, and linked artifacts.
  • Operations view: pipeline health, pass-rate trend, execution time, environment issues, and release risks.
  • Stakeholder view: escaped defects, critical journey coverage, release readiness, and remaining risk in plain language.

A release report can summarize the release scope, test runs, significant defects, coverage of critical paths, readiness decision, remaining risk, and next priorities. Keep failures traceable to a test case, owner, or work item so recurring problems can be followed through. Do not compress away uncertainty: say what was tested, what was not, and what evidence supports the decision.

9. Common problems and fixes

Problem Likely cause Practical fix
Pass rate drops but the team cannot identify why Results lack stable test identity, build context, or failure detail; the compared runs may have different scope. Standardize published results, retain build and environment metadata, and compare equivalent suites before grouping failures.
Every intermittent failure is dismissed after a rerun passes Rerun success is being treated as proof that the first failure was harmless. Track repeated outcomes, inspect execution context, and keep unresolved tests visible with an owner and review date.
Coverage rises but escaped defects do not fall The metric rewards exercised code without showing whether important behavior is asserted. Map tests to critical journeys and defect classes; add or improve assertions where risk warrants it.
Dashboard has many charts but no resulting work Metrics lack a decision, owner, threshold, or review cadence. Remove measures without a clear action and assign an owner and follow-up signal to each retained one.
Test history appears to reset after renames or migrations Analytics treats changed identifiers or result formats as new tests. Preserve stable test IDs where supported and document migration points before interpreting trend breaks.
Failure groups point to many unrelated causes Grouping dimensions are too broad or test names conceal different behaviors. Group by a meaningful combination such as test identity, failure signature, environment, and build, then drill into individual executions.
Artifacts are missing or hard to use Capture is not connected to failure handling, or storage and retention are unclear. Attach screenshots, logs, or traces at defined failure points and link them to the test result with clear access and retention rules.

10. A practical implementation checklist

  1. Choose one quality question that matters to a current engineering or release decision.
  2. Confirm that CI publishes complete, consistently identified results with timestamps, duration, build, environment, and failure context.
  3. Define a small set of metrics, including their formulas, scopes, and time windows.
  4. Set a review cadence that fits the release cycle and assign owners for investigations.
  5. Use trends to find a specific test, failure pattern, coverage gap, or time cost.
  6. Prioritize the finding against product risk and the maintenance cost of a proposed test change.
  7. Record the intervention and the later signal that will show whether it helped.
  8. Review unresolved flaky tests and quarantines so they do not become invisible permanent exceptions.
  9. Summarize readiness and remaining risk in language appropriate to each audience.

11. Or skip the browser setup

If a browser test needs a screenshot artifact, you can capture the page with your own browser tooling and attach the resulting file to the test report. Or use ScreenshotNeo, a website screenshot API and MCP server for developers. Its API returns an image or PDF from one GET request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

12. FAQ

How much testing is enough to qualify a software release?

There is no universal test count or coverage percentage that qualifies every release. Define the critical behavior and risk for the release, identify which tests provide evidence for those areas, and state any remaining gaps and known failures in the readiness decision.

Should every flaky test be quarantined?

No. Quarantine may reduce noise while an issue is investigated, but it can hide a product defect. Use it only with visible ownership, a remediation issue, and a condition for review or removal.

Can test analytics tell whether a failure is a product bug?

Analytics can reveal patterns and guide investigation, but a metric alone cannot establish the cause. Use failure details, code and environment changes, reproduction, and relevant artifacts to determine what happened.

Is code coverage a release quality score?

No. It shows which code was exercised according to a tool’s definition. It does not, by itself, show that tests checked the right behavior or that the most consequential paths are protected.

Sources