ScreenshotNeo

BlogEngineering

How Test Intelligence Can Find Patterns in Test Data

Learn how test intelligence reveals recurring failures, flaky tests, regressions, platform-specific issues, and coverage gaps from test history.

By the ScreenshotNeo team4 October 20269 min read

Test intelligence finds patterns by collecting comparable test results over time, then grouping and comparing them by test, build, code change, browser or device, environment, requirement, and failure signature. The patterns can show which failures recur, when they began, whether they are intermittent or platform-specific, and where intended coverage is missing. They help prioritize investigation; correlation alone does not prove a root cause.

A single isolated run rarely reveals a reliable trend. Preserve stable test identities and useful run context, accumulate published results, and drill into the evidence behind any pattern before deciding what caused it.

What patterns can test intelligence reveal?

Pattern What it may indicate What to inspect next
A test starts failing after a particular build or change A possible regression, environment shift, or dependency change Compare the last passing and first failing runs, their code changes, logs, and environment.
The same test passes and fails on the same code across repeated runs Flaky or nondeterministic behavior Compare retries, timing, ordering, shared state, network dependencies, and other run context.
Many tests fail with a similar signature A shared dependency, setup issue, or common failure mode Group by error, file, fixture, service, or environment; validate against underlying traces and logs.
A test fails only on a browser, device, or operating system A configuration-specific issue or difference in supported behavior Compare the same test and build across platforms, then reproduce on the affected configuration.
A requirement or changed area has little associated test evidence A potential coverage gap Trace requirements and code changes to tests and results; confirm the measure and its limits.

These are investigation leads, not diagnoses. A failure that begins after a release may be caused by the release, a changed test environment, test data, or an unrelated dependency.

Build a useful test history

  1. Publish results consistently. Send outcomes from CI runs to a history or analytics view. Microsoft describes Azure Pipelines Test Analytics as using published results accrued over time; trends need a comparable series, not just one run. Microsoft Learn: Test Analytics.
  2. Keep test identity stable. Use consistent names or identifiers across runs. Renaming or duplicating tests can fragment their history and obscure recurring failures.
  3. Retain context. Include build or commit, branch, time, browser/device, operating system, environment, test file, requirement link, and relevant failure output where available.
  4. Check comparability. If the runner image, browser version, test data, or pipeline changed, record that context so a trend is not mistaken for a code-only effect.
  5. Choose a time window that fits the question. A short window helps investigate a recent change; a longer history can reveal recurring or seasonal patterns. Note gaps in data and changes in test inventory.

Keep the raw run records accessible. Summary metrics and grouping are useful for finding a signal, while the underlying run, logs, traces, and configuration are what let a developer examine it.

Analyze the data step by step

  1. Start with the question. Decide whether you are looking for a new regression, repeat failures, flaky behavior, a platform-specific problem, or missing test evidence.
  2. Review totals and trends. Look at pass rate, failure count, duration, top failing tests, and changes by day or build. Check whether the test population changed; a pass-rate shift can reflect added or removed tests.
  3. Find concentration. Group failures by test, file, error signature, platform, environment, or other context available in the data. A cluster can reveal a shared lead for investigation.
  4. Open the individual histories. Follow a failing test across builds or days. Identify the last known passing outcome and first failing one, then compare the associated run details.
  5. Compare dimensions. Compare the same test and build across browsers, devices, or environments. Keep other variables as consistent as possible so the comparison is meaningful.
  6. Connect results to intent. Link tests and outcomes to requirements or changed code where the workflow supports it. Treat a coverage indicator as the tool’s defined measure, not proof that the product is adequately tested.
  7. Validate and record. Reproduce the suspected problem, inspect logs and traces, test the likely cause, and document what the evidence supports. Update the test or pipeline if the investigation finds a defect or reliability issue.

Distinguish a regression from a flaky test

Ask: “How can I tell whether a test failure is a regression or a flaky test?” Compare outcomes for the same test under repeated executions and across the surrounding builds.

  • Evidence favoring a regression: the test passes before a specific change, fails consistently afterward under comparable conditions, and the failure can be reproduced or tied to changed behavior.
  • Evidence favoring flakiness: the same test alternates between pass and fail on the same code and comparable setup, or failures correlate with timing, order, shared state, or unstable external services.
  • Evidence still needed: logs, traces, test data, environment details, and a reproduction. One failed attempt cannot classify the failure by itself.

Retries can help expose inconsistency, but a passing retry does not erase the first failure or prove the test is harmless. Record both outcomes and investigate why they differ. A 2022 survey of 335 professional developers and testers reported concern that flaky tests reduce trust in test results; the sample describes the study participants, not a universal prevalence estimate. Survey on how test flakiness affects developers.

Compare platforms, changes, and requirements

Browser and device comparisons

Compare the same test across configurations with the build and test data held steady when possible. If failure is limited to one browser or device, inspect configuration-specific behavior and verify the result with a reproduction. Sauce Labs documents Insights views for test histories and platform-specific comparisons. Sauce Labs Insights documentation.

Change-oriented investigation

For “Did failures begin after a particular change?”, locate the last passing and first failing runs, then review commits and pipeline or environment changes between them. The temporal link narrows the search; it does not establish causation. Validate by reproducing, bisecting where practical, or testing the suspected change under controlled conditions.

Requirement traceability and test gaps

For “Which requirements or changes have not been covered by tests?”, connect requirements and changed areas to test cases and execution outcomes. A missing link is a prompt to verify the intended coverage and whether the test ran, not proof that no relevant test exists. Qase describes analytics across cases, defects, runs, results, plans, and requirements, including requirement traceability integrations; these are vendor-documented capabilities. Qase Test Intelligence. J. Rott’s paper discusses analyses and visualizations that support software testing teams. Test Intelligence: analyses and visualizations.

Choosing an analysis view or tool

Evaluate a tool against the question your team needs to answer, rather than treating a single score or AI label as a complete diagnosis.

What to compare Questions to ask
Analysis question Can it show trends, flakiness, platform differences, requirement links, or grouped failures relevant to your workflow?
Dimensions and filters Can you filter by build, test, file, change, browser/device, environment, and requirement?
History and context How far back do results go, and can you open the underlying run and its evidence?
Workflow connections Can it receive CI results and connect to the issue or requirement systems your team uses?
Automated suggestions Can clustering or root-cause suggestions be checked against logs, traces, changes, and reproduction?

Azure Pipelines documents pass-rate and failure summaries, grouping, test history, and trend analysis for its published results. Microsoft Learn. Sauce Labs documents history and platform-oriented views. Sauce Labs documentation. Qase documents analytics and requirement traceability capabilities. Qase product page. TestMu AI describes flaky detection, failure clustering, root-cause analysis, and forecasting as vendor capabilities; treat generated results as suggestions to validate, not independent accuracy guarantees. TestMu AI Test Intelligence. The available evidence does not establish an objectively best test-intelligence vendor or an independent accuracy comparison.

Practical implementation checklist

  • Publish results from each relevant CI run to a consistent history.
  • Preserve stable test IDs and record commit, build, platform, and environment context.
  • Track test inventory and pipeline changes alongside pass/fail trends.
  • Group failures to find concentration, then inspect the individual run records.
  • Compare repeated outcomes before marking a test flaky.
  • Compare platform configurations with other variables controlled where possible.
  • Link tests to requirements or changed areas and understand the coverage measure used.
  • Keep a human review step for automated clusters and root-cause suggestions.
  • Record the evidence, suspected cause, reproduction, and resulting action.

Common mistakes and troubleshooting

Symptom Likely cause Fix
No trend or history appears Results are not being published, the time window is too short, or the selected project/filter excludes them. Confirm the CI publish step succeeded, widen the time range, and check project and branch filters.
A test appears as several unrelated tests Its name or identity changes between runs. Use a stable identifier and map renamed tests where the analytics system supports it.
Failure rate changes sharply after a pipeline update Runner, browser, dependency, test data, or test inventory changed along with code. Compare run context and isolate the changed variables before attributing the shift to an application change.
A test fails once and passes on retry Possible intermittent behavior, timing issue, or environmental variation. Retain both outcomes, inspect logs and ordering, and repeat under comparable conditions. Do not silently classify it as fixed.
Many tests fail at once A shared setup, service, dependency, or environment may be involved. Group by signature and environment, inspect representative traces, and verify the shared dependency.
A platform comparison is inconclusive Build, test data, browser version, or execution environment differs between groups. Align these variables or annotate the differences before comparing outcomes.
Coverage looks high but an important behavior is untested The displayed measure may not represent requirement or behavior coverage. Inspect the metric definition and map critical requirements or changed code to explicit tests.
An AI grouping or root-cause suggestion seems wrong Similar messages may have different causes, or the available context may be incomplete. Check raw failures, traces, logs, code changes, and reproduction; correct labels or groupings where possible.

Performance, reliability, and cost considerations

Test analytics adds work around result publishing, retention, and investigation. Keep publication reliable and monitor whether runs are missing; incomplete history can make apparent trends misleading. Use stable IDs and context fields to avoid fragmented histories. Large result volumes may require filtering by time, build, or test before drilling down, depending on the tool and dataset.

Interpret trend metrics in context: test additions, removals, retries, changed environments, and altered test selection can all affect counts or rates. Preserve enough run detail to investigate important failures, and follow your organization’s retention and access policies for logs and traces.

Costs and limits are product-specific and can change. Compare current pricing, retention, ingestion limits, integrations, and the effort needed to maintain result publishing for the tools under consideration. No independent vendor cost comparison is established by the research for this article. Human investigation remains necessary when the evidence is ambiguous.

Or skip the browser setup

If your investigation needs a rendered page capture as supporting evidence, ScreenshotNeo can return an image or PDF with one GET request. It is a website screenshot API and MCP server for developers, made by Yorker Media. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

FAQ

Can test intelligence identify the exact root cause automatically?

It can surface patterns and, in some products, vendor-described suggestions. Confirm the cause with the run evidence and reproduction.

How much history do I need?

Enough comparable published results to cover the behavior or period you are investigating. A single run cannot establish a trend, and sparse or changing data should be treated cautiously.

Does a high pass rate mean the software is well tested?

No. It describes outcomes for the tests that ran. It does not by itself establish that important requirements or behaviors were covered.

Should every flaky test be quarantined?

Not automatically. First determine impact and gather evidence. Quarantine policy depends on your team’s risk tolerance and CI workflow; keep the failure visible and assign follow-up.