ScreenshotNeo

BlogEngineering

How Machine Learning Helps Detect Anomalies and Defects in Software Testing

Learn how ML can prioritize defect-prone code, flag unusual test executions, and identify flaky tests—plus practical Python examples and limits.

By the ScreenshotNeo team4 October 202612 min read

Machine learning helps software testing in three different ways: it can rank code units by defect risk using project history, flag unusual test executions when expected results are hard to specify, and predict tests likely to be flaky from their prior behavior. These outputs help teams decide where to investigate. None proves that a bug exists, that unusual behavior is incorrect, or that a failed test can be ignored.

The practical starting point is to define which decision you need to improve, collect evidence relevant to that decision, and verify model alerts with tests, requirements, or human review. A single model should not blur these three tasks together.

Task Question it answers Typical evidence What its output means
Defect prediction Which code units may deserve earlier review or more tests? Historical defect labels, code metrics, change history, ownership or complexity features A relative risk estimate based on the project data
Execution anomaly detection Which execution looks unusual when a complete expected result is unavailable? Inputs, outputs, traces, runtime metrics, state transitions A deviation from learned or specified normal behavior
Flaky-test detection Which tests may produce unstable outcomes under nominally unchanged conditions? Pass/fail history, rerun outcomes, timing, environment and dynamic features A likelihood or observation of test instability

A defect predictor is not an anomaly detector: one estimates defect risk from project history, while the other flags an unusual execution or result. A flakiness detector focuses on inconsistent test outcomes, which can be caused by timing, shared state, network dependencies, or environment variation and does not necessarily mean the product is defective.

2. How defect prediction prioritizes code

A team labels past software units according to whether they were associated with defects, extracts features from code and project history, then trains a classifier or ranking model. The resulting score can help prioritize review or test effort. It cannot confirm a defect in a specific file or commit.

Labels and features determine what a model can learn. If a project records only a subset of defects, uses inconsistent definitions, or has too few examples for important defect types, its predictions inherit those limitations. A systematic review of software defect prediction research reports concerns with commonly used datasets, including feature adequacy, validation, and label detail. See the [systematic literature review](https://research.monash.edu/en/publications/a-systematic-literature-review-on-software-defect-prediction-usin/) and the [business-driven mapping study](https://www.sciencedirect.com/science/article/pii/S0950584922002373).

Practical workflow

  1. Choose the unit to rank: file, component, change, or module. Keep it consistent with the action that follows.
  2. Define a defect label and a time window. For example, a code unit is positive if a qualifying defect is linked to it within a stated period after release.
  3. Build features available at decision time only. Avoid including information that would only be known after the defect was found.
  4. Split evaluation data by time or release, so the model is evaluated on later work rather than near-duplicate examples from the same period.
  5. Measure precision and recall at the review capacity the team actually has. A score is useful only if the ranked items fit the available review budget.
  6. Inspect false positives and missed defects with engineers; revisit labels, features, and thresholds as code and workflows change.

For rare defects, accuracy alone can be misleading: a model that predicts “no defect” for everything may appear accurate while finding none of the cases the team cares about. Review the confusion matrix and the trade-off between missed positives and investigation workload.

3. How anomaly detection can help when the oracle is incomplete

A test oracle decides whether observed behavior is correct. For some systems, writing a complete expected output for every input is difficult. Anomaly detection can learn patterns from prior inputs, outputs, or execution traces and flag a new run that differs substantially. The learned pattern describes observed behavior; it is not automatically the intended behavior.

Research has evaluated semi-supervised and unsupervised approaches over execution data for this purpose. One empirical comparison reported that semi-supervised learning outperformed Daikon in most evaluated systems, while Daikon did better in at least one. This is evidence about those systems and methods, not a universal ranking. See the [IEEE comparison](https://ieeexplore.ieee.org/document/8990211/).

Using anomaly alerts responsibly

  • Capture the inputs, output, relevant trace features, build version, and environment alongside each alert.
  • Compare alerts with requirements, domain invariants, or a more authoritative oracle before filing a product defect.
  • Separate expected variation—such as timestamps, generated identifiers, or floating-point tolerance—from meaningful changes.
  • Review the baseline when the product changes. A once-normal behavior can become invalid, and a new valid behavior can look anomalous.
  • Track false alerts and missed incidents; tune detection thresholds to the cost of reviewing each alert.

Where a visual web interface is under test, a screenshot can be one piece of output evidence for a visual comparison. It does not by itself establish whether the page is correct: dynamic content, responsive layout, consent dialogs, and timing can all affect what is captured.

4. How teams identify flaky tests

A flaky test can pass or fail for the same test and program version under conditions intended to remain constant. Historical outcomes and dynamic features can help predict which tests are likely to be flaky. Reruns provide stronger evidence of instability, but each rerun consumes test execution time.

Parry and colleagues evaluated CANNIER, a combination of machine-learning and rerun-based techniques, on 89,668 test cases across 30 Python projects. In that evaluation, the authors reported an order-of-magnitude reduction in rerun time while maintaining better detection performance than ML alone. Treat this as a result from that study’s projects and setting, not a guarantee for a different language, suite, or CI system. See the [Empirical Software Engineering paper](https://doi.org/10.1007/s10664-023-10307-w).

Useful signals and confirmation

  • Outcome history: pass/fail sequences and failure frequency across repeated runs.
  • Runtime features: duration, timeout proximity, retries, resource use, or order in the suite.
  • Environment: worker, operating system, dependency versions, concurrency, and external service availability.
  • Rerun evidence: repeat the same test under controlled conditions and retain the environment metadata.

A model estimate and a rerun confirmation are different evidence. If a test is marked “likely flaky,” preserve its failure and investigate; do not automatically discard the result. Reruns can help distinguish instability from a repeatable product failure, but they may not reproduce timing-sensitive or environment-specific behavior.

5. A runnable Python example

This example trains a simple defect-risk classifier on made-up historical component records and an Isolation Forest on execution metrics. It prints ranked risk and unusual executions. The data is deliberately synthetic so the script runs as-is; its scores are not meaningful for a real project. Install the dependencies with python -m pip install scikit-learn, save as testing_ml_demo.py, and run python testing_ml_demo.py.

from sklearn.ensemble import IsolationForest, RandomForestClassifier

# Synthetic historical components. Features represent values available
# before a release; defect_found is the later label.
components = [
    # churn, complexity, prior_defects, lines_changed, defect_found
    [2, 3, 0, 12, 0], [4, 5, 0, 20, 0], [3, 4, 1, 18, 0],
    [8, 9, 2, 75, 1], [7, 8, 3, 61, 1], [9, 7, 2, 90, 1],
    [1, 2, 0, 5, 0], [6, 6, 1, 48, 0], [10, 10, 4, 110, 1],
    [5, 5, 1, 35, 0], [8, 8, 3, 80, 1], [2, 3, 0, 10, 0],
]
X = [row[:4] for row in components]
y = [row[4] for row in components]
names = [f"component_{i + 1}" for i in range(len(X))]

risk_model = RandomForestClassifier(
    n_estimators=200, class_weight="balanced", random_state=7
)
risk_model.fit(X, y)

new_components = [[7, 8, 2, 70], [2, 3, 0, 11]]
new_names = ["checkout", "health_check"]
probabilities = risk_model.predict_proba(new_components)
defect_class_index = list(risk_model.classes_).index(1)
ranked = sorted(
    zip(new_names, probabilities[:, defect_class_index]),
    key=lambda item: item[1], reverse=True
)
print("Defect-risk ranking (priority signal, not a bug verdict):")
for name, score in ranked:
    print(f"  {name}: {score:.2f}")

# Execution features: duration_ms, output_size_bytes, error_count.
# The first observations are a small baseline; the last is unusual in this toy set.
runs = [
    [101, 5200, 0], [98, 5100, 0], [105, 5300, 0],
    [99, 5150, 0], [103, 5250, 0], [650, 1200, 4],
]
run_names = [f"run_{i + 1}" for i in range(len(runs))]
anomaly_model = IsolationForest(
    n_estimators=200, contamination="auto", random_state=7
)
anomaly_model.fit(runs[:-1])  # Fit only on the assumed normal baseline.
labels = anomaly_model.predict(runs)  # -1 means flagged; 1 means inlier.
scores = anomaly_model.decision_function(runs)  # Lower scores are more unusual.
print("Execution anomaly review queue:")
for name, label, score in zip(run_names, labels, scores):
    if label == -1:
        print(f"  {name}: flagged (score {score:.3f}); inspect the run")

Interpretation: the defect example learns only from its toy labels; a production model needs representative project history and time-aware validation. The anomaly example assumes the first five runs are a valid normal baseline; if that assumption is wrong, it can learn bad behavior as normal. Isolation Forest’s contamination setting can influence the decision threshold when a numeric estimate of outlier share is justified. Do not choose it merely to make the alert count look convenient. The classifier probability is a model score, not a calibrated probability unless calibration has been evaluated.

The scikit-learn [IsolationForest documentation](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.IsolationForest.html) describes its anomaly scoring behavior; the [RandomForestClassifier documentation](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html) documents classifier options.

6. Choosing and validating an approach

Decision axis Questions to answer
Purpose Are you ranking code risk, flagging unusual behavior, or finding unstable tests?
Evidence Do you have defect labels, traces, input/output observations, test history, or rerun outcomes?
Representativeness Do the examples reflect the current project, release process, test environment, and defect definitions?
Detection quality Which misses and false alerts matter? Report task-appropriate precision, recall, alert volume, and review effort.
Cost How much time is needed for instrumentation, feature collection, training, reruns, and alert analysis?
Change resilience Does performance hold across releases, suite changes, environments, and changing data distributions?
Human verification Can an engineer relate each alert to a trace, requirement, code change, or reproducible test?

Start with a baseline your team can explain, record the data and model version behind each result, and compare the model with the existing workflow. Keep a review path for alerts and a way to suppress known benign cases without erasing their history. Reevaluate after meaningful changes to code, test infrastructure, or data collection.

7. Testing systems that contain machine learning

Using ML to improve tests is distinct from testing an ML system itself. For a product that contains a model, teams may need to test data, model behavior, and supporting frameworks against properties such as correctness, robustness, and fairness. A survey of 138 research papers organizes ML testing by properties, system components, workflows, and application scenarios ([Zhang et al.](https://discovery.ucl.ac.uk/id/eprint/10099087/)).

A Microsoft Research industry study reports 87 survey responses and interviews with 7 senior practitioners. It identifies data collection, execution, and result analysis as major activities; analysis combines quantitative metrics with practitioner judgment ([study](https://www.microsoft.com/en-us/research/publication/testing-machine-learning-systems-in-industry-an-empirical-study/)). These findings are a reminder to evaluate the full workflow, not just a model’s headline metric.

8. Visual evidence for web tests

For a web application, an image capture can support visual regression review or provide an artifact alongside an anomalous browser run. It is useful when a changed layout or missing element is part of the symptom. It does not replace semantic assertions, accessibility checks, API tests, or a defined expected result. Keep viewport, device scale, wait conditions, and page state consistent so captures are comparable.

You can produce captures with a browser automation tool already in your stack, or use a screenshot endpoint as a repeatable artifact source. ScreenshotNeo is a website screenshot API and MCP server for developers. Its API can return PNG, JPEG, WebP, or PDF; its documented options include full-page capture, element selection, device and viewport settings, wait conditions, custom CSS and JavaScript, and request controls. See [ScreenshotNeo](https://screenshotneo.com) and its [API documentation](https://screenshotneo.com/docs/).

Or skip the browser setup

Make one GET request to capture a page as WebP:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page info, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These captures can help build visual evidence for a test investigation, while your test suite still decides what behavior is expected.

Sign up for ScreenshotNeo’s free plan and start with 1,000 screenshots a month, no card required.

9. Troubleshooting model-assisted testing

Symptom Likely cause What to do
Many high-risk components are harmless Labels are noisy, features correlate with size, or the alert threshold ignores review capacity Inspect false positives, evaluate precision at the team’s review budget, and improve labeling or features
The model misses newly introduced defect types Training data does not represent new behavior, architecture, or defect categories Collect new examples, evaluate by release and defect type, and retain ordinary testing and review
Most executions are flagged as anomalous The baseline is too small, includes mixed operating modes, or has a mismatched feature scale Segment normal modes, inspect distributions, collect a representative baseline, and retune only with validation
A flagged execution is actually valid Unusual behavior is being mistaken for incorrect behavior Check specifications or domain invariants; refine the oracle or document an accepted variation
A suspected flaky test passes on rerun Flakiness may depend on order, load, timing, worker, or external state Preserve the original environment metadata and repeat under controlled, varied conditions; passing once is not proof of stability
Reruns make CI too slow Rerunning every test multiplies execution cost Use risk-based selection, parallelize only where safe, and compare saved investigation time with added runtime
Performance drops after a release Data drift, changed test behavior, feature extraction changes, or leakage in earlier evaluation Revalidate on later releases, review feature and label pipelines, and retrain only after checking data quality
Image comparisons vary between runs Viewport, fonts, content, wait timing, animations, or remote data differ Control capture conditions, wait for stable page state, hide irrelevant dynamic regions, and compare like with like

10. Performance, reliability, and cost

  • Feature collection: tracing and historical joins have an ongoing storage and maintenance cost. Prefer features that are available consistently and at the point decisions are made.
  • Training: many project datasets are modest, but training cost is only one part of the system; validation, triage, and data repair often consume more effort.
  • Reruns: confirmation improves evidence but increases CI minutes and feedback time. Selective reruns trade coverage for speed.
  • Reliability: predictions can shift as teams, architecture, release cadence, and environments change. Monitor alert volume and quality over time, and keep a non-ML fallback test process.
  • Visual captures: network resources, page load time, and capture settings affect duration and image consistency. Capture only the artifacts needed for a decision and standardize conditions.
  • Business value: compare the time spent collecting data and reviewing alerts with the testing or review effort saved. The cited research does not establish a universal accuracy rate or cost saving across projects.

FAQ

Can machine learning identify bugs without a test oracle?

It can flag unusual behavior or rank risk without a complete expected output, but a requirement, invariant, or human investigation is still needed to determine whether behavior is wrong.

Does a flaky test mean the application is defective?

No. It describes outcome instability. The cause may be the test, environment, dependencies, or product, so investigate before assigning blame.

Should teams build one model for all three tasks?

Usually these tasks need different labels, evidence, and evaluation. Keep their data and success criteria explicit even if they share infrastructure.

Can the model replace code review or reruns?

No. It can prioritize attention, while review and controlled reruns provide different forms of evidence.