ScreenshotNeo

BlogEngineering

How Machine Learning Is Used in Test Automation

Learn how machine learning can generate tests, propose expected results, improve test suites, and help analyze outcomes—and how to evaluate its limits.

By the ScreenshotNeo team4 October 202612 min read

Machine learning (ML) can help automate software testing by proposing test inputs and executable tests, suggesting expected results, improving test suites, and helping analyze test outcomes. It does not make a generated test correct by itself: developers still need to check that the test expresses intended behavior, then measure whether it finds faults and remains useful over time.

There are two related but distinct problems. One is using ML to test ordinary software. The other is testing software that itself uses AI or ML, where expected outputs may be difficult to specify or may vary between runs. The second problem makes trustworthy test oracles—ways to decide whether a result is correct—especially important.

Where machine learning fits in test automation

A 2023 systematic mapping study reviewed 124 publications on ML-assisted automated test generation. It describes ML being used to generate inputs or expected-result oracles, and to improve the effectiveness or efficiency of existing generation methods. Its sample covers research literature; it is not an estimate of industry adoption. Read the mapping study.

1. Generate inputs, steps, and test cases

A model can propose values, sequences of actions, or whole tests. The target might be a unit, a GUI, a complete system, a performance scenario, or a combinatorial configuration. The useful output depends on the target: a plausible unit-test input is not the same thing as a realistic sequence of user actions.

Microsoft Research describes transformer models trained on developers’ code to generate tests intended to be accurate and readable. The project page lists C# in Visual Studio and Java in VSCode as supported contexts. These are stated capabilities for those contexts, not a guarantee that generated tests will work for any repository. See Microsoft Research’s AI for Testing project.

2. Propose expected results and assertions

Generating inputs answers “what should we try?” Oracle generation addresses “what should happen?” A model may propose an assertion, expected value, or pass/fail verdict. This can help when a test framework needs an expected result, but an assertion that matches the current implementation can still be wrong if the implementation does not match the requirement.

TOGA is a research example of neural test-oracle generation integrated with EvoSuite. Its authors report 96% overall accuracy on a held-out test dataset and 57 real-world bugs found in large Java programs, including 30 that other automated methods in that evaluation did not find. Those are results from the study’s evaluated data and setup, not a general success rate for test-generation products. Read the TOGA paper summary.

3. Improve an existing test suite

ML may help prioritize tests, tune generation strategies, adapt to system-specific feedback, or filter similar tests. For example, a prioritization model might put tests likely to expose a regression earlier in a run. A filtering model might reduce redundant cases. In either case, measure what was gained and what was lost: a smaller or faster suite is not an improvement if it stops detecting important failures.

4. Analyze execution results

Models can assist with classifying failures, grouping similar results, or evaluating execution outcomes. ETSI describes AI-assisted test generation, test-data creation, execution-result evaluation, and continuous monitoring as areas of activity. Its working-group overview describes work in progress; consult the linked standards for detailed requirements rather than treating the overview as a conformance checklist. See the ETSI MTS AI working group.

Which testing tasks can use ML?

The mapping study describes research across several testing targets. Technique choice should follow the task and the available evidence: models suited to generating unit-test inputs may not suit GUI exploration or performance testing.

Testing target Possible ML contribution What to review
Unit Propose method inputs, tests, or assertions Whether cases and assertions reflect requirements and meaningful edge conditions
GUI Propose actions, paths, or test data Whether paths represent user behavior and remain stable as the interface changes
System Generate data or scenarios across components Whether scenarios represent real integrations and important failure boundaries
Performance Help create or tune workloads and scenarios Whether the workload matches the question being measured and results are repeatable
Combinatorial Help explore combinations of options or parameters Whether important interactions and constraints are represented

The reviewed work includes supervised and reinforcement learning frequently, and also includes unsupervised methods such as filtering similar tests. That describes the surveyed publications, not a universal ranking of techniques. The study details its methods and evaluation measures.

How to evaluate ML-generated tests

Do not accept a model’s confidence score or a high prediction accuracy as proof that its tests protect the product. Evaluate the resulting testing outcome as a whole. The mapping study reports conventional measures such as fault detection, coverage, efficiency, and test size, alongside ML-specific measures such as prediction accuracy, adaptivity, training-data needs, and sensitivity.

  1. Start from a requirement or risk. State the behavior or failure mode the test should address. This gives reviewers a basis for judging generated inputs and assertions.
  2. Review the oracle. Check every generated expected value or assertion against the requirement, specification, or independently understood behavior. A test that merely reproduces current implementation behavior may lock in a defect.
  3. Run the tests and inspect failures. Determine whether a failure reveals a product fault, an invalid generated test, an environmental issue, or a flaky test.
  4. Measure useful outcomes. Track faults found, meaningful coverage, regressions caught, execution time, suite size, and review or maintenance effort. Interpret coverage alongside behavior; coverage alone does not establish correctness.
  5. Check input diversity and robustness. Include boundary values, unusual but valid inputs, invalid inputs, and relevant stress conditions. Google Research notes that tests limited to held-out data assumed to follow the training distribution may miss robustness failures and corner cases. Read the Google Research paper summary.
  6. Keep behavior-changing tests reviewable. Have a person approve tests and assertions that encode product behavior, and keep generated artifacts editable so the team can correct them.
  7. Compare against a baseline. Use a fixed existing suite or generation approach as a point of comparison. Record the target system, data, tool configuration, and evaluation conditions so results can be interpreted later.

These steps are practical guidance inferred from the documented oracle and evaluation challenges; they are not presented as a prescribed workflow from a standard.

Testing software that contains AI or ML

Testing an ML-enabled system creates an additional challenge: the output may not be deterministic, and expected behavior may be difficult to specify precisely. ISO/IEC TR 29119-11:2020 identifies the test-oracle problem—deciding what results should be and whether observed results pass—as a main challenge in testing AI-based systems. The ISO page describes the report as edition 1, published in November 2020, and currently under review, so check its status before relying on it for a formal process. See ISO/IEC TR 29119-11:2020.

In practice, define acceptance criteria that fit the system. Depending on the product, these may include allowable ranges, invariants, behavior across related inputs, or human review of representative outputs. Avoid assuming that one exact expected string or value is always suitable for a system whose valid outputs can vary. Add edge cases and stress conditions rather than relying only on average-case scores or a test set assumed to match training data.

Choosing an ML-assisted testing approach

There is no universal winning technique or tool. Compare approaches against the actual target and the cost of maintaining the output.

Decision area Questions to ask
Target Does it address unit, GUI, system, performance, or combinatorial testing?
Output Does it produce input data, executable tests, assertions, prioritization, or result classifications?
Adaptation Can it use relevant code, requirements, documentation, execution traces, or system-specific feedback?
Evidence Are faults found, meaningful coverage, input diversity, and regressions measured?
Operating cost What runtime, training data, labeling, integration, flakiness, review, and maintenance does it require?
Human control Can developers inspect, edit, and approve generated tests and expected behavior?

Browser-based visual checks with ScreenshotNeo

When a test needs to inspect a rendered page, a screenshot can be one artifact for visual review or comparison. ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. It is not an ML test generator; it can supply screenshot or PDF captures to a browser-based test or review workflow. Its API accepts one GET request with a URL and returns a PNG, JPEG, WebP, or PDF. See ScreenshotNeo and its API documentation.

Capture a page using cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Capture a page using Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Capture a page using Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Replace the example URL and provide your API key. For automated capture, check the response status before saving the body as an image: an error response should not be mistaken for a valid screenshot.

Options for visual test captures

ScreenshotNeo lists 63 options. Relevant choices for visual checks include full-page capture with lazy images loaded, a single element selected by CSS, dark mode, 12 device presets or a custom viewport, retina scale, image resizing, transparent backgrounds, and custom CSS or JavaScript. You can click an element before capture, hide selectors, or wait for a selector, a delay, or network idle.

For repeatable browser states, configure custom headers, cookies, user agent, timezone, or geolocation as needed. You can block ads, trackers, requests, or resource types. Caching supports a TTL you choose; disable or adjust it when a test needs a fresh page state. ScreenshotNeo also supports PDF options including paper size, margins, landscape, and page ranges; HTML/CSS-to-image; async jobs with signed webhooks; bulk capture of up to 100 URLs per call; signed links for public <img> tags; a usage API; and an OpenAPI spec. Parameter names used by other screenshot APIs also work to ease migration. Consult the docs for parameters and response details.

Use captures as test evidence carefully

  • Use a fixed viewport, device scale, color mode, and relevant browser state when comparing captures.
  • Wait for a meaningful page condition, such as a selector becoming available, rather than relying on a short arbitrary delay where possible.
  • Decide whether dynamic regions should be hidden or stabilized; otherwise timestamps, rotating content, or personalized data may create noisy comparisons.
  • For full-page shots, account for lazy-loaded content and long pages. For component checks, capture the relevant element to reduce unrelated changes.
  • Store capture settings with the test result so a later difference can be traced to a page change or a changed capture configuration.

Or skip the browser setup

Use one API call to capture a page. The ScreenshotNeo API docs describe the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor, then removes 60+ known consent platforms, newsletter popups, and chat widgets before the capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the shot was billed. Its MCP server lets AI agents such as Claude, Cursor, or any MCP client use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan, and yearly billing gives two months free. Sign up for 1,000 free screenshots a month with no card.

Performance, reliability, and cost considerations

Performance

Generated tests can reduce manual authoring in some workflows, but generation, model inference, setup, and review all have costs. Measure end-to-end time, including failures investigated and tests maintained, rather than only generation speed. For browser captures, page weight, full-page length, wait conditions, and batch size affect the work involved; choose capture scope and waits that match the check.

Reliability

Generated tests can be brittle if they depend on unstable UI details or encode accidental behavior. Review failures, measure flakiness and maintenance, and make generated tests reproducible enough to diagnose. For AI systems, use acceptance criteria suited to potentially variable outputs, and include stress conditions and corner cases.

Cost

Account for compute or service charges, training and labeling needs, engineering integration, runtime, and the human cost of reviewing and maintaining tests. The cited research does not establish a universal return on investment or representative production adoption rate. For ScreenshotNeo, only clean shots are billed; response headers identify verdict and billing status. Its published plans are Free with 1,000 shots per month, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Check the product site for current plan details.

Troubleshooting generated tests

Symptom Likely cause What to do
A generated test passes but a bug remains The test does not assert the intended requirement, or exercises only ordinary inputs Review the oracle against requirements; add boundary, invalid, and risk-based cases; measure faults found as well as coverage
A generated assertion fails repeatedly The expected result may be wrong, the behavior may be variable, or the environment may differ Check the requirement and execution conditions; define an appropriate acceptance criterion for variable output
The test suite grows but finds no additional defects Generated cases may be redundant or target already-covered behavior Measure marginal fault detection and meaningful coverage; consider filtering similarity and review suite size against value
Tests fail intermittently Timing, external dependencies, unstable UI state, or non-deterministic output can vary Control inputs and environment where possible, wait on explicit conditions, and distinguish product variation from test flakiness
Model evaluation looks strong but edge cases fail Evaluation data may reflect the training distribution and omit stress conditions Add representative corner cases and stress tests; do not rely on held-out average-case accuracy alone

Troubleshooting screenshot captures

Symptom Likely cause What to do
Saved file is not a valid image The request returned an error body or unsuccessful status Check the HTTP status before writing the body; inspect response headers and consult the API docs
Page appears blank or incomplete The page did not load successfully, or capture occurred before content appeared Use an appropriate selector, delay, or network-idle wait; check the returned page verdict and billing headers
Screenshot differs between runs Viewport, page state, dynamic content, cache, or timing differs Fix capture settings and state, wait for a stable condition, and choose an intentional cache TTL
Expected banner or widget is missing Consent and other known overlays are removed before capture by default Turn off the relevant cleanup step when the overlay itself is what the test needs to inspect
Full-page capture omits content Lazy content may not have loaded or the page may depend on interaction Use full-page capture with lazy images loaded and, when needed, click an element or wait for the relevant selector
Large capture workflow is slow Many independent URLs or large pages are being captured individually Consider bulk capture for up to 100 URLs per call, or async jobs with signed webhooks; tune page scope and wait behavior

Frequently asked questions

Does machine learning replace conventional automation?

No. ML can assist generation, selection, and analysis, while conventional test frameworks execute tests and teams still define behavior and review results.

Does a generated test prove the software is correct?

No. It checks selected behavior under selected conditions. Confidence depends on the quality of the test, oracle, and coverage of relevant risks.

Can screenshot tests verify application logic?

A screenshot records rendered appearance. Use assertions and other tests to check logic, data integrity, and behavior that a visual capture cannot establish.

Are the published research results product benchmarks?

No. The TOGA figures are results from a scoped research evaluation; they do not establish typical results across products or codebases.

Can AI agents request screenshots?

ScreenshotNeo provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and any MCP client.

Further reading