How Machine Learning Is Used in Software Testing
Learn how machine learning supports test generation, regression prioritization, and defect prediction—and how to evaluate its limits in your workflow.
Machine learning (ML) can help software teams generate test cases, order regression tests, and estimate which components may carry higher defect risk. It learns patterns from code, existing tests, execution history, or other project data, then offers suggestions or predictions for developers to review. Those outputs can guide testing, but they do not prove a program correct or guarantee that a bug will be found.
There are two related meanings to separate. Using ML to test conventional software means applying learned methods to activities such as test generation and prioritization. Testing software that contains ML models means evaluating properties such as correctness, robustness, and fairness in a system whose behavior depends on learned parameters and data. This guide covers both, with emphasis on the first.
1. What machine learning does in software testing
Traditional test automation executes tests written by people or generated through fixed rules. ML adds a learned estimate: given examples from code, tests, and past runs, which inputs might be useful, which tests should run first, or which components deserve more attention?
The result is decision support in a testing workflow. A generated test still needs review; a risk score is not a discovered defect; and a prioritized suite still needs an appropriate full-suite strategy. ML works best when its inputs are relevant to the project and its outputs can be checked against observable behavior.
| Testing task | How ML may help | What it does not establish |
|---|---|---|
| Test generation | Suggest test structures, inputs, properties, or expected outputs from code and examples. | That the tests are correct, maintainable, or comprehensive. |
| Test selection and prioritization | Estimate which regression tests are useful or should run earlier after a change. | That omitted tests can be skipped permanently or that a fault will be caught. |
| Defect prediction | Estimate which components may be more fault-prone based on past data. | That a predicted component contains a fault. |
| ML-system testing | Help evaluate model-containing systems against requirements such as robustness or fairness. | That one generic test suite is sufficient for every ML application. |
2. Test generation: proposing inputs and checks
A test-generation model can use source code, examples, existing tests, or related project information to propose inputs and test structures. Research describes applications across unit, GUI, system, performance, and combinatorial testing, as well as property-based tests, verdicts, and expected outputs. The useful output depends on the task: a plausible input is not enough if the assertion checks the wrong behavior.
One documented example is Microsoft Research’s AI for Testing project. Its project description says it trains transformer models on developer code and currently supports C# in Visual Studio and Java in VSCode. The stated aims include finding bugs, increasing coverage on existing methods, and supporting test-driven development for methods not yet implemented. These are project goals and scope statements, not evidence that generated tests universally improve results or a statement of commercial availability. Microsoft Research: AI for Testing.
Generated tests should be treated as candidate code. Review whether each test:
- Exercises a meaningful behavior or boundary rather than merely executing a line.
- Has an assertion that would fail for a relevant regression.
- Uses valid setup and deterministic inputs.
- Would remain understandable and maintainable after the implementation changes.
- Does not encode an accidental current output as the intended specification.
For test-driven development, a generated test can help make a requirement concrete, but the developer still has to decide whether its behavior matches the requirement. If the method is not implemented yet, an expected result can be especially uncertain unless the requirement is explicit.
3. Regression testing: selecting and ordering tests
As a regression suite grows, running every test after every change can delay feedback. An ML-based approach can use test attributes and project history to estimate which tests are likely to be useful or should run earlier. A University of Luxembourg repository summary describes combining partial and imperfect sources to predict test selection and prioritization for earlier continuous-integration feedback. University of Luxembourg research repository.
Prioritization changes the order in which tests execute. Selection may choose a subset for an early run. These are different operational choices:
- Prioritize: run the predicted high-value tests first, then continue with the remaining suite.
- Select: run a subset for a particular feedback window, while retaining a policy for broader validation.
A practical workflow is to use ranking to reduce time to an early signal, while keeping full-suite execution in a later CI stage, on a schedule, or before release where appropriate. Track faults found by early tests, runtime, skipped-test exposure, and cases where the ranking missed a useful test. If a wrong prediction can cause a costly release defect, keep safeguards that do not depend on the prediction.
4. Defect prediction: estimating where to look
Defect-prediction models learn associations between code or project characteristics and previously labeled defects. They estimate which components may deserve extra review or testing attention. A software-quality-assurance survey describes this as predicting components likely to contain more faults in a future release, supporting planning and corrective action. Software quality assurance survey.
Use such scores as a risk signal alongside code ownership, recent changes, incident history, and engineering judgment. A high score does not mean a defect has been found, and a low score is not evidence that a component is safe. Results can transfer poorly when the new project differs from the training data, coding practices change, or historical defect labels are incomplete or biased.
5. Testing software that contains ML models
This is a different testing problem from using ML to generate or prioritize tests. In a model-containing system, behavior may depend on data and learned parameters rather than a short, fixed set of rules. Test design should start from the system’s requirements and failure costs.
An IEEE survey organizes ML-system testing around properties including correctness, robustness, and fairness; components including data, the learning program, and its framework; and workflow stages including test generation and evaluation. Its survey covers 144 papers. IEEE, “Machine Learning Testing: Survey, Landscapes and Horizons”.
- Correctness: define expected behavior for the application and assess model outputs against that requirement.
- Robustness: evaluate whether relevant input changes or perturbations cause unacceptable behavior.
- Fairness: define the applicable criteria and assess outcomes across relevant groups. The criteria depend on the domain and use case.
- Data and pipeline: check the data and the components that prepare, train, serve, or evaluate the model, as well as the model itself.
There is no universal fairness or robustness test that substitutes for application-specific requirements. The test plan needs to state what acceptable behavior means, which data and populations it covers, and how failures are handled.
6. Learning approaches and what research counts mean
Different tasks and datasets call for different methods. A 2023 systematic mapping study examined 124 publications and reported supervised learning and reinforcement learning among common approaches to automated test generation; it also identified unsupervised and semi-supervised work. A separate 2024 systematic review examined 40 studies from 2018 through March 2024 and classified supervised, unsupervised, reinforcement, and hybrid methods. These numbers describe the samples in those reviews; they are not counts of the entire field or evidence that one method is best.
These approaches are broad categories, not product guarantees:
- Supervised learning learns from examples paired with labels, such as historical test outcomes or defect records.
- Unsupervised learning looks for patterns in data without the same kind of labeled outcome.
- Reinforcement learning learns through action and feedback, which can suit iterative search or generation settings.
- Hybrid methods combine techniques or learned components with other methods.
The 2023 mapping study covers 124 publications; the 2024 review covers 40 studies over its stated period; and the IEEE survey covers 144 papers. Their scopes and inclusion methods differ, so do not compare these sample sizes as measures of field growth. The reviews do not establish a general percentage improvement in quality, coverage, or cost for every team.
Sources: Fontes et al., 2023 mapping study; 2024 systematic review of machine learning methods in software testing; IEEE survey, 2022.
7. A practical adoption process
- Choose one task. Decide whether the problem is generating candidate tests, getting earlier regression feedback, estimating component risk, or validating an ML-containing system.
- Record the current baseline. Capture suite runtime, feedback delay, relevant faults found, coverage measures appropriate to the task, and maintenance effort. Avoid relying on a single metric.
- Inspect the inputs. Identify what the approach learns from: source, existing tests, execution history, defect labels, data, or documentation. Check whether that information is representative and current.
- Run suggestions alongside existing practice. Compare generated tests or ranked tests with the current workflow before making a prediction a gate or excluding tests.
- Review failures and misses. Inspect false positives, incorrect expected outputs, flaky tests, and faults found only by lower-ranked or omitted tests.
- Set a fallback and ownership. Define who maintains the tests and model inputs, how the full suite runs, and what happens when the ML component is unavailable or its output is low confidence.
- Reassess when the project changes. New frameworks, code patterns, test conventions, or defect-labeling practices can make historical data less relevant.
8. Choosing an approach or tool
Compare approaches against the job they perform rather than the label “AI testing.” Useful evaluation questions include:
| Axis | Questions to answer |
|---|---|
| Task | Does it generate tests, rank or select tests, predict component risk, or evaluate an ML system? |
| Inputs | Does it need code, existing tests, execution history, labels, test data, or documentation? Can those inputs be used safely? |
| Integration | Which languages, IDEs, frameworks, and CI steps are supported today? Distinguish current support from roadmap statements. |
| Evidence | Are the evaluation projects representative? Are fault models, metrics, datasets, and reproduction details explained? |
| Human review | Can developers inspect generated tests and understand recommendations before relying on them? |
| Failure cost | What happens if an oracle is wrong, a prediction misses a risky component, or a ranking delays an important test? |
Published evidence depends on datasets, suites, fault models, and workflow. Ask for results relevant to the team’s own languages and CI setup, and keep a way to measure misses as well as successes. The cited literature provides a map of approaches and one official project example; it does not establish a market-wide comparison or prove that a particular tool will help every team.
9. Troubleshooting common problems
| Symptom | Likely cause | What to do |
|---|---|---|
| Generated tests pass but catch no regressions. | Tests may execute code without checking meaningful behavior, or assertions may reflect current implementation details. | Review assertions against requirements and mutation or fault scenarios relevant to the project. Measure useful fault detection, not just execution or line coverage. |
| Generated expected outputs are wrong. | The model inferred an oracle from examples that do not fully specify intended behavior. | Have an owner validate expected results against requirements, add explicit examples, and reject unsupported assumptions. |
| Ranked tests miss a regression. | Historical signals may not represent the change, or selection may omit a useful test. | Retain the rest of the suite for broader runs, inspect the miss, and update the ranking inputs or policy. |
| Defect scores are misleading on a new project. | Training data, labels, language, or development practices differ. | Validate locally before using scores to allocate scarce review time; treat scores as estimates and monitor drift. |
| Tests are flaky or hard to maintain. | Generated tests may depend on unstable timing, external state, random values, or overly specific outputs. | Make setup deterministic, isolate dependencies where practical, add stable waits only when behavior requires them, and edit or discard low-value tests. |
| Results improve in a paper but not in CI. | The evaluation may use different projects, test suites, labels, fault assumptions, or execution constraints. | Compare on the team’s own workflow and report the evaluation conditions and limitations. |
| Teams disagree about fairness or robustness. | The acceptance criteria are not defined for the application. | Specify the relevant groups, input variations, thresholds, and failure response with the system’s stakeholders before choosing metrics. |
10. Performance, reliability, and cost
ML can add a model inference step, data preparation, integration work, and review effort. Whether it saves time depends on the task and workflow. Test prioritization may shorten time to an early signal while still requiring the rest of the suite later. Generated tests can increase coverage candidates while also adding review and maintenance cost. Defect prediction may help focus attention, but a bad estimate can divert it.
For reliability, avoid making unvalidated predictions the sole release gate. Keep a fallback when the model or its inputs are unavailable; monitor how often suggestions are accepted, useful, incorrect, or missed; and reassess after substantial changes in code or test practices. For cost, include model operation and integration where applicable, plus developer review, CI runtime, test maintenance, and the cost of missed faults. The research summaries cited here do not supply a universal benchmark or savings figure.
11. Capture visual states as test evidence
Some testing workflows need a rendered page image to inspect or compare visual behavior. A screenshot is evidence of one captured state; it does not by itself validate application logic, accessibility, or correctness. For a team building a screenshot step into its workflow, record the page URL, viewport, relevant state, and capture timing so a later image can be interpreted in context.
ScreenshotNeo is a website screenshot API and MCP server for developers. Its API can return PNG, JPEG, WebP, or PDF captures, and its MCP server exposes screenshot and page-information tools to AI agents. For capture details and parameters, see the ScreenshotNeo site and API documentation.
12. Or skip the browser setup
To capture a page without setting up a browser, make one GET request. The examples below use the supplied Stripe URL; replace it with the page you need. Create an API key through the service before substituting YOUR_API_KEY. See the ScreenshotNeo API documentation for request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
13. Frequently asked questions
Does machine learning replace software testers?
No. It can assist with generating candidates or estimating priorities, but people still define expected behavior, review tests, and decide how to respond to risk.
Does more code coverage mean better tests?
Not by itself. Coverage can show which code ran; it does not show whether assertions would catch the failures that matter.
Can a team use ML testing without historical data?
Some approaches can use code, examples, or other information, but the needed inputs depend on the task and method. Check the actual input requirements and validate results locally.
Are research sample sizes performance results?
No. The publication counts above describe the papers included in particular reviews; they do not measure accuracy, savings, or improvement.
What is the first useful pilot?
Choose one bounded task, preserve the existing validation path, and compare the ML-assisted output with your current process on representative changes.


