Challenges of Generative AI in Software Testing
Generative AI can speed up test creation, but generated tests still need review. Learn the risks, evidence, and safeguards that help teams assess them.
Generative AI can help developers draft test cases, suggest edge cases, and explore behavior. But a generated test is only a candidate: its assertions may be wrong or weak, and its execution may be unstable. Teams should review expected behavior, run tests repeatedly, and assess bug-detection strength instead of treating test volume or coverage as proof of quality.
The evidence is promising but bounded. A 2025 study of Java test oracles found average mutation scores of 43% for generated oracles and 45% for human-designed ones in its particular setup. A 2026 study of four database systems found that 72 of 115 flaky generated tests examined relied on an order that was not guaranteed. Neither result is a universal rate for all models, languages, or teams.
What generative AI changes in software testing
LLMs can turn code, requirements, bug reports, or examples into test ideas and test code. They can also help explain failures or propose data combinations. This can reduce the effort of getting a first draft, but it does not remove the central testing question: how do you know whether the observed result is correct?
A test oracle defines the expected behavior against which a result is judged. A test can execute successfully and still be ineffective if it asserts the wrong thing, checks only a trivial condition, or misses the defect it was meant to catch. AI may generate both the test and its expected result, so plausible-looking output is not independent evidence of correctness.
Key challenges, with evidence and limits
1. Test oracles can be weak or incorrect
In an ASE 2025 study, Davide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst, and Mauro Pezzè evaluated 13,866 test oracles from 135 Java projects. The project oracles postdated the tested models’ training cutoffs. Generated oracles achieved an average mutation score of 43%, compared with 45% for human-designed oracles. The authors also identify limits for complex oracles and describe thorough oracle generation as an open problem. These measurements apply to that study’s models, projects, language, and methods; they do not establish that generated tests generally match human quality. Read the oracle study.
Mutation score measures how many intentionally introduced program changes (“mutants”) a test suite detects. It provides evidence about whether assertions distinguish some faulty behaviors from the original, but it is not a complete measure of real-world correctness.
2. Tests can depend on unstable ordering or state
A 2026 study of four database systems found a slightly higher proportion of flaky cases among LLM-generated tests than among existing tests. Of 115 flaky generated tests the researchers examined, 72 (63%) relied on an order that was not guaranteed, for example by omitting an explicit SQL ORDER BY. This is a finding in the studied database settings, not a general flakiness percentage for AI-generated tests. Find the study.
Generated tests may accidentally depend on iteration order, current time, random values, shared state, network availability, or unspecified runtime behavior. These dependencies can make a test pass locally and fail in CI, or fail inconsistently on repeated runs.
3. Benchmark results can overstate generalization
Public evaluation benchmarks can overlap with model training data, which threatens the independence of an evaluation. The oracle study addressed this threat for its reported experiment by using project oracles created after the tested models’ training cutoffs. That design reduces one specific concern; it does not show that all test-generation benchmarks are contaminated, nor does it eliminate every evaluation bias.
4. Plausible output can contain reasoning errors
ISTQB’s 2025 sample-exam materials state that hallucinations and reasoning errors are intrinsic challenges of current AI technologies and that testers cannot prevent them from occurring, but should identify and mitigate their risks. This is certification guidance, not a measured prevalence rate. The reviewed evidence does not establish a universal hallucination percentage for software testing. See ISTQB AI testing materials.
5. User-perceived help is not the same as measured test quality
An observational study involving 12 undergraduate students reported perceived time savings and help with test ideation, alongside diminished trust, quality concerns, and a lack of ownership. It found no significant effects of prompting strategies on measured test effectiveness or test code quality. Because this was a small novice-student sample, it should not be treated as evidence about all professional teams. Read the student study.
6. Coverage does not establish fault detection
Line or branch coverage shows which code executed, not whether a test would detect an incorrect result. A 2024 study introduced MuTAP, which uses mutation testing to assess whether generated tests expose seeded faults. Mutation testing is one useful evaluation approach, not a guarantee of real-world bug detection or a universally accepted single metric. Read about MuTAP.
A practical review workflow for AI-generated tests
- Define the behavior first. Write down the requirement, invariant, or defect being tested. Identify the expected result independently where possible.
- Review setup and assertions. Check that the test exercises the intended path, that assertions distinguish correct from incorrect behavior, and that expected values do not merely repeat the implementation.
- Inspect assumptions. Look for implicit ordering, shared mutable state, dependence on wall-clock time, random seeds, environment variables, external services, or data that may not exist in CI.
- Run repeatedly in relevant environments. Rerun tests locally and in CI, including under different orderings or supported runtime versions where practical. Reruns can reveal instability but cannot prove that a test is never flaky.
- Measure more than coverage. Use mutation testing or representative fault injection where feasible. Review which mutations survive and whether they expose weak assertions or missing cases.
- Keep evaluation independent. When comparing models or prompts, record the model version, language, project type, dataset, and evaluation method. Prefer data whose relationship to model training is understood.
- Keep a human owner. A developer should be able to explain what a test protects and why its expected behavior is correct before it becomes a relied-upon regression check.
How to evaluate a generated test suite
| Question | Useful evidence | Limit |
|---|---|---|
| Are expected results trustworthy? | Trace assertions to requirements, contracts, independently computed examples, or reviewed reference behavior. | A plausible assertion can still encode a mistaken requirement. |
| Does the suite catch faults? | Mutation score, seeded fault detection, and inspection of surviving mutants. | Mutants are proxies; a score does not predict every production defect. |
| Is execution stable? | Repeated runs, order variation, clean test data, and CI runs across supported environments. | Passing reruns do not guarantee future stability. |
| Does the model generalize? | Evaluation on projects and tasks with a known relationship to training data; report model and dataset details. | Post-cutoff data addresses one contamination threat, not all sources of bias. |
| Is the workflow useful? | Task-level outcomes such as defect detection, error tracing, localization, review effort, and maintainability. | Perceived speed or more generated tests alone does not establish effectiveness. |
Results from different languages, repositories, participant groups, models, and evaluation methods are not directly interchangeable. State those details whenever reporting a comparison.
Where AI assistance fits
AI is often most useful for expanding a developer’s test ideas: boundary values, alternate inputs, malformed data, and candidate regression cases. A reviewer can then select cases that correspond to specified behavior and improve their assertions. This keeps the tool’s contribution concrete while preserving human responsibility for the oracle and the test’s purpose.
For browser-facing applications, screenshots can be one artifact in a visual regression workflow, but a screenshot alone does not prove application behavior. Pair visual comparisons with assertions about state and interaction. ScreenshotNeo is a website screenshot API and MCP server made by Yorker Media; see ScreenshotNeo and its API documentation for capturing pages as PNG, JPEG, WebP, or PDF.
Or skip the browser setup
For a page screenshot, call the API directly:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. See the API options and documentation, then sign up for 1,000 free screenshots a month with no card.
Performance, reliability, and cost considerations
- Review effort: Count the time needed to verify assertions and remove unstable assumptions, not just the time to produce test code.
- Execution cost: Larger generated suites can increase CI runtime. Keep cases that add meaningful behavioral or fault-detection coverage and remove duplicates.
- Reliability: Isolate test data and external dependencies where practical; make order requirements explicit. A test that passes only under one execution sequence is a maintenance cost.
- Model and data cost: The reviewed sources do not quantify inference costs or organizational savings. Track those in the team’s own workflow rather than projecting a general return from these studies.
- Privacy and security: The reviewed research does not establish general rates or impacts for privacy exposure, security vulnerabilities, intellectual-property disputes, or organizational costs specific to generative-AI testing. Assess these according to the code, prompts, and policies in your own environment.
Troubleshooting common problems
| Symptom | Likely cause | What to do |
|---|---|---|
| Test passes, but a known defect remains | The assertion is too weak, checks the wrong result, or does not reach the faulty path. | Trace the assertion to the requirement, add a targeted failing case, and inspect surviving mutations where feasible. |
| Test fails intermittently | Unspecified ordering, shared state, time, randomness, or an external dependency. | Make ordering explicit (for SQL, use ORDER BY when order matters), isolate state, control randomness, and capture failure context. |
| Generated expected value looks plausible but is wrong | The model inferred behavior rather than deriving it from a trusted specification or independent oracle. | Verify against a contract, reviewed example, or independently calculated result; do not approve by fluency alone. |
| Many tests, little added protection | Tests duplicate paths, assert implementation details, or increase coverage without meaningful assertions. | Group by behavior, remove duplicates, and evaluate fault detection as well as coverage. |
| Benchmark result is unexpectedly strong | Dataset overlap, narrow task selection, or an evaluation that rewards superficial success may be involved. | Document the dataset and model version, use independent or post-cutoff data where possible, and report the evaluation method and limits. |
| Prompt changes do not improve test quality | Prompting may not be the limiting factor; oracle ambiguity or task difficulty can dominate. | Clarify expected behavior and measure outcomes such as effectiveness and code quality instead of assuming a prompt strategy will help. |
FAQ
Can an LLM-generated test be trusted without review?
No. Review the test’s purpose, setup, and expected result. Generated tests are candidates, not ground truth.
Does higher code coverage mean better AI-generated tests?
Not by itself. Coverage records execution; mutation testing or fault-based evaluation can offer additional evidence about whether assertions detect incorrect behavior.
Are AI-generated tests always more flaky?
The reviewed database study found a slightly higher proportion of flaky generated tests in its four systems. It does not establish a universal result across software domains.
Do these studies show that AI testing saves professional teams time?
No general conclusion follows. The student study reported perceived help among 12 undergraduates but did not establish a general professional productivity effect.
Sources and scope
- Molinelli, Di Grazia, Martin-Lopez, Ernst, and Pezzè (2025), empirical study of generated test oracles and mutation scores: paper.
- “On the Flakiness of LLM-Generated Tests…” (2026), study of four database systems: paper search.
- ISTQB (2025), AI testing certification materials: official materials.
- Ardic, Le Dilavrec, and Zaidman (2025), observational study of undergraduate testing workflows: paper.
- MuTAP (2024), mutation-testing approach for assessing generated tests: paper.
All study findings above are limited to their reported models, projects, participants, datasets, and tasks. The reviewed evidence does not support universal rates for hallucinations or flakiness, nor broad claims about privacy, security, intellectual property, or organizational impact.


