How Large Language Models Are Changing Software Testing: Part 2
LLMs can draft tests and power applications that need testing. Learn how to validate generated tests, measure their value, and evaluate variable model behavior.
Large language models (LLMs) are changing software testing in two distinct ways: developers use them to draft or improve tests for conventional software, and teams must test applications that use LLMs as components. In both cases, a generated test is a candidate to inspect and validate, not proof that the code is correct.
A useful workflow combines requirements, human review, conventional automated checks, coverage, and—where appropriate—mutation testing. For LLM-enabled applications, it also accounts for variability across repeated runs, inputs, prompts, configurations, and model versions.
1. Where LLMs fit in software testing
An LLM can help propose test cases, reason about a targeted code path, turn ambiguous requirements into questions and examples, and assist with debugging-related work. These tasks can reduce the effort of getting a first draft, but they do not remove the need to decide what correct behavior means.
Keep two testing problems separate:
- Testing conventional software with an LLM: the model proposes or improves tests for code whose expected behavior should be established from requirements and trusted examples.
- Testing software that contains an LLM: the system under test includes a component whose outputs may vary. The test strategy must assess individual responses and behavior across a set of cases or repeated runs.
Confusing these roles can lead to misplaced confidence. A test suite generated by a model may encode the same mistaken assumption as the implementation it is supposed to check.
2. Generated tests need more than a successful run
A generated test can compile and pass while being incorrect, redundant, hard to understand, or weak at detecting defects. Evaluate test quality on separate dimensions:
| Dimension | Question to ask |
|---|---|
| Correctness | Does the test assert the behavior the requirement actually specifies? |
| Readability | Can a maintainer see the setup, input, expected result, and reason for the case? |
| Coverage | Does it execute the intended statements, branches, or paths? |
| Bug detection | Would it fail when relevant behavior is changed incorrectly? |
Coverage and bug detection are related but not interchangeable. Executing a branch does not establish that the assertion would catch a faulty result on that branch. Likewise, a passing test demonstrates consistency with its assertions, not that those assertions represent the right specification.
The 2024 ASE study record from Aalto describes an evaluation across 690 Java classes, four LLMs, five prompting techniques, and 216,300 generated tests. It assessed correctness, readability, coverage, and bug detection, and its abstract says correctness still needs improvement. Those are the scope and conclusion of that study, not a universal comparison of LLMs with conventional generators or a guarantee about other models and projects. See the Aalto research record.
3. Ask for targeted behavior, not just more tests
Test generation includes different tasks: broad coverage, reaching a particular line or branch, and satisfying the conditions along a path. A test that reaches a selected branch may require reasoning about the program’s inputs and control flow; a plausible sample is not enough.
The 2025 TESTEVAL paper describes these coverage tasks and a dataset of 210 Python programs from LeetCode. Its results are tied to that benchmark and setup. Read TESTEVAL in Findings of NAACL 2025.
For example, suppose a function has a branch guarded by amount >= limit. Ask the model to propose inputs that exercise both sides, including the boundary value, and explain which path each input should take. Then run coverage and inspect whether the assertions check the intended result. This is an explanatory example, not a reported experiment.
A practical test-generation workflow
- Provide the relevant source code, nearby trusted tests, and the behavioral requirement. State the language, test framework, and constraints.
- Ask for a small set of cases grouped by behavior: normal inputs, boundaries, invalid inputs, and the target branch or path. Request a brief explanation of each case.
- Review the expected results against the requirement before running the tests. Change assertions that merely repeat the implementation’s assumptions.
- Run the project’s formatter, type checks, and test command as appropriate. Treat syntax errors and failing tests as feedback to investigate, not as reasons to accept a revised test automatically.
- Measure statement or branch coverage for the target and inspect whether the test assertions would detect a wrong result.
- Where the behavior matters, use mutation testing or known defect examples to check whether the suite catches meaningful changes.
- Keep useful tests, remove duplicates, and record any unresolved ambiguity in the requirement.
Providing surrounding context helps the model produce tests that fit the existing codebase. It does not make the output trustworthy by itself; review remains necessary.
4. Use mutation testing to check whether tests can catch changes
Mutation testing makes small changes to a program and checks whether the test suite detects them. A surviving mutation can reveal a behavior that is executed but not meaningfully asserted, or an untested path. A killed mutation indicates that some test noticed the change, though it does not prove that the suite covers every relevant fault.
The 2024 MuTAP article describes adding mutation-testing feedback to prompts and reports a 93.57% average mutation score in its experimental setup. That figure is specific to the study’s models, programs, prompts, and mutations. It is not an expected production score or a cross-project guarantee. Mutation score is a proxy for fault detection under the selected mutations, not a complete measure of test usefulness. See the MuTAP article record.
Use mutation results diagnostically. Review surviving mutations that represent meaningful behavior, and consider whether equivalent mutations or changes outside the selected mutation operators make the score misleading.
5. Tests can clarify requirements and help select code
Tests can be part of an interactive workflow for clarifying intent before code is accepted. TiCoder uses test-driven interaction to help users clarify requirements while generating code. Microsoft Research reports an average absolute improvement of 45.97% in pass@1 across four LLMs and two Python datasets within five user interactions. The paper describes its user feedback as an idealized proxy, so this result is evidence for that bounded setting, not a forecast for every team. Read the TiCoder research summary.
Tests may also help select among candidate programs. An ISSTA 2024 study describes checking candidates for consistency with an LLM-generated test suite, while acknowledging that generated programs can be incorrect. The selection oracle is only as reliable as the expected behavior encoded in the tests. If the test generator and implementation share an incorrect assumption, consistency can select a wrong program. See the ISSTA 2024 proceedings abstract.
6. Testing applications that contain LLMs
For an LLM-enabled application, exact-string snapshots can be brittle when equivalent responses differ, yet overly broad checks can miss important failures. Define what must be true about a response and the surrounding system, then choose checks that match that contract.
A 2025 taxonomy paper highlights variability in testing goals, systems under test, and inputs. It distinguishes atomic oracles, which assess an individual result, from aggregated oracles, which assess behavior across multiple results. It also identifies weaknesses in how current tools capture repeated runs, model versions, and configurations. This supports treating variability and reproducibility as explicit testing concerns; it does not validate a particular testing product. Read the 2025 testing taxonomy.
Build an evaluation set around the contract
- Correctness criteria: Use deterministic assertions where possible. For outputs that can vary, define semantic checks and document their limitations, including any evaluator that uses a model.
- Behavior coverage: Include typical inputs, boundary cases, malformed or adversarial inputs where relevant, safety constraints, and targeted scenarios.
- Variability: Repeat cases where variation matters. Record the model version and the prompt, configuration, and input conditions needed to understand a result.
- Regression value: Ask whether a failure represents a behavior change that matters. A text difference alone may be harmless; identical output can still violate an external requirement.
- Review and reproduction: Preserve failing examples and enough context to reproduce or explain them. Have a person review whether the evaluation judgment matches the intended behavior.
These axes synthesize concerns in the cited taxonomy and empirical work; they are a practical checklist, not a checklist validated as a single standard by one paper. A 2024 software-engineering perspective organizes research, practice, tools, and benchmarks for testing LLMs as components, while a 2025 roadmap groups collaboration into preparation, interaction, and validation stages. Together they illustrate that this is a broader testing discipline with technical and social challenges, not just a prompt-writing task. See the 2024 software-engineering perspective and the 2025 research roadmap.
7. Record enough to investigate regressions
When an evaluation fails, preserve the case, the expected criterion, and the relevant execution context. At minimum, record the input, system or prompt version, model version, configuration, and the observed result. For repeated evaluations, keep the number of runs and outcomes so a change in aggregate behavior is distinguishable from one unusual response.
Make failures actionable: store examples that a developer can inspect, identify whether the check was deterministic or semantic, and note known evaluator limitations. If the model or prompt changes, rerun the relevant cases rather than assuming prior results still apply.
8. Common mistakes and how to fix them
| Symptom | Likely cause | What to do |
|---|---|---|
| Generated tests do not compile | The model guessed an API, import, fixture, or framework version. | Provide the project’s existing test style and actual signatures; inspect and correct setup before evaluating behavior. |
| Tests pass but a known bug remains | Assertions are weak, cover the wrong case, or encode the same mistaken assumption as the code. | Write down expected behavior independently, add a failing regression case, and inspect mutation survivors. |
| Coverage rises without confidence rising | Tests execute code but do not assert meaningful outcomes. | Review assertions and use targeted mutants or known defects to test detection. |
| Boundary branch stays uncovered | Inputs do not satisfy the path conditions, or the model inferred the condition incorrectly. | Trace the condition manually, include values immediately below, at, and above the boundary when valid, then check branch coverage. |
| LLM evaluation fails inconsistently | The output varies across runs, or model/configuration context was not controlled or recorded. | Repeat relevant cases, retain run context, and define whether the criterion applies per response or across a set. |
| Snapshot tests produce noisy failures | The test expects exact wording where the contract allows variation. | Assert stable semantic properties and retain exact snapshots only for behavior that truly requires exact output. |
| A model-based evaluator disagrees with reviewers | The evaluator is an imperfect oracle or the requirement is ambiguous. | Inspect examples, refine criteria, and use human review for consequential or unclear judgments. |
| Results cannot be reproduced after a release | The model version, prompt, configuration, input, or run count was not retained. | Version and store those evaluation inputs alongside the results. |
9. Performance, reliability, and cost
LLM assistance adds a generation and review step. Keep requests focused on the code and behavior in question, and ask for a manageable set of cases rather than many near-duplicates. Measure the workflow in your own project: the cited studies use specific datasets and setups and do not establish general time savings, adoption rates, or defect reductions.
For conventional code, make the test suite’s normal checks reproducible and run generated cases through them before relying on the result. For an LLM component, repeated runs can reveal variability but also add evaluation cost; choose run counts based on the behavior and risk being assessed, and report the count with the result. A passing evaluation on one prompt, model version, or run does not guarantee behavior under another configuration.
There is no general cost or performance figure in the cited research that applies across models, providers, and projects. Estimate costs from the model and evaluation configuration you actually use, including repeated calls and human review, and compare that with the value of the behavior being checked.
10. Capture website output while documenting tests
Some testing and regression workflows need a rendered website capture alongside the test result—for example, to inspect a visual state or attach a page artifact to a report. You can capture pages with a browser you control, or use ScreenshotNeo, a website screenshot API and MCP server for developers. The examples below capture Stripe; replace the target URL with a page you are authorized to access. The API supports PNG, JPEG, WebP, and PDF output. See the ScreenshotNeo API documentation for request options and response details.
DIY: capture a page with Playwright in Python
Install Playwright and its Chromium browser in an environment where browser downloads are permitted:
python -m pip install playwright
python -m playwright install chromium
Save this as capture.py and run python capture.py. It writes a full-page PNG and closes the browser even if capture fails.
import asyncio
from pathlib import Path
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page(viewport={"width": 1440, "height": 900}, device_scale_factor=1)
try:
response = await page.goto("https://stripe.com", wait_until="networkidle", timeout=60000)
if response is None:
raise RuntimeError("Navigation did not return an HTTP response")
if response.status >= 400:
raise RuntimeError(f"Page returned HTTP {response.status}")
await page.screenshot(path="shot.png", full_page=True)
print(f"Saved {Path('shot.png').resolve()}")
finally:
await browser.close()
asyncio.run(main())
networkidle can be unsuitable for pages with persistent network connections or frequent background requests. If it times out, wait for a meaningful selector instead, such as await page.locator("main").wait_for(), or use an explicit delay when the site’s rendering behavior requires it. An HTTP error check is useful but does not establish that the page is the right content; inspect the captured result as part of the workflow.
Or skip the browser setup
Make one GET request. Replace YOUR_API_KEY with your ScreenshotNeo key:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say which page verdict and billing status applied. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
11. Frequently asked questions
Can an LLM-generated test suite prove code is correct?
No. Tests check selected behavior against assertions. Their value depends on whether the cases and expected outcomes represent the requirement and expose relevant faults.
Should I replace a conventional test generator with an LLM?
The cited study evaluates particular LLMs, prompts, classes, and measures. It does not establish that one approach should replace another across projects. Compare candidates on your codebase using correctness, readability, coverage, and bug detection.
How many repeated runs should an LLM application test use?
The sources do not establish one universal count. Choose a count appropriate to the variability and risk you need to assess, and report it with the model and configuration used.
Does mutation testing measure every kind of test quality?
No. It measures detection of the selected mutations. It does not replace review of requirements, readability, coverage, or whether the mutations represent realistic faults.
Research references
- TESTEVAL, Findings of NAACL 2025.
- Aalto research record, ASE 2024 unit-test generation evaluation.
- MuTAP, Information and Software Technology, 2024.
- TiCoder, Microsoft Research.
- Testing LLM-enabled systems taxonomy, 2025.
- Software-engineering perspective on testing LLMs as components, 2024.
- Research roadmap for testing LLM-based software, 2025.
- ISSTA 2024 candidate-program selection abstract.


