How AI Is Improving Software Testing and Quality
AI can speed up test drafting, edge-case discovery, and end-to-end testing. Learn where it helps, how to verify its output, and how to measure results.
AI is improving software testing by helping developers draft unit and integration tests, suggest edge cases, scaffold end-to-end tests, and propose debugging or repair ideas. These tools can reduce the effort of getting checks started, but they do not guarantee software quality: run the tests, verify that assertions reflect intended behavior, and review what risks remain uncovered.
The practical model is AI-assisted test preparation followed by deterministic execution and human review. A generated test is useful only if it checks meaningful behavior and would fail when that behavior is wrong.
1. What AI contributes to software testing
AI-assisted testing covers more than generating test code. Research has examined several software-testing activities; a 2023 survey of 102 studies identified test-case preparation and program repair as representative uses. A 2024 systematic review examined 55 AI-based test automation tools, while empirically evaluating two selected tools on two open-source projects. Those findings describe an active and varied field, not universal proof that a particular tool improves every team’s results. 2023 survey; 2024 systematic review.
| Testing activity | How AI can help | What still needs checking |
|---|---|---|
| Unit test preparation | Draft cases, inputs, assertions, and test scaffolding from code or a behavior description. | Whether assertions express the requirement rather than merely repeat the implementation. |
| Edge-case discovery | Suggest boundaries, unusual inputs, and failure conditions a developer can evaluate. | Whether each proposed case is relevant and whether important cases are missing. |
| Integration and end-to-end testing | Generate test scaffolding or scenarios spanning components and user flows. | Whether tests run reliably in the real environment and assert outcomes that matter. |
| Debugging and repair | Explain a suspected defect or propose a code change. | Whether the change is correct, reviewed, and protected by regression tests. |
| Test maintenance | Help update tests when interfaces or workflows change. | Whether a changed test preserves the original intent instead of weakening it. |
Official GitHub documentation describes Copilot assistance for unit and integration test generation. It advises giving more detailed prompts for complex scenarios, reviewing generated tests, and adding tests as needed. Google Cloud described a Firebase App Testing agent intended to generate, manage, and execute end-to-end tests; its announcement said the agents were in preview at that time, so check current availability before relying on that status. GitHub Docs: Writing tests with GitHub Copilot; Google Cloud announcement.
2. How to use AI to draft a test safely
- State the behavior first. Write down the requirement, inputs, expected results, and important failure cases. If the requirement is ambiguous, resolve that before asking for test code.
- Give the assistant relevant context. Include the function or interface, nearby tests, framework conventions, and any constraints the test must preserve. Avoid sharing code or test data with a service unless its handling fits your organization’s rules.
- Ask for cases and rationale. Have the assistant list normal, boundary, and failure scenarios before or alongside the code. Ask it to explain what each assertion proves.
- Review the test as code. Check imports, fixtures, setup, cleanup, naming, determinism, and consistency with the project’s patterns. Look for assertions that merely check a value derived from the same implementation.
- Run it in the project environment. Use the normal local and CI workflow. A generated test that compiles but never runs does not verify the change.
- Check what remains uncovered. Compare tested behaviors against requirements and risk areas. Add focused tests where needed, and do not treat line coverage as a substitute for meaningful checks.
Prompt pattern
Write tests for the behavior described below using this project's existing test framework and conventions.
Behavior:
[State the intended behavior and observable outcomes.]
Relevant code or interface:
[Paste the smallest useful code context.]
Include:
- normal behavior
- boundary inputs
- invalid inputs and expected errors
- relevant state changes or side effects
For each test, explain the behavior it checks and why its assertion would fail if that behavior were broken. Do not change production code. Call out assumptions or missing requirements instead of guessing.
This prompt is a starting point, not a guarantee of complete coverage. For complex cases, provide more specific context and constraints, then review and supplement the result.
3. Measure test quality, not test volume
More generated cases or a higher line-coverage percentage can coexist with weak verification. For each test, ask: does it encode an expected behavior? Can it fail if the behavior is wrong? Does it exercise a meaningful boundary or failure mode? Is its result stable across runs?
To evaluate an AI testing tool, run a bounded pilot on representative work and compare it with your normal process. Track:
- Generated tests accepted, edited, or discarded, and the review time required.
- Defects or regressions caught before release and escaped defects afterward.
- Flaky-test rate, test execution time, and CI feedback time.
- Change failure rate and delivery stability alongside throughput.
- Developer experience and whether the tool fits your language, framework, IDE, and review process.
Keep the baseline and evaluation period clear. A before-and-after change alone does not show that AI caused an improvement if the team also changed its process, platform, or workload. Consider context access, workflow fit, verification, coverage quality, and governance when comparing tools. Confirm vendors’ current terms and organizational approval requirements before using source code or test data with an assistant.
4. What the evidence says about AI and quality
DORA’s 2025 report announcement describes a survey of nearly 5,000 technology professionals and more than 100 hours of qualitative data. It reports that 90% of respondents used AI at work, more than 80% believed AI increased productivity, and 30% reported little or no trust in AI-generated code. These are survey findings, not controlled measurements showing that generated tests are more effective.
The same announcement reports a positive relationship between AI adoption and throughput and product performance, alongside a negative relationship with delivery stability. Treat these as reported associations, not proof that AI directly caused either outcome. DORA emphasizes the surrounding engineering system, including platform quality, clear workflows, team alignment, testing, version control, and fast feedback. As DORA Lead Nathen Harvey put it, “AI doesn’t fix a team; it amplifies what’s already there.” DORA 2025 report announcement.
GitHub’s U.S. 2024 developer survey summary reports that 92% of U.S. respondents used AI coding tools to generate test cases at least some of the time. This is self-reported tool usage, not evidence that those test cases improved test effectiveness. GitHub survey summary (PDF).
5. Failure modes and safeguards
| Failure mode | Why it happens | Safeguard |
|---|---|---|
| Tests mirror a bug in the implementation | The assistant sees code but lacks an independent, precise statement of intended behavior. | Provide requirements and observable outcomes. Review expected values independently. |
| Tests pass without checking much | Assertions can be missing, too broad, or unrelated to the behavior at risk. | For each test, identify the defect that would make it fail. Add assertions for the required outcome. |
| Important boundaries are absent | A short or underspecified prompt may produce only the obvious happy path. | Ask for boundary and invalid-input cases, then compare against requirements and known risks. |
| Generated tests are flaky | Tests may depend on timing, shared state, network services, or unstable data. | Run repeatedly where appropriate, isolate dependencies, control data and time, and use the project’s normal reliability practices. |
| Repair suggestions introduce regressions | A plausible patch may fix one example while breaking another requirement. | Review the diff, run regression tests, and add a test for the original defect. |
| Coverage rises but confidence does not | Line coverage counts execution, not whether assertions verify correct behavior. | Review behavior and risk coverage, not only coverage percentages or generated-test counts. |
6. Troubleshooting AI-generated tests
| Symptom | Likely cause | What to do |
|---|---|---|
| Code does not compile or imports are wrong | The assistant lacks framework or project context, or guessed a version-specific API. | Provide the relevant existing test file and dependency conventions. Verify imports against the project. |
| Test fails immediately for an unrelated reason | Fixture setup, environment assumptions, or test data do not match the repository. | Use established fixtures and run commands. Separate setup failures from behavior assertions. |
| Test passes when expected behavior is broken | Assertions are weak, vacuous, or derived from the implementation. | Write down the expected observable result and strengthen the assertion; confirm with a known failing case where practical. |
| Test passes locally and fails in CI | Timing, ordering, environment, or external dependencies differ. | Control nondeterministic inputs and dependencies, then reproduce with the CI configuration. |
| Assistant produces repetitive happy-path tests | The prompt did not identify boundaries, invalid inputs, or risk areas. | Ask for a case matrix by input class and explicitly request expected outcomes for each class. |
| Suggested fix makes tests pass by weakening them | The assistant optimized for a passing result rather than preserving the requirement. | Review test and production changes separately. Reject assertion removal unless the requirement itself changed. |
7. Performance, reliability, and cost
AI may reduce the time needed to draft tests, but generation adds review work and does not replace test execution. Measure total cycle time, including prompting, editing, review, reruns, and CI feedback. Keep deterministic tests in the same fast feedback loop that protects ordinary changes; reserve slower end-to-end checks for the workflows and risks they cover.
Reliability depends on both the generated test and the test system around it. Track flaky failures, rerun rates, execution time, and escaped defects. Keep human review for security-sensitive behavior, release decisions, and assertions tied to important requirements. DORA’s findings support attention to testing and fast feedback when AI makes it easier to increase change throughput.
Tool pricing and data-handling terms vary and can change; verify current vendor details before a purchase or rollout. Include inference or subscription costs, CI compute, reviewer time, and maintenance in a pilot’s cost picture. A tool that creates many low-value tests can cost more to review and maintain than it saves in drafting time.
8. Capture website states as test evidence
For browser-based test workflows, screenshots can record what a page looked like at a particular point in a run. A screenshot is useful evidence for visual changes, but it does not replace semantic assertions, accessibility checks, or tests of application behavior. If screenshots are part of your testing workflow, ScreenshotNeo is a website screenshot API and MCP server for developers.
DIY browser capture with Playwright
Install Playwright and its Chromium browser, then save this as capture.mjs. Run node capture.mjs https://example.com. It writes a full-page PNG to page.png.
import { chromium } from 'playwright';
const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
For repeatable captures, use a stable test environment, fixed viewport, controlled test data, and a deliberate readiness condition. networkidle may not occur on pages with persistent network activity; in that case wait for a meaningful selector or application signal and choose an explicit timeout. Full-page captures can be large and may trigger lazy-loaded content differently from a viewport shot. For visual regression, account for dynamic content such as timestamps and animations.
9. Or skip the browser setup
ScreenshotNeo can return a screenshot from one GET request. See the API documentation. This runnable cURL example saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners are accepted like a visitor and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card.
10. FAQ
Can AI improve software quality?
It can help teams prepare checks and get feedback faster. Whether quality improves depends on test adequacy, execution, review, and the engineering workflow around the tool.
Are AI-generated tests reliable?
They are drafts that need validation. Reliability comes from correct assertions, repeatable execution, and coverage of relevant behaviors and risks.
Should AI-generated tests be merged automatically?
Use the same review and CI controls as other test code. Teams may automate routine checks, but should inspect whether generated tests preserve intended behavior and project conventions.
Does more test coverage mean better quality?
No. Coverage indicates which code ran; it does not establish that tests would detect incorrect behavior.
Sources
- Software Testing with Large Language Models: Survey, Landscape, and Vision.
- AI-powered test automation tools: A systematic review and empirical evaluation.
- GitHub Docs: Writing tests with GitHub Copilot.
- Visual Studio Code Docs: Testing code.
- Google Cloud: An application-centric, AI-powered cloud.
- DORA 2025 report announcement.
- GitHub: AI and automation: advancing code security (PDF).


