How AI Is Used in Software Testing
See where AI assists software testing, how to evaluate AI-powered systems, and why human review and risk-based testing still matter.
AI is used in software testing to help draft test cases, generate text-based test data, prepare reports, and augment testing automation. Separately, software systems that use AI need to be tested as software systems, with approaches chosen according to risk. In both cases, generated output is an aid for testers to review—not proof of coverage, correctness, or quality.
That distinction matters: using AI to test software describes AI assisting a testing workflow; testing AI systems means evaluating a product or component that uses AI. The same project may involve both.
1. Where AI fits in software testing
AI can help with specific tasks in the test lifecycle. Applause’s 2025 survey of more than 4,400 independent software developers, QA professionals, and consumers worldwide reported these top AI uses among QA professionals:
| Reported use | Respondents reporting it | What a tester still needs to check |
|---|---|---|
| Test case generation | 66% | Whether each case traces to a requirement or risk, includes relevant boundary conditions, and asserts the intended behavior. |
| Text generation for test data | 59% | Whether data is valid for the scenario, covers useful edge cases, respects privacy constraints, and is safe to use in the target environment. |
| Test reporting | 58% | Whether the report matches observed results, distinguishes failures from environment problems, and preserves enough detail to reproduce an issue. |
These are Applause survey findings, not universal usage rates or evidence that AI improved software quality or testing speed. The survey reports uses; it does not establish that generated cases are complete or correct. [Applause’s 2025 AI survey release]
Generating test cases
A tester can give an AI tool a requirement, interface description, or existing test and ask for candidate scenarios. The useful output is a reviewable draft: for example, normal input, missing input, invalid input, permission boundaries, and state transitions. Compare each proposed case to the acceptance criteria and risk model. Remove duplicates, add omitted conditions, and define an observable expected result before adding a case to a test suite.
Do not infer coverage from the number of generated cases. A long list can miss a high-impact failure mode, repeat the same scenario in different words, or assert behavior the product was never meant to have.
Generating test data
AI can draft textual fixtures or representative examples for test scenarios. Specify the format, constraints, and purpose. Review that the data is syntactically valid and exercises the intended conditions. Avoid submitting personal, confidential, or production data unless the tool and your organization’s policies explicitly permit that processing. Use synthetic or approved test fixtures where possible.
Generated data can be plausible but invalid, internally inconsistent, or biased toward common cases. Validate it with the same schema, privacy, and business rules as other test data.
Drafting test reports
An AI tool can turn structured test output or tester notes into a summary. Keep the underlying logs, test identifiers, environment, and reproduction steps available. Check that the summary does not turn a timeout into a product defect, omit a failed assertion, or describe an unobserved result as fact. Treat the report as a draft and retain a trace back to the execution evidence.
Augmenting test automation
Some tools use AI to assist automation work or its maintenance. Fit the tool to the task—case creation, data, reporting, execution, or evaluation—and assess how generated changes are reviewed and maintained. A 2025 literature review describes test automation as requiring substantial design, development, maintenance, and evolution effort, while discussing AI augmentation at different levels of automation. AI assistance does not remove that lifecycle work. [Schieferdecker’s 2025 literature review]
2. How to put AI assistance into a test workflow
- Choose a bounded task. Start with one repeatable activity, such as drafting cases from a stable requirement or summarizing a structured test run. Define what a useful result looks like before choosing a tool.
- Provide context and constraints. Include the relevant requirement, supported inputs, expected outputs, and exclusions. Do not include secrets or sensitive data unless policy and the tool’s handling terms allow it.
- Ask for reviewable output. Request cases with a scenario, setup, action, and expected result, or a report that cites test IDs and observed outcomes. Keep the source material alongside the draft.
- Review against requirements and risk. Check traceability, edge cases, privacy, correctness, and whether the output validates the intended behavior. Add tests the model missed and discard unsupported assumptions.
- Run through the established test process. Execute approved tests in the right environment, record results, investigate failures, and maintain the suite as product behavior changes.
- Evaluate with local evidence. Track whether the tool’s output is accepted, corrected, or discarded for the task. Compare it on your own workflow; do not treat vendor claims or survey responses as a controlled result for your team.
3. How to test software that uses AI
Testing an AI-enabled system is a separate activity from asking AI to help write tests. ISO/IEC TS 42119-2:2025 describes applying established software testing processes to AI systems and components through a risk-based approach. It addresses risk identification, test approaches, and documentation, and connects to the ISO/IEC/IEEE 29119 software testing series. The public standard page describes its scope; the complete standard may require access. [ISO/IEC TS 42119-2:2025]
A practical evaluation plan starts with the system’s intended use and potential harms. Identify what the system receives, what it produces, who relies on the output, and what happens when it is wrong or unavailable. Then choose tests that produce evidence about those risks. There is no single protocol in the cited sources that applies to every AI product.
- Behavior and output quality: define representative inputs and review outputs against task-specific acceptance criteria. Include ambiguous, unusual, and boundary inputs that matter to the intended use.
- Prompt and response evaluation: Applause reports prompt and response grading as a human-involved AI testing activity. Define a grading rubric and use qualified reviewers for judgment calls; preserve examples and rationale.
- User experience: Applause also reports UX testing among human-involved activities. Evaluate whether people can understand the system’s role, interpret outputs, and recover when it fails, using scenarios tied to the product’s users.
- Accessibility: Accessibility testing is another activity Applause lists. Include accessibility requirements and appropriate human evaluation; do not assume a plausible AI response means the interface is accessible.
- Failure and fallback behavior: test what the surrounding software does when outputs are missing, delayed, malformed, or unsuitable. Check that safeguards and escalation paths work as intended.
- Documentation and traceability: connect test results to risks, requirements, system versions, and the test approach so reviewers can understand what was evaluated.
Applause’s 2025 survey reported 61% for prompt and response grading, 57% for UX testing, and 54% for accessibility testing among its listed human-involved AI testing activities. These percentages describe that survey’s respondents, not a universal test standard. [Applause’s survey findings]
4. What adoption figures do and do not show
Katalon’s State of Software Quality Report 2025 says 76% of respondents used AI-powered tools in software testing activities, and 56% of QA teams still struggled to keep up with testing demands. These are findings reported on Katalon’s report page. The accessible page does not establish a population-wide adoption rate, and the figures do not show that AI use caused, prevented, or failed to prevent testing workload challenges. [Katalon’s 2025 report]
Applause also reported that respondents saw productivity potential, but the sources in this guide do not provide a controlled before-and-after estimate for speed or software quality. Treat survey results as descriptions of respondent reports, not causal evidence. Gartner’s public abstract for its 2024 Market Guide says the market is evolving and flags security and legal risks; the full vendor analysis is access-restricted, so it does not support a vendor ranking here. [Gartner’s public abstract]
5. Evaluate an AI testing tool before adopting it
| Decision area | Questions to answer |
|---|---|
| Task fit | Does it address the actual need: test cases, data, reports, execution or automation, or evaluation of AI outputs? |
| Coverage and control | Can the team map output to requirements, risks, and edge cases? Can a qualified person review and approve changes? |
| Integration and maintenance | How does it fit the current test process? Who maintains generated tests and handles changes as the application evolves? |
| Security and legal handling | What information is sent to the tool, how is it handled, and do those controls meet organizational and legal requirements? Gartner’s public abstract specifically flags security and legal risks. |
| Evidence | Are claims based on vendor statements, survey self-reports, or results observed on your own systems? What evidence would justify continued use? |
These are decision criteria for an internal evaluation, not a ranking of products. The available sources do not support a vendor-by-vendor recommendation.
6. Reliability, performance, and cost considerations
- Reliability: Generated content can be incomplete or wrong. Keep deterministic checks, test execution evidence, and human review for decisions with meaningful risk. Record the inputs and tool version where needed to explain a result later.
- Performance: The cited sources do not establish a general speed improvement. Measure the full task in your environment, including review, correction, integration, and maintenance time—not just generation time.
- Cost: Compare tool and usage costs with the effort required to review outputs, integrate them, protect data, and maintain resulting tests. The source set provides no general return-on-investment figure.
- Change management: AI-assisted tests and AI-enabled products can change over time. Revisit cases when requirements, interfaces, models, prompts, or operating conditions change, and retain enough documentation to interpret comparisons.
7. Troubleshooting common problems
| Symptom | Likely cause | Fix |
|---|---|---|
| Generated tests look complete but miss important failures | The prompt or source requirement omitted risks, constraints, or boundary conditions. | Map cases to requirements and risk scenarios; add missing boundaries and negative cases, then have a tester review coverage. |
| Many generated cases are duplicates | The request asked for volume without defining distinct behaviors or expected outcomes. | Specify scenario categories and ask for unique conditions; deduplicate against the existing suite before adoption. |
| Test data is plausible but fails validation | The output did not capture schema, cross-field, or business constraints. | Provide approved format rules and validate generated fixtures with the application’s normal validators. |
| A report states a failure inaccurately | The summary inferred a cause from incomplete logs or confused infrastructure errors with product behavior. | Require references to test IDs and observed evidence; verify every conclusion against logs and rerun when appropriate. |
| AI output varies between runs | The task permits more than one acceptable response or the system’s output is variable. | Define a rubric and acceptable ranges, retain representative cases, and evaluate behavior against risk-specific criteria rather than a single exact string. |
| Teams cannot tell whether the tool helped | There is no baseline or the evaluation counts generation while ignoring review and maintenance. | Choose a bounded task, record the existing workflow, and compare total effort and defect-relevant evidence on local work without assuming causation from surveys. |
| Sensitive information is exposed to a tool | Data handling was not assessed before sending prompts, code, or fixtures. | Stop sending the information, follow incident and data policies, and assess tool controls and approved data boundaries before resuming. |
8. Using screenshots in visual QA
For interface testing, screenshots can provide visual evidence of a rendered page or component across routes and viewport sizes. A browser automation setup can capture screenshots for comparison, but the capture itself does not decide whether a visual difference is a defect; define expected states and review differences in context.
To make a browser screenshot with Playwright in Node.js, install Playwright and its Chromium browser with npm install -D playwright and npx playwright install chromium. Save the following as capture.mjs, then run node capture.mjs https://example.com:
import { chromium } from 'playwright';
const url = process.argv[2];
if (!url) throw new Error('Pass a URL, for example https://example.com');
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
For repeatable visual checks, pin the viewport and browser version, wait for the application state you intend to capture, use stable test data, and keep authentication and secrets out of source control. A full-page screenshot can be large; element captures can be more focused. Dynamic content, animation, fonts, and third-party requests can create noise, so control or disable them in the test environment where possible.
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its API returns a screenshot or PDF from one GET request; the API supports capture settings such as full-page capture, element selection, viewport and device presets, waits, custom CSS and JavaScript, and request blocking. See the ScreenshotNeo API documentation.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, and failed loads are never billed; response headers say the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
10. FAQ
Does using AI for testing mean the software itself uses AI?
No. AI may assist a conventional test workflow, while an AI-enabled product is software that needs to be tested. A team can do either or both.
Do survey adoption numbers prove AI improves quality?
No. The cited figures describe survey respondents’ reported use or activities. The source set does not provide a controlled causal estimate for quality or speed gains.
Can AI-generated tests replace a QA review?
No. Reviewers still need to determine whether cases reflect requirements, risks, valid data, and observable expected behavior.
Is there one standard test protocol for every AI system?
The cited ISO/IEC technical specification describes a risk-based application of established software testing processes. The appropriate tests depend on the system, intended use, and risks.
Sources
- ISO/IEC TS 42119-2:2025, Testing of AI systems — Part 2: Guidelines for testing AI systems and components.
- Applause, 2025 AI Survey results.
- Katalon, State of Software Quality Report 2025.
- Gartner, Market Guide for AI-Augmented Software-Testing Tools, public abstract.
- Ina K. Schieferdecker, “Navigating the growing field of research on AI for software testing,” 2025 preprint.


