How AI Is Used in Quality Engineering
AI can assist with requirements, test design, automation, and reporting. Learn how to evaluate its output and test AI-enabled systems by risk.
AI is used in quality engineering in two distinct ways: teams use AI to assist testing work, and teams test products that contain AI. Generative AI can help analyze requirements, draft tests, assist with automation, and summarize results. Those outputs are proposals to review, not proof that software is correct. When the product itself includes AI, testing should address risks in the model, data, and use context as well as conventional software behavior.
This distinction shapes the workflow: keep requirements and risk visible, check generated artifacts against them, and collect evidence from actual execution. This guide covers practical uses, a reviewable workflow, testing AI-enabled systems, measurement, and common failure modes.
1. AI for testing and testing AI are different
| Approach | What it means | Example question |
|---|---|---|
| AI for testing | Using AI tools to assist quality activities. | Can a model suggest test scenarios from this acceptance criterion? |
| Testing AI | Evaluating an AI component or product that uses AI. | Does the model behave acceptably on representative inputs in its intended context? |
A team may do either without doing the other. Generated test ideas can be wrong even when the application under test is ordinary deterministic software. An AI-enabled product may also produce variable outcomes or depend on training and input data, so its risks may call for data, model, and use-context evaluation in addition to ordinary functional testing. ISO/IEC TS 42119-2:2025 applies the ISO/IEC/IEEE 29119 testing series to AI systems and components through a risk-based approach. ISO/IEC TS 42119-2:2025
2. Where AI can help quality engineers
Requirements and acceptance criteria
Ask a model to identify ambiguous terms, assumptions, missing actors, boundary conditions, and questions that need stakeholder decisions. It can restate a requirement or suggest scenarios, but product owners and domain experts must confirm intended behavior and business rules.
Test design and test data ideas
AI can draft candidate positive, negative, boundary, and state-transition tests from requirements. Review each candidate for a clear precondition, action, expected result, meaningful coverage, duplication, feasibility, and traceability to a requirement or risk. Treat synthetic test data as a proposal: check its validity, privacy, and suitability before using it.
Automation authoring and maintenance
A model can translate a test description into a candidate script, explain existing automation, suggest refactoring, or help locate a likely change after an interface update. Run the script, review assertions and selectors, and inspect whether it tests the intended behavior. A plausible script can encode the wrong expected result or pass without checking the important outcome.
Test execution, defect triage, and reporting
AI can summarize logs, cluster similar failures, draft a defect report, or turn execution artifacts into a status summary. Verify every conclusion against the underlying logs, environment, test data, and artifacts such as screenshots. A summary that omits a failure or confuses a product defect with an environment problem is not release evidence.
Continuous improvement
AI can help identify repeated failure patterns or propose improvements to a regression suite. Compare proposals with an agreed baseline and check whether they improve useful coverage or reduce maintenance. More generated cases do not automatically mean better coverage or lower risk.
These are assistance patterns, not a transfer of quality accountability to a model. ISTQB’s CT-GenAI syllabus covers prompt engineering, evaluation of generated outputs, and applying generative AI through the testing lifecycle. ISTQB CT-GenAI resources
3. A reviewable workflow for AI-assisted testing
- Start from a source of truth. Provide the requirement, acceptance criteria, relevant constraints, and the intended audience for the result. Do not ask the model to infer business rules that are not documented.
- Ask for bounded outputs. Request a small set of scenarios in a defined structure, with assumptions and unresolved questions separated from test steps.
- Review against requirements and risks. Remove unsupported expectations, duplicates, infeasible cases, and cases without a meaningful oracle. Ask a domain owner to settle ambiguities.
- Link accepted tests to their source. Preserve requirement or risk identifiers so reviewers can see why each test exists and what it covers.
- Execute and inspect evidence. Run automation in the target environment. Review failures, logs, and artifacts; do not treat generated explanations as a substitute for execution evidence.
- Measure before expanding use. Compare the AI-assisted workflow with the existing one using agreed quality and cost measures.
Example prompt for candidate test cases
Given the requirement and acceptance criteria below, draft up to 8 candidate tests.
For each test include: requirement reference, risk addressed, preconditions,
actions, expected result, and any assumptions.
Separate unresolved questions from test cases. Do not invent business rules.
Requirement: [paste reviewed requirement]
Acceptance criteria: [paste acceptance criteria]
Review the result line by line. The prompt structure can improve clarity, but it cannot guarantee correctness, completeness, or coverage. Keep the accepted tests in the team’s normal review and version-control process.
4. How to test an AI-enabled system
Choose test work from the system’s risks and requirements. ISO/IEC TS 42119-2:2025 describes identifying risks, considering likelihood and consequence, prioritizing exposure, and selecting test treatments. It also treats requirements as an important input to a risk-based strategy. The relevant test level, type, design technique, review, and coverage measure depend on the system and risk; there is no single checklist that fits every model.
- Describe intended use and boundaries. Identify users, inputs, outputs, operating context, dependencies, and what the system must not be relied on to do.
- Identify risks and requirements. Consider potential harm or business impact, likelihood, affected groups, data concerns, and conventional software failure modes. Record assumptions and prioritize exposure.
- Select suitable evidence. Depending on risk, this can include functional tests, model-level testing, data-representativeness checks, static review, non-functional tests, or continuous testing where behavior may change in production.
- Choose meaningful coverage measures. Match the measure to the risk and test design. A count of cases alone does not show that important data, behaviors, or failure conditions were covered.
- Reassess after change. Changes to the model, data, prompts, surrounding software, or use context can change risk. Decide what needs to be repeated and what new evidence is required.
ISO’s overview says: “Risk-based testing (RBT) is a core concept in the ISO/IEC/IEEE 29119 series, which expects risks to be used as the prime driver for determining the test approaches included in the test strategy and therefore the consequent software testing.” (ISO/IEC TS 42119-2:2025, section 5.4.) Read the ISO overview
Standards and guidance
- ISO/IEC TS 42119-2:2025 is a published overview of applying the testing series to AI systems and components.
- ISO/IEC TS 25058:2024 provides guidance for evaluating AI system quality using an AI system quality model.
- ISO/IEC 25059:2023 is the published edition identified in the research. The second-edition FDIS page is a draft-stage item, so verify its current status before describing it as published: ISO second-edition FDIS listing.
- NIST AI Resource Center provides AI risk management and testing, evaluation, verification, and validation (TEVV) resources. NIST describes the AI Risk Management Framework as voluntary and says version 1.0 is being revised.
5. Evaluate whether AI assistance helps
Set a baseline before expanding adoption. Compare work on outcomes relevant to the team, such as:
- How many suggested tests reviewers accept as useful, and how many require correction or removal.
- Requirement and risk coverage, including important gaps found during review.
- Time spent reviewing and correcting generated tests, scripts, or reports, not only time spent drafting them.
- Defects found and defects that escape, interpreted with the limits of the comparison.
- Automation maintenance burden and whether tests remain understandable and reliable.
- Traceability and the quality of execution evidence available for release decisions.
These are suggested measures, not published performance claims. A 2025 secondary study mapping industry-context research reported that proposed use cases outnumbered actual implementations and observed benefits in the literature it reviewed. That qualifies the evidence; it does not establish that no organizations use AI for testing. Avoid universal adoption or productivity claims. 2025 study on AI adoption in software testing
6. Common failure modes and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Tests contain unsupported business rules | The prompt left gaps and the model filled them with plausible assumptions. | Provide approved criteria, require assumptions and questions to be separated, and get stakeholder confirmation. |
| Many tests repeat the same behavior | Generation optimized for quantity without a coverage model. | Deduplicate by requirement, risk, state, and expected result; request a bounded set and review gaps. |
| An automation script passes but misses the defect | Assertions check incidental UI details or no meaningful outcome. | Review the oracle and assertion against the acceptance criterion; execute with known failure cases where appropriate. |
| A report says a run passed while logs show failures | The generated summary omitted or misclassified evidence. | Check the source run and artifacts; keep summaries as drafts until verified. |
| AI-system testing focuses only on ordinary functions | Model, data, or use-context risks were not included in planning. | Revisit risk identification and choose model, data-representativeness, functional, review, or continuous testing as warranted. |
| Teams cannot show why a test exists | Generated tests were not linked to requirements or risks. | Retain source references and review history with each accepted test. |
| Reported efficiency gains disappear after rollout | Review, correction, maintenance, and failure costs were excluded. | Measure the full workflow against a baseline and include rework and ongoing maintenance. |
7. Capture visual evidence for web quality work
For web testing, screenshots can help document a rendered state, compare an expected result, or attach context to a defect report. A screenshot is one artifact: it does not prove that an interaction, accessibility requirement, backend behavior, or full test passed. Capture the relevant state reproducibly, record the environment and viewport, and keep the associated test result and other evidence together.
DIY: capture a page with a browser
For a manual capture, open the page in a browser at the target viewport, wait for the state under test, and save a screenshot. For repeatable automation, use a browser automation tool already approved in your project and explicitly set the viewport and wait condition. Validate that the capture shows the intended state rather than a loading frame or consent overlay. A browser capture is useful when you need interactive setup or a local environment.
8. Or skip the browser setup
For a URL that can be reached by the API, ScreenshotNeo takes a screenshot with one GET request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the shot was billed. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo and the API docs. Sign up for 1,000 free screenshots a month, no card required.
9. Reliability, performance, and cost considerations
- AI-assisted work: include review and correction time in estimates. A fast draft that takes substantial effort to validate may not reduce total effort.
- Repeatability: generated content can vary. Preserve the prompt, source requirements, accepted output, reviewer, and any relevant tool or model version when reproducibility matters.
- Test reliability: flaky automation undermines evidence. Use explicit state and wait conditions, stable assertions, and inspect failures before attributing them to the product.
- AI-system change: model or data updates can alter behavior. Risk-based plans should define triggers for reassessment and repeat testing.
- Privacy and security: follow organizational rules before sending requirements, logs, customer data, source code, or test data to an AI service. Minimize sensitive content and use approved environments.
- Evidence cost: balance coverage depth and frequency against the impact of missed failures. Select test levels and evidence based on risk rather than maximizing test count.
10. Frequently asked questions
Does using AI mean a team needs fewer testers?
The cited research does not support a general conclusion about staffing. AI may assist tasks, while teams still need people to validate requirements, judge risk, review outputs, and make accountable release decisions.
Can AI-generated tests replace a test oracle?
No. A test needs a defensible expected result or other evaluation criterion. A model can suggest one, but the team must verify it against requirements and domain rules.
Is NIST AI RMF mandatory?
NIST describes the AI Risk Management Framework as voluntary guidance. Whether other obligations apply depends on the organization and context.
Which standard should a team read first?
For testing AI systems, start with ISO/IEC TS 42119-2:2025’s overview of risk-based application of the testing series. For AI system quality evaluation guidance, consult ISO/IEC TS 25058:2024. Select based on the question and verify publication status when citing standards.
Conclusion
AI can assist quality engineering across analysis, test design, automation, execution reporting, and improvement. The dependable pattern is to keep people accountable for requirements and risk, review AI-generated work, and base conclusions on execution evidence. When the product contains AI, plan tests around its model, data, and operating context as well as ordinary software behavior. Use risk and measurable outcomes to decide where AI assistance and additional testing belong.


