How Generative AI Can Improve QA Testing
Generative AI can draft tests and surface gaps, but useful QA depends on clear requirements, reviewed assertions, and running tests in the real environment.
Generative AI can help QA teams draft tests, expand scenarios, and analyze failures. It works best as an assistant: give it clear requirements, relevant code, and existing test conventions; review every assertion; then run the tests in the project’s real environment. A test that passes can still check the wrong behavior.
This guide covers practical uses, a runnable Python example, a review workflow, risks, and troubleshooting. It focuses on testing software with AI assistance, not testing the safety or correctness of an AI model itself.
What generative AI can improve in QA
AI tools can turn source code, specifications, and existing tests into candidate test cases. They can also suggest boundary conditions, help explain failures, and point to scenarios missing from a suite. These are drafting and analysis tasks: execution and QA judgment remain essential.
| QA task | Useful assistance | What a reviewer must verify |
|---|---|---|
| Test drafting | Propose unit tests from a function, contract, or requirement. | Tests compile, follow project conventions, and assert required behavior. |
| Scenario expansion | Suggest boundaries, invalid inputs, state transitions, and combinations. | Cases are relevant, feasible, and not duplicates of existing coverage. |
| Failure analysis | Summarize an error and suggest likely causes or follow-up cases. | The explanation matches the failure and is confirmed against code and requirements. |
| Test maintenance | Suggest updates when an interface or behavior changes. | Updates preserve the intended contract instead of merely accommodating implementation changes. |
A 2025 practitioner playbook describes test generation, continuous feedback, failure analysis, prototyping, and varied-user or condition simulation as potential applications. These are workflow possibilities, not controlled estimates of time saved or defects prevented. IEEE Computer practitioner playbook
Why specifications and code context matter
A model needs enough context to distinguish intended behavior from an implementation detail. Provide the relevant function, its callers or types where they affect behavior, written requirements, existing tests, and the test framework conventions. State what should happen for valid, invalid, and boundary inputs.
Google Research’s 2026 evaluation on production bugs found that an agent which first documented preconditions, postconditions, and undefined behavior improved bug detection by 9.8 percentage points and branch coverage by 2.5 percentage points versus its traditional test-generation agent baseline. The study also reported that spec-driven suites were judged superior in 77.8% of cases versus baseline suites and 56.7% versus human-authored tests using an LLM-as-a-Judge. Those evaluator preferences do not prove that AI universally writes better tests, and the measured improvements belong to that study’s setup. Google Research study
A practical workflow for AI-assisted test generation
- Start with the contract. Write down the behavior users or callers rely on. Identify preconditions, postconditions, edge cases, and any undefined behavior.
- Supply focused context. Include the target code, relevant specification, nearby types or dependencies, existing tests, and the command used to run them. Remove secrets and unrelated data.
- Ask for a test plan first. Have the assistant list cases and expected outcomes before it writes test code. Correct misunderstandings early.
- Generate a small set. Ask for readable tests in the project’s actual framework, with one behavior per test and explicit expected results.
- Review the assertions. Trace each assertion to a requirement. Reject plausible guesses that are not specified behavior.
- Run tests in the real environment. Check imports, fixtures, versions, data, and side effects. Confirm that failures indicate real defects rather than setup errors.
- Check effectiveness and gaps. Review branch coverage and, where practical, introduce a known defect or mutation to see whether relevant tests fail. Coverage alone does not establish test quality.
- Repeat for variable behavior. For AI-backed or otherwise nondeterministic features, try varied inputs and repeated runs; evaluate behavioral criteria or distributions instead of a single pass/fail result.
Runnable example: generate and review Python tests
This example uses only the Python standard library and unittest. It demonstrates the reviewable output to request from an AI assistant; no model SDK or paid service is required. Save the files in one directory and run python -m unittest -v.
1. Define behavior in code
# pricing.py
def total_cents(unit_price_cents: int, quantity: int) -> int:
"""Return the total in cents; price and quantity must be nonnegative."""
if unit_price_cents < 0:
raise ValueError("unit price must be nonnegative")
if quantity < 0:
raise ValueError("quantity must be nonnegative")
return unit_price_cents * quantity
2. Ask for a test plan before code
Given pricing.py and the contract in its docstring, list test cases before writing code.
Include zero price, zero quantity, an ordinary multiplication, and each negative input.
Do not infer behavior beyond the stated contract. Then write unittest tests that follow
these cases, with expected values derived from the contract. Explain any ambiguity.
3. Review and run the resulting tests
# test_pricing.py
import unittest
from pricing import total_cents
class TotalCentsTests(unittest.TestCase):
def test_multiplies_price_by_quantity(self):
self.assertEqual(total_cents(250, 3), 750)
def test_zero_price_returns_zero(self):
self.assertEqual(total_cents(0, 3), 0)
def test_zero_quantity_returns_zero(self):
self.assertEqual(total_cents(250, 0), 0)
def test_negative_price_is_rejected(self):
with self.assertRaisesRegex(ValueError, "unit price"):
total_cents(-1, 3)
def test_negative_quantity_is_rejected(self):
with self.assertRaisesRegex(ValueError, "quantity"):
total_cents(250, -1)
if __name__ == "__main__":
unittest.main()
Each assertion corresponds to behavior the example explicitly states. If the real requirement instead allows negative values, caps quantities, or uses a different error contract, change the specification and tests accordingly. A generated test should not silently decide those product requirements.
Review checklist for generated tests
- Does each test map to a documented requirement or an agreed quality property?
- Are expected values independently derived, rather than copied from the implementation?
- Do tests cover boundaries, invalid inputs, and important state or dependency interactions?
- Are mocks limited to external boundaries, so the test still checks meaningful behavior?
- Does the suite pass for the right reason and fail when a relevant defect is introduced?
- Can another developer understand and maintain the tests without the prompt?
- Were secrets, personal data, and proprietary context handled according to the organization’s policy?
A 2024 study of 290 Copilot-generated tests for 53 sampled tests from open-source Python projects found 45.28% passed within an existing test suite; 54.72% were failing, broken, or empty. Without an existing suite, 92.45% were failing, broken, or empty. These figures describe that study’s sample and setup, not universal or current Copilot performance. They reinforce the need to execute and review generated code. TU Delft study record
Testing AI features and nondeterministic behavior
When the software under test itself produces variable outputs, one fixed expected string may be a poor oracle. Define stable behavioral criteria, use a representative range of inputs, repeat runs where variation matters, and track distributions or rates against an agreed threshold. Record the model or configuration and test data needed to reproduce a result. The practitioner playbook discusses nondeterminism and cautions that generated assertions can be incorrect or biased. IEEE Computer practitioner playbook
For ordinary deterministic code, do not weaken exact assertions just because an AI drafted them. Match the oracle to the contract: exact values for exact contracts, invariant or range checks only where the requirements allow them.
Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Generated test does not import or compile | The assistant guessed module names, framework versions, or project layout. | Provide the real file tree, imports, and a neighboring test; request code for the installed framework version and run it locally. |
| Test fails because a fixture or service is missing | Environment setup and dependencies were omitted from context. | Share the test command and fixture pattern, or keep the test isolated with the project’s established test doubles. |
| Test passes but a bug remains | The assertion mirrors current implementation or checks the wrong property. | Compare it to the requirement, derive expected results independently, and try a relevant mutation or known defect. |
| Test fails on every run | Generated expectations may be wrong, or the environment differs from assumptions. | Separate setup failures from behavior failures; verify the contract and actual inputs before editing production code. |
| Flaky test | Timing, shared state, randomness, network dependencies, or nondeterministic output. | Control seeds and clocks where appropriate, isolate state, avoid arbitrary sleeps, and use repeated-run criteria for intentionally variable behavior. |
| Many tests duplicate existing coverage | The model lacked the current suite or coverage context. | Include nearby tests and ask it to identify uncovered behaviors before drafting additions. |
| Generated test encodes an unstated requirement | The prompt left behavior ambiguous and the model filled the gap. | Mark the case as a product decision. Clarify the contract before adding an assertion. |
Performance, reliability, and cost
Generation can reduce the effort of producing a first draft, but the cited practitioner guidance does not establish a general time saving. The actual cost includes providing context, reviewing output, running it, repairing setup errors, and maintaining tests as software changes. Keep generated batches small enough to review and prioritize high-risk behavior over raw test count.
Reliability comes from a repeatable execution environment and traceable requirements. Store accepted tests in the same version control and CI workflow as other tests. When prompts or model versions change, review resulting diffs and rerun the suite. For variable systems, capture enough configuration and input information to investigate changing outcomes.
Financial cost depends on the AI tool and its plan; the research here does not establish comparable pricing. Check the provider’s current terms and your organization’s data handling requirements before sending code. Human review and test execution remain necessary costs in any workflow.
Capture reproducible browser evidence for QA
Browser screenshots can make visual regressions and bug reports easier to review. A do-it-yourself workflow can launch a browser, navigate to a test URL, wait for the relevant state, and save a screenshot. Keep the browser version, viewport, test data, and readiness condition consistent so captures are comparable. For automated suites, capture only after the page reaches the state under test; a screenshot of a loading or consent overlay can obscure the result.
Or skip the browser setup
With ScreenshotNeo, a website screenshot API and MCP server for developers, one GET request returns an image or PDF. Its capture options include full-page screenshots with lazy images loaded, CSS element capture, device and viewport settings, dark mode, custom CSS and JavaScript, selector or network-idle waits, and custom headers or cookies. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, and failed loads are not billed, and cache hits cost nothing; response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.
FAQ
Can AI-generated tests be trusted if they pass?
A passing result only shows that the test’s assertions held for that run. Review whether those assertions represent the intended behavior, and where practical verify that a meaningful defect makes the test fail.
Should AI write an entire test suite at once?
Usually, start with a small, specified behavior. Smaller batches are easier to inspect, run, and correct before mistakes spread through a suite.
Is line coverage enough to measure generated test quality?
No. Coverage shows which code ran, not whether assertions would detect a defect or reflect the contract. Combine coverage with assertion review and defect-oriented checks where practical.
Where can I learn more about testing with generative AI?
The German Testing Board lists an English CT-GenAI syllabus, version 1.1 (2026), among its syllabi. The listing establishes the resource’s availability; consult it for its scope and details.


