ScreenshotNeo

BlogEngineering

Examples of Generative AI in Software Testing

See how generative AI can draft tests, refine them with execution feedback, suggest repairs, and help assess defects—plus how to verify the results.

By the ScreenshotNeo team4 October 20267 min read

Generative AI can help prepare test cases from code or requirements, suggest a repair after a test fails, refine tests using execution feedback, assess outputs, and identify likely defects in source code or binaries. Treat its output as a candidate artifact: run it, check whether it reflects intended behavior, and review any proposed code change. Research surveys and reviews identify these as active task categories, but do not establish a universal accuracy or productivity figure. A 2024 software-testing survey covers test preparation and program repair; a 2025 review also discusses feedback guidance, output assessment, and static defect detection.

What generative AI does in software testing

In a testing workflow, a generative model produces or revises an artifact from supplied context. That artifact might be a test scenario, executable test code, a suggested patch, or an explanation of a possible defect. The model does not establish by itself that the requirement is correct, that a test is meaningful, or that a proposed fix is safe.

The examples below are task categories and study approaches described in the research dossier, not guarantees that any model or tool will succeed on a project.

1. Draft test cases from code or requirements

Provide a function, a user story, or a structured requirement and ask for candidate cases. Starting from code can surface boundary conditions visible in the implementation. Starting from a requirement can keep attention on intended behavior rather than only the current implementation. A 2025 preprint studies high-level test generation from business requirements and model evaluation; treat its results as specific to that study, not settled industry evidence. Read the preprint record.

Example prompt

Given the requirement and function below, propose test cases only.
For each case, include: name, input, expected result, and the requirement it checks.
Include normal, boundary, invalid, and empty-input cases.
Do not assume behavior that the requirement does not define.

Requirement: [paste the requirement]
Function: [paste the relevant function]

Review the cases against the specification before turning them into executable tests. A model can confidently invent behavior when the requirement is ambiguous. Record unresolved decisions, such as whether a missing value should be rejected or defaulted, and ask the product owner or specification author to settle them.

2. Generate executable tests as candidates

Once cases are agreed, ask for tests in the project’s existing framework and style. Supply relevant fixtures, dependency versions, and conventions, but avoid sending secrets or unrelated proprietary context. Compile and run the generated tests, then inspect their assertions: a test that merely calls a function without checking a meaningful outcome can pass while providing little protection.

Practical review checklist

  • Does every test correspond to an explicit requirement or an intentional defensive behavior?
  • Do expected values come from the specification rather than being copied from the implementation?
  • Are boundary and invalid inputs represented where relevant?
  • Do assertions fail when the behavior under test is wrong?
  • Are shared fixtures isolated so one test cannot affect another?
  • Can a maintainer understand and update the test without relying on the prompt?

3. Suggest a repair after a test failure

Program repair is another representative LLM-assisted testing task in the 2024 survey. A useful workflow gives the model a failing test, the failure output, and a narrowly relevant code region, then asks for a proposed change and its rationale. The survey identifies this task category; it does not imply that suggested repairs are correct.

  1. Reproduce the failure with the project’s normal test command.
  2. Check that the failing assertion represents intended behavior.
  3. Ask for the smallest plausible code change and an explanation tied to the failure.
  4. Review the diff for unrelated changes, security implications, and compatibility concerns.
  5. Run the focused test and the relevant broader suite; add a regression case if needed.

Do not accept a patch solely because it makes the visible test pass. It might weaken an assertion, special-case one input, or break an untested path.

4. Refine tests with execution feedback

Dynamic workflows can use execution feedback to guide test generation or evaluate output. For example, run a candidate test, inspect whether it compiles and what it exercises, then ask for a revision that addresses a specific gap. This is an iterative assistant workflow, not evidence of autonomous reliability.

Give feedback that is concrete: a compile error, an uncovered branch, an unexpected result, or a test that passes after a known behavior change. Keep the intended behavior fixed while revising the test. Otherwise, iteration can drift toward a test that is easy to satisfy instead of one that checks the requirement.

5. Assess test quality beyond coverage

Coverage describes which code ran; it does not show that assertions would catch a defect. A 2024 study in Information and Software Technology discusses this limitation and uses mutation testing to assess generated tests’ ability to expose faults. See the study record.

Mutation testing evaluates a test suite against deliberately altered versions of a program. If a mutation changes behavior but the tests still pass, that can reveal a missing assertion or untested behavior. A surviving mutation is a prompt for investigation, not automatic proof that a test is defective: some mutations may be equivalent for the specified behavior or unreachable in the test setup.

Assess a generated suite with several signals together:

  • Execution: tests compile and run reproducibly.
  • Coverage: relevant paths and branches are exercised.
  • Fault detection: tests fail for meaningful behavior changes, including selected mutations.
  • Assertion quality: checks distinguish correct from incorrect outcomes.
  • Human review: cases match the requirement and remain maintainable.

6. Analyze source code or binaries for likely defects

Static detection work in the 2025 review covers approaches aimed at source code and binary defects. These analyses can help prioritize suspicious regions or explain a possible issue. Verify findings with ordinary code review, static analysis, reproduction, and tests. A generated explanation is a lead to investigate, not proof of a defect or its absence. The review describes these research categories.

How to choose an approach and evaluate it

Decision Questions to ask
Input context Will the task use source code, structured requirements, or natural-language user stories? Is the context complete and current?
Output level Do you need high-level scenarios, executable tests, repair suggestions, or defect-analysis results?
Evaluation Will you check execution, coverage, mutation or fault detection, assertion quality, and human review?
Feedback loop Can execution results guide revisions while the expected behavior stays anchored to a requirement?
Evidence maturity Is a claim based on a survey or review, an individual experiment, or a preprint? Does its context match your project?

Do not compare approaches using a single coverage number or assume one study result transfers to another codebase. The reviewed evidence does not support a comparable cross-industry accuracy, adoption, or productivity figure.

Reliability, privacy, and cost considerations

Generative AI adds review and evaluation work. Budget time to validate requirements, inspect generated assertions and patches, reproduce failures, and maintain tests as the code changes. The benefit depends on whether the generated artifact reduces that work without weakening the checks.

Keep prompts scoped to the task and follow your organization’s rules for source code, customer data, and secrets. Use synthetic or redacted examples where possible. Pin the model and prompt configuration when reproducibility matters, and retain the generated artifact and relevant execution result so reviewers can see what was proposed and what passed.

The research summarized here provides no universal cost or performance benchmarks. Measure your own workflow: time to review and repair output, test execution time, useful faults detected, and maintenance burden. Include the cost of model use and human review rather than counting generated lines alone.

Or skip the browser setup

When a test workflow needs a screenshot of a page state, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns an image or PDF; see the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free.

FAQ

Can generative AI generate test cases?

Yes. It can draft candidates from code or requirements. Check each case against intended behavior and run it before relying on it.

Does high coverage mean an AI-generated test suite is good?

No. Coverage records executed code, not whether assertions catch faults. Fault-oriented evaluation such as mutation testing adds a useful signal.

Can AI replace software testers?

The cited surveys and reviews describe assistance tasks and research approaches; they do not establish that testers can be replaced.

Should AI-generated tests go straight into a production branch?

Only after the same review, execution, and maintenance checks applied to other tests. The output is a proposal, not validation.