ScreenshotNeo

BlogEngineering

How Generative AI Is Changing Software Testing

Generative AI can speed up test ideation and implementation, but generated tests still need careful evaluation. Here’s what the evidence shows and how teams can use it responsibly.

By the ScreenshotNeo team4 October 202611 min read

Generative AI is changing software testing by helping developers brainstorm test cases and draft test code. It can reduce the effort of getting started, especially for unit tests, but generated tests still need to be checked for correctness, usefulness, and fit with the existing test suite. The available research supports cautious use for unit-test generation; it does not establish equivalent reliability across end-to-end, GUI, acceptance, security, or other testing types.

For developers, the practical change is a shift in work: AI can draft candidates, while people remain responsible for deciding what behavior matters, supplying relevant context, running the tests, and evaluating what they actually prove.

1. What generative AI changes in software testing

In this article, generative AI in software testing means using a generative model to propose test ideas or produce test implementations from a prompt and available code context. That is different from testing an AI system itself, which may involve checking a model’s outputs, robustness, or safety.

Test generation can help with several parts of a developer’s workflow:

  • Test ideation: Suggest input classes, boundary conditions, and failure cases a developer can evaluate.
  • Test implementation: Draft test functions, setup code, and assertions in the project’s framework.
  • Test maintenance: Help interpret failures or propose updates when code changes, subject to review.
  • Test explanation: Summarize what a test appears to exercise and identify assumptions to verify.

These are possible uses, not guarantees. A plausible-looking test may not run, may encode the wrong expected behavior, or may pass without detecting defects. The research summarized here is strongest for unit-test generation; it should not be treated as proof that AI can reliably automate every testing layer.

2. A practical workflow for using AI to generate tests

Use a model as a drafting assistant inside a reviewable testing process. Supply the code and relevant project conventions, ask for focused cases, then execute and assess the results.

  1. Choose the behavior to test. State the contract, important inputs, expected outputs, and failure behavior. Avoid asking for tests for an entire system without defining what correctness means.
  2. Provide useful context. Include the function or component, its dependencies, the test framework, related tests, and any fixtures or project conventions the model needs. Existing suite context can matter: one empirical study found markedly different outcomes depending on whether generation took place within an existing suite.
  3. Request cases before code. Ask for a compact list of normal cases, boundaries, invalid inputs, and relevant regressions. Correct omissions or wrong assumptions before asking for implementations.
  4. Generate a small batch of tests. Ask for runnable code using the project’s actual framework and style. Keep each test tied to a specific behavior so failures are interpretable.
  5. Review the assertions. Check that each expected result comes from the product contract or an independently verified requirement, not an invented assumption.
  6. Run the tests in the project environment. Fix compile, import, fixture, and environment issues; do not count unexecuted generated code as a passing test.
  7. Assess value, not just volume. Look for meaningful assertions and behavior exercised. Where appropriate, use mutation testing or other measures aligned with the quality claim.
  8. Keep human ownership. A developer or tester should be able to explain why each test exists and what regression it would catch.

NIST’s 2025 GenAI Code Challenge evaluation plan describes a pilot to measure and evaluate AI-generated unit tests for elementary Python code. That is evidence that evaluation is a necessary part of the work; the plan itself is not a finding that generated tests are effective. NIST’s evaluation plan.

3. What empirical studies say about generated unit tests

GitHub Copilot study: suite context and execution mattered

El Haji, Brandt, and Zaidman’s 2024 peer-reviewed study examined 290 GitHub Copilot-generated Python tests associated with 53 sampled tests from open-source projects. In the study’s setting, about 45.28% of generated tests were passing when generation happened within an existing test suite. Without an existing suite, 92.45% were failing, broken, or empty.

These are results from that study’s sample, tool, language, and 2024 setting. They are not current benchmarks for all Copilot versions, models, languages, or teams. They do illustrate two practical points: context can affect generated output, and generated code must be executed and classified before anyone treats it as a useful test. See the TU Delft research record for the study.

Student study: perceived help came with trust concerns

A 2026 observational study by Ardıç, Le Dilavrec, and Zaidman involved 12 undergraduate students using ChatGPT running GPT-3.5 for unit-testing tasks. Participants reported time-saving, reduced cognitive load, and help with test ideation. They also raised diminished trust, test-quality concerns, and lack of ownership. The study abstract reports that interaction and prompting strategies did not significantly affect test effectiveness or test-code quality as measured by mutation score or test smells.

This small student study describes those participants’ experience; it does not establish professional productivity gains or universal effects. Its measures also underline why counting generated tests is not enough. See the study in Empirical Software Engineering.

4. How to evaluate AI-generated tests

Evaluate a generated test at several levels. Passing is necessary for a test intended to join a passing suite, but a passing test can still assert little or test the wrong thing.

Check Question to answer Warning sign
Execution Does the test compile, load, and run under the project’s normal command? Empty output, skipped test, import error, or hidden dependency on an unavailable environment.
Expected behavior Are expected values grounded in a contract, requirement, or independently checked result? The model guessed behavior that has not been specified.
Assertion strength Would the test fail if the target behavior regressed? Assertions only check that a value exists, is truthy, or does not raise when that is not the contract.
Suite fit Does it follow project conventions and use the right fixtures, mocks, and isolation? Duplicated setup, brittle timing, network dependence, or conflicting assumptions.
Coverage of behavior Does it exercise meaningful paths, boundaries, and relevant failures? Many similar cases while an important branch or boundary remains untested.
Effectiveness measure Does the measure match the claim being made? Using test count or line coverage alone as proof of defect detection.

Mutation score, used in the student study, is one possible effectiveness measure: mutations represent small changes to code, and tests that detect them can provide evidence that assertions respond to behavioral changes. It is not a complete measure of test quality, and mutation results need interpretation. Test smells can also flag maintainability concerns, but do not by themselves establish whether a test catches important defects.

A review checklist

  • Can a reviewer state the behavior this test protects?
  • Does the test fail when a relevant behavior is deliberately changed?
  • Are boundary and invalid cases chosen based on the actual contract?
  • Are mocks limited to external boundaries rather than hiding the behavior under test?
  • Does the test pass consistently in the normal CI environment?
  • Is the generated test distinct from existing coverage and worth maintaining?

5. How developers’ and testers’ roles are changing

Generative AI can move effort from writing every test line toward specifying intent, supplying context, reviewing assumptions, and checking effectiveness. Developers still need to understand the code and expected behavior; testers still bring risk-based thinking, test design, and independent scrutiny.

The student observations suggest a real tension: participants described reduced cognitive load and ideation support, alongside lower trust and weaker ownership. Teams can address that tension by requiring authors to explain assertions, keeping generated tests reviewable, and making the person approving a test responsible for its purpose and maintenance.

AI output can also shape what people think to test. Treat generated ideas as prompts for human judgment, not as a complete inventory of risk. A model may omit unusual state transitions, concurrency conditions, security properties, or domain-specific invariants unless context and explicit direction bring them into view.

6. Risks and controls for teams

Gartner’s August 2025 abstract identifies hallucinations, skills atrophy, intellectual property, and regulatory infringement as risks associated with GenAI-assisted testing. This is an industry advisory summary, not a quantified experimental result. Its risk categories suggest controls teams can make concrete:

  • Hallucinated behavior: Require assertions to map to a documented contract or reviewed expectation.
  • Skills atrophy: Keep test design and review in the team’s practice; use AI output as material to reason about, not as a substitute for understanding.
  • Intellectual property: Apply organizational rules to source code and prompts sent to external services, and use approved tools and settings.
  • Regulatory concerns: Track what generated artifacts are used in regulated work and apply the organization’s required review and documentation process.

See Gartner’s advisory on risks of GenAI-assisted testing. The source gives risk categories; the controls above are practical recommendations, not quoted Gartner requirements.

7. What the evidence does not establish

The cited work does not establish that generated tests are generally reliable across all languages, models, organizations, or testing levels. The Copilot study concerns Python test generation in a specific 2024 setting. The human-interaction study concerns 12 undergraduates using GPT-3.5 on unit-testing tasks. NIST’s document describes an evaluation pilot rather than a positive effectiveness result. Keep those boundaries attached to the conclusions.

In particular, these sources do not settle how well generative AI handles end-to-end browser tests, GUI exploration, acceptance testing, security testing, or production monitoring. Those tasks involve different oracles, environments, and failure modes and need evidence specific to them.

8. Where screenshots fit in browser-testing workflows

Visual regression and browser debugging sometimes need screenshots as artifacts. A screenshot can help compare what a page rendered, but it does not replace assertions about behavior, accessibility, or security. For a browser test that needs a reproducible local capture, a developer can use an automation framework such as Playwright. The following Node.js example captures a full-page PNG from a public URL.

import { chromium } from 'playwright';

const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

Install Playwright and its browser in the project, then run the script with the target URL as its first argument. For dynamic sites, replace networkidle with a selector or application-specific ready condition; network activity can continue indefinitely on some pages. Keep screenshots out of pass/fail logic unless the comparison method has defined tolerances for fonts, animation, timestamps, and other sources of visual variation.

Or skip the browser setup

For a screenshot artifact without maintaining browser automation, ScreenshotNeo offers a website screenshot API and MCP server. Its website screenshot service can return an image or PDF from one GET request. For the full set of request options, see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo removes known cookie and consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

9. Cost, speed, and reliability considerations

AI-assisted test generation adds costs and review work that should be considered alongside any saved drafting time. Model usage may have service-specific pricing and data-handling terms; check the provider’s current terms rather than assuming a universal cost. The studies cited here do not provide a general cost or speed benchmark.

Generated tests can be unreliable in practical ways even when the model responds quickly: code may not compile, tests may be empty, assumptions may be wrong, and output may vary with supplied context. Improve reliability by generating small batches, including the existing suite and framework conventions, executing tests automatically, and requiring review before merge. Track the team’s own measures, such as review time, failure rate, maintenance burden, and an effectiveness measure suited to the test objective.

For screenshot artifacts, ScreenshotNeo’s plan prices are Free for 1,000 shots per month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Browser-based DIY capture avoids an API charge but requires browser setup and maintenance. Choose based on how often screenshots are needed and whether that operational work is worthwhile.

10. Common problems and fixes

Problem Likely cause Fix
Generated test does not compile or import The model guessed project paths, names, framework version, or fixtures. Provide the relevant files and exact test command; ask for a minimal patch that matches existing conventions, then run it.
Test passes but catches no meaningful regression Assertions are weak, or the test follows implementation details rather than the contract. State the behavior under test and check that a relevant code change makes the test fail.
Expected value seems plausible but is wrong The model inferred an undocumented requirement. Verify expected behavior against a source of truth before accepting the test.
Tests fail only in CI or intermittently Time, randomness, concurrency, external services, or environment assumptions are uncontrolled. Use deterministic inputs, isolate external boundaries, and follow the project’s established fixture and cleanup patterns.
Many generated tests duplicate one another The prompt rewards quantity or lacks suite context. Provide existing tests and request only uncovered behaviors; review for distinct failure conditions.
AI-generated browser screenshot is blank or incomplete The page has not reached its ready state, requires authentication, or renders content after the chosen wait condition. Use an application-specific ready selector, valid test credentials, or a deterministic fixture; inspect the capture independently.
Visual snapshots change between runs Fonts, animation, timestamps, ads, or responsive layout vary. Stabilize viewport and test data, disable animation where appropriate, and use comparison tolerances that match the purpose.

11. Frequently asked questions

Can generative AI replace software testers?

The cited evidence does not show that. AI can assist with test ideas and drafts, while people remain responsible for risk selection, expected behavior, and judging whether tests provide useful evidence.

Does a passing AI-generated test prove the code is correct?

No. It shows that the code passed that test under that environment. The test may omit important behavior or encode an incorrect expectation.

Is this research evidence about AI testing AI systems?

No. The empirical studies and NIST pilot described here focus on generating or evaluating unit tests for code, not on a general method for validating AI models.

Should generated tests be committed?

They can be committed when they meet the same project standards as other tests: reviewed intent, meaningful assertions, stable execution, and a clear maintenance owner.

Do these findings apply to every programming language?

No. The cited empirical generation study examined Python tests in its specific setting. Other languages and testing tasks need their own evaluation.

Sources