ScreenshotNeo

BlogGuides

How to Generate Software Test Cases with AI

Use AI to draft useful software tests from code or requirements. Learn how to prompt, review, run, and improve the cases without trusting guessed behavior.

By the ScreenshotNeo team4 October 202610 min read

AI can draft software test cases when you give it a clear source of expected behavior: code, a requirement, an acceptance criterion, or concrete examples. Ask for a focused set covering normal behavior, boundaries, invalid inputs, exceptions, and important branches. Then check every assertion against the source, add the tests in your project’s existing style, and run them in the normal test environment. Generated tests are proposals; a test that passes is useful only if it checks the intended behavior.

This guide shows how to go from test basis to reviewed tests, with runnable examples in Python and JavaScript. The same workflow applies to other languages and frameworks.

1. Choose the test basis before prompting

A test basis is the material that defines what the software should do. Choose the narrowest useful source and provide enough context to distinguish intended behavior from implementation details.

Test basis Best fit Include
Function or module Unit tests for existing code Relevant code, dependencies, framework, and nearby test examples
User story or acceptance criteria Behavior cases before or during implementation Actors, preconditions, expected outcomes, and constraints
API or data contract Validation, serialization, and compatibility cases Schema, required fields, types, error behavior, and representative payloads
Examples of input and output Clarifying a rule or generating test data Valid examples, invalid examples, and any boundary values already known

If a requirement is ambiguous, tell the AI to list questions and assumptions before writing tests. Do not let it silently invent business rules. GenAI can help analyze requirements, derive candidate test objectives, propose expected results, and create test data, but those outputs still need validation against the actual basis. See the ISTQB CT-GenAI syllabus.

2. Ask for scenarios before asking for code

First request a scenario list with the reason each case matters. A useful set usually includes:

  • Ordinary valid inputs and expected results
  • Minimum, maximum, just-below, and just-above boundaries where applicable
  • Empty, missing, null, or default values where the interface permits them
  • Malformed or unsupported inputs
  • Exceptions, timeouts, or dependency failures the code is meant to handle
  • Each important conditional branch, including alternate success paths
  • Repeated calls, ordering, or state transitions when behavior depends on state

Not every category applies to every function. Ask the model to explain why a proposed case belongs, and to identify categories that do not apply. GitHub’s guidance recommends detailed scenario prompts, including edge cases, exception handling, and data validation; complex cases need more context. GitHub: Writing tests with GitHub Copilot.

A reusable prompt

You are helping draft tests for a software project.

Test basis:
[Paste the relevant requirement, code, contract, or examples.]

Expected behavior:
[State known input/output rules. If any behavior is unclear, say so.]

Project conventions:
- Language and version: [e.g. Python 3.12]
- Test framework: [e.g. pytest]
- Relevant existing test style or fixtures: [paste a small example]
- Dependencies that should be mocked: [list, or say none]

First, list focused test scenarios for ordinary behavior, boundaries, invalid
inputs, exceptions, and important branches. For every scenario, state which
requirement it checks and its expected result. List unclear behavior and
assumptions separately; do not invent undocumented rules. Do not write or
modify files yet.

Review the scenario list before asking for implementation. Remove duplicates, resolve open questions, and make sure the expected result comes from a requirement or an explicit decision.

3. Generate tests that match the project

Once the cases are agreed, provide the framework and a nearby test file. Request descriptive names, focused assertions, minimal setup, and mocks only at external boundaries. Ask for the complete test file or a patch, and ask it to explain fixture and mock assumptions.

Example: Python with pytest

This small function has a defined contract: negative quantities are rejected; valid quantities are multiplied by a nonnegative unit price.

# pricing.py

def line_total(quantity: int, unit_price: float) -> float:
    if quantity < 0:
        raise ValueError("quantity must be non-negative")
    if unit_price < 0:
        raise ValueError("unit_price must be non-negative")
    return quantity * unit_price

Save this test module as test_pricing.py. Install pytest with python -m pip install pytest, then run python -m pytest -q from the directory containing the files.

# test_pricing.py
import pytest

from pricing import line_total


def test_line_total_multiplies_quantity_by_price():
    assert line_total(3, 2.5) == 7.5


def test_line_total_is_zero_for_zero_quantity():
    assert line_total(0, 2.5) == 0


@pytest.mark.parametrize("quantity,unit_price", [(0, 0), (1, 0)])
def test_line_total_accepts_zero_values(quantity, unit_price):
    assert line_total(quantity, unit_price) == 0


@pytest.mark.parametrize(
    "quantity,unit_price",
    [(-1, 2.5), (1, -0.01)],
)
def test_line_total_rejects_negative_values(quantity, unit_price):
    with pytest.raises(ValueError):
        line_total(quantity, unit_price)

The example makes equality appropriate because these inputs and results are exactly representable in the stated cases. For computations with rounding behavior, use a tolerance assertion and derive that tolerance from the requirement instead of selecting one arbitrarily.

Example: JavaScript with Node’s built-in test runner

Save the following module as pricing.mjs and tests as pricing.test.mjs. Run them with node --test on a supported Node.js version with the built-in test runner.

// pricing.mjs
export function lineTotal(quantity, unitPrice) {
  if (quantity < 0) throw new RangeError('quantity must be non-negative');
  if (unitPrice < 0) throw new RangeError('unitPrice must be non-negative');
  return quantity * unitPrice;
}

// pricing.test.mjs
import test from 'node:test';
import assert from 'node:assert/strict';
import { lineTotal } from './pricing.mjs';

test('multiplies quantity by unit price', () => {
  assert.equal(lineTotal(3, 2.5), 7.5);
});

test('returns zero for zero quantity or price', () => {
  assert.equal(lineTotal(0, 2.5), 0);
  assert.equal(lineTotal(3, 0), 0);
});

test('rejects negative quantity', () => {
  assert.throws(() => lineTotal(-1, 2.5), RangeError);
});

test('rejects negative unit price', () => {
  assert.throws(() => lineTotal(1, -0.01), RangeError);
});

These examples are deliberately small. For a real function, ask the AI to retain the public contract, import paths, fixtures, async conventions, and test naming style used by the repository. Avoid tests that merely assert private implementation details unless those details are themselves part of a stable contract.

4. Review every proposed case

Before adding generated tests, walk through this checklist:

  • Traceability: Can you point to the requirement or documented behavior behind every expected value?
  • Coverage of behavior: Are relevant branches and failure modes represented, rather than only the happy path?
  • Correct oracle: Does the assertion check the externally meaningful result, or just mirror the current implementation?
  • Realistic setup: Do fixtures, mocks, clocks, and data represent the conditions the code actually sees?
  • Independence: Does each test make its own result clear, without relying on order or leaked state?
  • Useful failure: Would a failure help someone locate the violated behavior?
  • Assumptions: Has every inferred rule been confirmed or removed?

High line coverage or a large number of tests does not show that assertions are meaningful. A test can pass while checking the wrong expectation. GitHub cautions that generated tests may not cover everything, and recommends reviewing them. Read the guidance.

5. Run tests and investigate failures

Run the suite with the repository’s normal command, environment variables, fixtures, and dependency versions. Microsoft’s workflow for AI-assisted testing includes comparing proposals with existing tests, adding agreed cases, running them, and investigating failures. Microsoft: Test existing code with AI.

  1. Run the new tests alone to catch syntax, import, and fixture issues quickly.
  2. Run the full relevant suite to detect interactions and regressions.
  3. Read each failure. Separate a bad test assumption from a product behavior defect.
  4. Confirm any corrected expected result against the test basis; do not edit assertions just to make the run green.
  5. Keep the final test names and comments focused on behavior, not on the fact that AI drafted them.

When a test fails, ask the assistant to explain the failure using the requirement and output, then propose possible causes. Do not have it automatically weaken the assertion or alter production code without review.

6. Add property-based tests where they fit

Example-based tests check selected inputs. Property-based tests check a general invariant across many generated inputs, which is useful when the invariant is simple to state and a broad input range matters. For example, if a normalization function promises idempotence, a property might assert that normalizing an already normalized value does not change it.

Ask the AI to propose candidate properties and counterexamples, then validate each property against the specification. Use the project’s chosen property-testing library and seed or reproduce failing examples according to its conventions. Property-based testing complements selected boundary and example cases; it does not establish that the property itself is correct. Anthropic describes an AI agent using this approach to find bugs: Finding bugs with Claude and property-based testing.

7. Protect code, requirements, and test data

Before sharing repository material with an AI service, follow your organization’s rules for source code, customer data, credentials, and confidential requirements. Remove secrets and personal data from examples, and use synthetic fixtures where possible. ISTQB’s CT-GenAI coverage includes hallucinations, bias, privacy, and security risks; its current certification page lists syllabus version 1.1 and states that CTFL is a prerequisite, so verify current details there before relying on certification information. ISTQB CT-GenAI.

For broader test-writing guidance, see ScreenshotNeo, a website screenshot API and MCP server for developers. Screenshot capture can be useful when a test workflow also needs visual artifacts of rendered pages; it is separate from generating assertions for application logic.

Or skip the browser setup

If your test workflow needs a rendered-page screenshot, ScreenshotNeo returns an image or PDF from one request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.

Performance, reliability, and cost

  • Keep generation focused: Work one function, requirement, or cohesive behavior at a time. Smaller proposals are easier to check and less likely to mix unrelated assumptions.
  • Reuse project context: Provide concise nearby examples and conventions instead of asking the model to infer framework setup.
  • Keep tests deterministic: Control time, randomness, network calls, and external state with the project’s established fixtures or mocks. Avoid adding brittle sleeps when a condition can be awaited directly.
  • Use failure reproduction: When a generated test exposes a real defect, retain the minimal reproducible input and verify the fix against the same requirement.
  • Manage service costs and privacy: Costs and data handling vary by AI provider and plan; check the provider’s current terms and follow organizational policy. The research does not establish a general productivity or defect-detection percentage, so do not use one as a planning assumption.

Troubleshooting

Symptom Likely cause What to do
Generated test does not compile or import The prompt omitted runtime, module, or repository layout details Provide the language version, exact test command, import conventions, and a nearby test file; request a corrected test only.
Test passes but the feature is still wrong The assertion encodes an invented or overly weak expectation Trace expected values to requirements and add an assertion for the user-visible behavior that matters.
Many tests duplicate each other The prompt asked for exhaustive output without a boundary or behavior model Ask for a compact scenario matrix and one test per distinct behavior or partition.
Mocks make the test unrealistic The AI mocked internal details or omitted an important integration boundary Show the dependency boundary and existing fixtures; mock only external effects that need isolation.
Test fails only in the full suite Shared state, ordering, environment, or fixture leakage Run the case alone and in different orders; isolate state and ensure cleanup.
AI invents behavior for unclear requirements The prompt did not forbid assumptions or request clarification Ask it to enumerate uncertainties first, then obtain a product or engineering decision before asserting them.
Generated test data contains sensitive values Real data was included in the prompt or copied into fixtures Replace it with synthetic examples, remove credentials and personal data, and follow the applicable data policy.

FAQ

Can AI generate tests from a requirement without seeing the code?

Yes. It can propose behavior scenarios from requirements, but implementation details and existing project conventions will still need to be supplied or checked later.

Should AI write the test file directly?

It can draft a patch, but reviewing a scenario list first makes guessed expectations easier to catch before they become code.

Does more coverage mean better tests?

No. Coverage indicates which code ran; it does not prove that assertions check the correct behavior.

What if there is no test framework yet?

Choose a framework supported by the project’s language and runtime, establish one minimal working test, then ask for cases that follow that convention.

Sources and further reading