ScreenshotNeo

BlogHow-to

How to Generate Software Tests With AI

Learn how to use AI to draft unit, integration, and end-to-end tests, then review, run, and improve them safely.

By the ScreenshotNeo team4 October 20266 min read

AI can draft unit, integration, and end-to-end tests when you give it the code, repository conventions, and specific behaviors to protect. Treat the output as a candidate: inspect its assertions, run it in your project, investigate failures, and add cases it missed. A passing suite or high coverage number alone does not show that the tests protect the behavior you care about.

1. Decide what the tests should protect

Start with observable behavior, not a request for “complete coverage.” List what the code should do for valid inputs, boundary values, invalid inputs, failures, and important interactions. Clarify ambiguous requirements before prompting; implementation code alone may not reveal the intended behavior.

For a function that parses a page count, for example, identify expected behavior for a normal positive integer, zero, a negative number, a nonnumeric string, and a missing value. For a component that calls a service, identify both the successful response and the relevant failure behavior.

2. Give the assistant repository context

Open or reference the implementation and nearby tests. Tell the assistant the language, test framework, naming style, fixture setup, and mocking conventions. Existing test files help it follow project patterns; without that context, it may choose unfamiliar imports, invent helpers, or use a different assertion style.

VS Code documents prompts for unit, integration, and end-to-end test generation, along with running and debugging tests in its editor. GitHub’s guidance likewise describes supplying relevant code and existing test context. See VS Code testing documentation and GitHub’s guide to writing tests with Copilot.

3. Ask for a focused draft

Name the unit or behavior, framework, project conventions, scenarios, and assumptions you want surfaced. A reusable prompt is:

Write tests for [function or module] using [test framework] and the conventions in [existing test file]. Cover [normal cases], [boundary cases], and [failure behavior]. Test public behavior rather than private implementation details. Use the existing fixtures and mocking approach. Return the test code and list any assumptions or cases that remain unclear.

Ask for a small, reviewable set first. You can then request missing cases or a separate integration or end-to-end layer. Avoid asking the model to certify that coverage is complete.

4. Choose the right test level

Test level Good fit Context to provide
Unit A function or module’s behavior in isolation Inputs, outputs, errors, framework, and existing unit-test examples
Integration Interactions between modules, storage, or a service boundary Dependencies, test environment, fixtures, and which boundaries should be real or mocked
End-to-end A user-visible workflow across the application Runner, setup instructions, entry point, user actions, expected visible results, and cleanup

Generated tests should match the risk and boundaries of the behavior. A unit test does not establish that a real integration works, and an end-to-end test may not pinpoint which internal behavior failed.

5. Review the generated tests before trusting them

For every test, check that it calls the real code under test and asserts an outcome that matters. Confirm that it does not simply repeat the implementation’s logic or depend on incidental private details. Review imports, fixtures, mocks, setup, teardown, and test names. Ask whether a plausible bug would make the test fail.

GitHub explicitly cautions that generated tests may miss scenarios and recommends reviewing them and adding tests as needed. In a peer-reviewed 2024 study sampling 53 tests from open-source Python projects, researchers evaluated 290 Copilot-generated tests. In that study setup, about 45.28% passed when an existing suite was available; without an existing suite, 92.45% were failing, broken, or empty. These are results for that sample, tool, language, and setup—not general failure rates for current AI tools. Read the study.

6. Run, debug, and refine

  1. Run the tests with the project’s usual command or IDE test runner.
  2. Separate syntax and setup errors from failures that indicate unexpected behavior.
  3. For each behavior failure, check the requirement and expected result before changing the test.
  4. Ask the assistant to address a specific error or add a named missing scenario, providing the failure output and relevant code.
  5. Run the suite again and review the final diff, including any changes the assistant made to production code.

Do not weaken an assertion just to make the suite green. VS Code documents running and debugging discovered tests through Test Explorer and the editor. VS Code testing documentation.

7. Use coverage as a map, not a verdict

Coverage can point to code that tests have not reached, but it does not show whether assertions would catch regressions. Review tests against requirements and plausible incorrect behavior. Benchmark tests themselves also need scrutiny: OpenAI’s 2026 audit reported material test-design or problem-description issues in 59.4% of 138 difficult SWE-bench Verified tasks. That audit concerns benchmark evaluation quality; it is not a measured rate for everyday AI-generated tests. OpenAI’s explanation of the audit.

8. Troubleshoot common problems

Symptom Likely cause What to do
Imports or test APIs do not exist The prompt omitted the project’s framework, version, or nearby examples. Provide the package and existing test file, then ask for a framework-conforming revision.
Fixture or setup errors The draft invented a fixture or skipped required setup and teardown. Point to the project’s actual fixtures and setup pattern; check cleanup for shared resources.
Tests fail on the first run The code may have syntax/setup problems, wrong assumptions, or an incorrect expected result. Classify the error, inspect the behavior specification, and share the exact failure for a targeted repair.
Tests pass but miss an obvious case The prompt did not enumerate that behavior, or the assertion is too weak. Add the case explicitly and verify that a plausible regression would fail it.
Mocks make the test pass regardless of behavior The mock replaces too much of the code path or the test asserts only that the mock was called. Mock only the intended boundary and assert the externally meaningful result.
End-to-end test is flaky Timing, shared state, external dependencies, or incomplete cleanup may be involved. Use the project’s established waits and isolated fixtures; avoid arbitrary delays where the runner supports waiting for a condition.
Coverage rises without confidence Lines execute but assertions do not distinguish correct from incorrect outcomes. Review requirements and failure sensitivity instead of optimizing only for the percentage.

9. Performance, reliability, and cost

AI test generation adds a review and execution loop; generated code still needs to be run with the same dependencies and environment as the project. Keep prompts focused to reduce irrelevant output, and generate suites in small groups when a module has many behaviors. For reliability, preserve the project’s ordinary test command in the workflow and rerun after revisions. The research cited here does not establish a universal time saving or accuracy rate, so evaluate usefulness in your own repository rather than assuming one.

Model and IDE costs depend on the tool and plan you choose; the cited documentation does not provide a basis for a general price comparison. Do not treat coverage gains or benchmark scores as proof of correctness.

Or skip the browser setup

If your tests or QA workflow need website screenshots, ScreenshotNeo is a screenshot API and MCP server for developers. One GET request returns an image or PDF; see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before capture; known consent platforms, newsletter popups, and chat widgets are removed.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents.
  • The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, no card required.

FAQ

Can AI write unit tests for my code?

Yes. It can draft candidates when you provide the code, framework, conventions, and behaviors. Review and run them before relying on them.

How do I get AI to test edge cases?

Name the boundary and failure scenarios explicitly in the prompt, then inspect the assertions to confirm each scenario is actually distinguished.

Should I ask AI for complete test coverage?

Ask for specific behaviors and gaps. No prompt or coverage percentage establishes that a suite is complete or meaningful.

Can generated tests replace QA review?

No. They are draft code that still needs engineering review, execution, and comparison with the intended behavior.