ScreenshotNeo

BlogHow-to

How to Generate Test Automation with ChatGPT

Generate a useful first draft of automated tests with ChatGPT by sharing code, expected behavior, and project conventions—then review and run every test.

By the ScreenshotNeo team4 October 20268 min read

Short answer: Give ChatGPT the code or repository context, language and test framework, expected behavior, existing test conventions, and the edge cases you care about. Ask for a small, focused set of tests, inspect every assertion, and run the tests with the project’s normal command. Treat generated tests as a draft: they can encode a mistaken assumption or miss a requirement, and passing tests do not prove the software is correct.

1. What ChatGPT can help you generate

ChatGPT can help draft tests for a function, module, API behavior, or a narrowly scoped repository task. Match the test type to the behavior you want to check:

  • Unit tests check a small unit, such as a function, with its dependencies controlled where appropriate.
  • Integration tests check behavior across components or boundaries that need to work together.
  • Property-based tests check a rule across many generated inputs, rather than relying only on a few hand-picked examples.

These are different approaches, not interchangeable labels. Choose based on the behavior under test, compatibility with your language and existing project, and how the test fits the project’s usual local and CI workflow. OpenAI’s coding guidance discusses these test categories and examples such as empty inputs, maximum lengths, null inputs, and invalid states. See OpenAI’s coding use cases.

2. Gather the context before prompting

A vague request such as “write tests for this” leaves important decisions unstated. Before asking, collect:

  • The function, module, or behavior to test. Include relevant code, signatures, types, and dependencies.
  • The language, test framework, and the project’s command for running the relevant tests.
  • One or two existing tests that show naming, setup, fixtures, mocks, and assertion conventions.
  • Expected behavior for ordinary inputs, boundary values, unusual but valid states, and invalid inputs.
  • Failure behavior: what should happen when a dependency fails, input is malformed, or a precondition is unmet.
  • Constraints, such as avoiding network access, keeping tests deterministic, or not changing production code.

Do not paste secrets, credentials, private customer data, or unrelated source code into a prompt. Share the smallest context that makes the behavior and conventions clear.

3. Use a focused prompt

State the task, context, coverage, and output you expect. For example:

Using the existing [framework] conventions, write tests for this function:

[function and relevant types]

Expected behavior:
- [ordinary case and expected result]
- [boundary case and expected result]
- [invalid or failure case and expected result]

Follow the style in these existing tests:
[relevant test examples]

Cover documented behavior, boundary values, empty and invalid inputs, and
failure cases where applicable. Do not change production code. Explain what
each test asserts and list any assumptions you had to make. If behavior is
unspecified, ask a question or label the assumption instead of inventing a
requirement.

Replace the placeholders with real project details. If the expected behavior is unclear, settle that first; otherwise the generated test may merely lock in an accidental implementation detail.

4. Generate tests in manageable steps

  1. Start with one behavior. Avoid asking for tests for an entire application in a single prompt. A focused scope makes assumptions and omissions easier to spot.
  2. Ask for the test draft. Request tests that follow the examples you supplied, and ask for an explanation of each assertion and any assumptions.
  3. Review the proposed cases. Check that they cover normal behavior, meaningful boundaries, unusual valid states, and relevant failures. Not every category applies to every function.
  4. Check the assertions against requirements. An assertion should express intended behavior, not just reproduce whatever the current code happens to do.
  5. Save the tests in the project and run the normal command. Use the project’s actual dependencies, configuration, and environment.
  6. Use failures as evidence. Give ChatGPT the exact failing output and ask whether the test, implementation, or environment could explain it. Review any suggested edits before applying them.

When working with a repository-aware coding experience, provide a narrow task and ask it to inspect the existing test patterns before writing files. OpenAI describes Codex as a coding-focused experience for working with repositories, writing or debugging code, and running tests and commands. Available products, plan access, and workspace features can change; check the current OpenAI Help Center for setup and availability. Review generated changes and verify them in your own project before integration.

5. Review generated tests before trusting them

Read each test as a requirement. Ask:

  • Does the expected result come from a specification, documented behavior, or an agreed requirement?
  • Does the test check the outcome that matters, or an implementation detail that could change without breaking behavior?
  • Are mocks and fixtures appropriate, or do they hide the behavior the test should exercise?
  • Can the test pass for the wrong reason, such as because it never reaches the assertion or catches every exception?
  • Is it deterministic and isolated from network, clock, random, and shared-state effects unless those are explicitly under test?
  • Does the test name say what behavior is being checked?

Generated tests can miss cases, make up unspecified behavior, or contain incorrect setup. OpenAI’s Codex launch guidance says users should manually review and validate agent-generated code before integrating and executing it. A green test suite is evidence about the cases that ran; it is not proof of overall correctness.

6. Diagnose failures systematically

When a generated test fails, first classify the failure instead of immediately changing the implementation:

  1. Test expectation: Compare the assertion with the documented behavior. If they differ, correct the test or clarify the requirement.
  2. Production defect: Reproduce the behavior and check whether the implementation violates the requirement. Fix the code only when the evidence supports that conclusion.
  3. Test setup: Check fixtures, mocks, async handling, cleanup, and required environment variables against nearby working tests.
  4. Environment: Confirm the intended runtime, dependency versions, configuration, and test command. Compare with the project’s normal local or CI setup.

For a follow-up prompt, include the exact command, complete relevant failure output, the test, and the behavior requirement. Ask for a diagnosis and possible explanations before requesting a code change. Do not ask the model to “make the tests pass” without preserving the expected behavior; that can encourage weakening the assertions.

7. Common problems and fixes

Problem Likely cause What to do
The test imports the wrong module or uses nonexistent helpers The prompt omitted project structure or examples. Share the relevant import path and a nearby test; ask for a patch limited to the test file.
The test passes but checks the wrong thing The prompt described the current code rather than the intended behavior, or the assertion is too weak. Write down the requirement first, then revise the assertion to fail when that requirement is violated.
A boundary or failure case is missing The prompt listed only ordinary behavior or left error handling unspecified. Name the exact boundary and expected outcome. If the outcome is unknown, clarify it rather than guessing.
The test is flaky It depends on time, randomness, network access, shared state, or execution order. Control the source of variability or isolate the test, following existing project patterns.
Mocks make the test pass while integration is broken The test replaces the boundary whose behavior needs verification. Keep unit tests focused, and add an integration-level check where the real interaction is part of the requirement.
The proposed fix removes an assertion or catches broad errors The model is optimizing for a passing run instead of the requirement. Reject the weakening unless the requirement itself changed; ask for a fix that preserves the assertion.
Tests fail only in CI Local and CI runtimes, dependencies, configuration, or services may differ. Compare the exact command and environment inputs. Avoid assuming the test or CI is at fault before reproducing.

8. Reliability, speed, and cost

There is no reliability percentage established here for ChatGPT-generated tests. OpenAI examples describe useful test-generation tasks, not a universal accuracy or defect-detection rate. Evaluate a draft by reviewing its requirements, running it, and checking whether it catches representative regressions.

Keep the context focused: relevant code and a small number of representative tests are easier to review than an entire repository pasted into a prompt. Generate and validate a small group of tests at a time. The test suite’s runtime is determined by the project and test types; a network-dependent integration test, for example, has different setup and runtime needs from a focused unit test. Follow the project’s normal commands and CI constraints.

ChatGPT plan access, Codex availability, and workspace capabilities may vary and can change. Check current official product information for applicable access and pricing rather than relying on stale setup advice. Regardless of the tool or plan, review and execute the tests in the project environment before relying on them.

9. ScreenshotNeo for screenshot-based checks

If a test workflow needs a website screenshot as an input or visual artifact, ScreenshotNeo is a website screenshot API and MCP server for developers. It can return PNG, JPEG, WebP, or PDF captures. It does not generate or validate software tests; use it when the workflow needs to capture a page. Its API supports options such as full-page or CSS-selector capture, viewport and device presets, waiting for a selector or network idle, and custom headers or cookies. See the ScreenshotNeo API documentation.

Or skip the browser setup

One GET request captures a URL. This cURL example saves a WebP response:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

The equivalent Python request:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Replace the example URL and API key with your target and key. Cookie banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, and failed loads are never billed. The removals can be turned off individually. An MCP server lets AI agents use screenshot tools, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card.

10. FAQ

Can ChatGPT write unit tests for my code?

It can draft them when you provide the code, behavior, language, framework, and relevant conventions. You still need to inspect and run the tests.

Should I ask it to test an entire application at once?

Usually start with a narrow behavior. That makes assumptions, failures, and missing cases easier to identify before expanding coverage.

Do generated tests prove the code is correct?

No. They check the cases and assertions that were written and executed. Requirements review and other forms of verification may still be needed.

What should I provide when a generated test fails?

Share the exact command and failure output along with the relevant test and expected behavior. Ask for diagnosis first, then review any proposed changes.