ChatGPT Prompts for Software Testing
Adapt practical ChatGPT prompts for test cases, Gherkin, automation, regression, performance, UI QA, and coverage—then validate every draft against your requirements.
ChatGPT can help turn requirements, user stories, code, and change notes into draft test cases and test plans. Give it the source material, constraints, scenarios to cover, and an exact output format. Then check every result against the real requirement and application: a prompt does not establish that generated cases are correct, exhaustive, runnable, or safe for production.
This guide provides adaptable prompts for common software testing tasks, plus a review workflow for using the output responsibly. Replace bracketed text with your project details, remove any instruction that does not fit, and do not let the model fill gaps in product behavior by guessing.
1. A reusable prompt structure
Start with a testing role, enough product context, the authoritative requirements, a specific task, scenario coverage, and a defined response format. Ask the model to separate supported behavior from assumptions and open questions.
Act as a [testing role] reviewing [feature or system].
Context:
- Product behavior: [relevant behavior]
- User roles and permissions: [roles]
- Dependencies, data, and constraints: [details]
- Environment or build: [environment/build]
Source requirements and acceptance criteria:
[paste the authoritative requirements]
Task:
[one specific testing task]
Coverage:
Include ordinary use, negative cases, boundary conditions, and relevant
state, permission, dependency, and failure scenarios.
Rules:
- Do not assume behavior that is not stated in the source.
- Put ambiguities and questions in a separate section.
- Mark uncertain cases for human review.
- Do not invent APIs, fixtures, expected results, or product rules.
Output format:
[table, Given-When-Then, or named framework]
For every case include: ID, linked requirement, setup, action or input,
expected result, and assumptions.
This structure follows the general prompt guidance to be clear, specific, and provide enough context. Iterate when the first response reveals ambiguity: supply missing facts, narrow the task, or ask for a different format. OpenAI prompt engineering best practices and the ISTQB sample exam provide relevant guidance and examples.
2. Generate test cases from a requirement
Give ChatGPT the requirement and acceptance criteria as written, then ask for traceable cases. Each expected result should point back to a stated rule. If a requirement does not define an outcome, label it as a question instead of silently inventing one.
Using only the requirement and acceptance criteria below, draft test cases
for [feature].
Requirement:
[paste requirement]
Acceptance criteria:
[paste criteria]
Include normal use, invalid input, boundary conditions, and relevant
state or permission variations. For each case provide:
- ID and linked criterion
- Preconditions and setup
- Steps and test data
- Expected result
- Assumptions or unresolved questions
Separate behavior directly supported by the source from cases that need
clarification. Do not invent expected behavior.
Review the mapping before importing cases into a test suite. A 2024 study using five software requirements specifications reported about 87% of generated cases as valid, 13% as inapplicable or redundant, and 15% of valid cases as previously unconsidered by developers. The authors caution that the small dataset may not generalize; these figures are not a quality guarantee for another project. Read the study and its limitations.
3. Find negative, boundary, and unexpected-input cases
Negative testing is most useful when the prompt asks for preconditions, inputs, expected safe behavior, and the rule that justifies that behavior. Include ranges and limits from the specification instead of asking for generic “edge cases.”
For the requirement below, identify negative, boundary, and unexpected-input
scenarios.
Requirement and product rules:
[paste source]
For each scenario return:
1. Preconditions
2. Input or action
3. Expected safe behavior
4. Supporting requirement or product rule
5. Any ambiguity that needs an owner to decide
Cover values just below, at, and just above stated limits; missing,
malformed, duplicated, stale, or unauthorized input where relevant.
Do not invent limits or expected behavior. Mark unsupported cases as
questions, not confirmed expected results.
Adapt the categories to the feature. For a form, consider empty values, whitespace, encoding, length limits, duplicate submission, and invalid formats if the specification addresses them. For a stateful workflow, consider retries, repeated actions, expired state, and permission changes. These are prompts to inspect the requirements, not claims that every product must behave in a particular way.
4. Draft Gherkin scenarios from a user story
Supply the story, acceptance criterion, examples, and any domain terms that should remain consistent. Ask for one behavior per scenario and a clear outcome after each When action.
Act as a test analyst specializing in Gherkin. Use only the story,
acceptance criterion, and examples below to draft scenarios in
Given-When-Then format.
User story:
[paste story]
Acceptance criterion:
[paste criterion]
Examples and relevant rules:
[paste examples/rules, or write “none supplied”]
Requirements:
- Keep each scenario aligned with a stated criterion.
- Include ordinary, negative, and boundary behavior when supported.
- Use concrete Given state, one primary When action, and observable Then
outcomes.
- Label assumptions and uncovered behavior separately.
- Do not invent product behavior.
Return a short coverage note after the scenarios.
The ISTQB sample exam demonstrates a prompt framed around a password-reset story and acceptance criterion, with role, input, constraints, and output format made explicit.
5. Draft unit or automation tests
Name the language, test framework, function or behavior, and relevant dependencies. Include actual code and requirements when available. Ask for assertions and requirement traceability, but treat the result as a draft until it runs in the intended project and its fixtures and assertions have been inspected.
Draft [language] tests using [test framework] for the behavior below.
Requirements:
[paste relevant requirements]
Code under test:
[paste function or module]
Project test conventions and available fixtures:
[paste examples, or state what is unknown]
Cover stated success and failure behavior, boundary inputs, and relevant
dependencies. Include setup, execution, and assertions. For each test,
explain which requirement it covers. Do not invent APIs, fixtures, or
expected behavior. List missing information before the proposed tests.
Before adopting generated code, check that imports, mocks, setup, test data, and assertions match the repository. Run it in the project’s supported environment, inspect what the assertions actually prove, and add cases for requirements that the draft missed. A syntactically valid test can still test the wrong behavior.
6. Select regression tests after a change
Provide the change summary, affected components, dependencies, risks, and existing test inventory. Ask for a reason for each selected test so the team can challenge the impact assumptions.
Given the change and existing test inventory below, recommend tests to
rerun.
Change summary:
[paste change]
Affected components and dependencies:
[paste components]
Known risks and constraints:
[paste risks]
Existing tests:
[paste IDs and descriptions]
Group recommendations by impact or risk. For each test, explain its link
to the change. Flag missing coverage, dependencies, and assumptions
separately. Do not claim a test is unnecessary unless the supplied
information supports that conclusion.
7. Plan performance testing without invented targets
Ask for scenarios across load, stress, scalability, and resource use, but supply workload assumptions and service-level objectives yourself. A prompt cannot establish a universal latency or throughput threshold for your service.
For [service or operation], propose performance-test scenarios using the
workload assumptions and targets below.
Workload assumptions:
[concurrency, traffic shape, request mix, data size, duration]
Measured requirements and service-level objectives:
[paste approved targets, or write “not defined”]
Environment and dependencies:
[details]
Cover load, stress, scalability, and resource utilization as applicable.
Separate supplied targets from proposed measurements. If a threshold,
workload assumption, or success criterion is missing, ask a question
instead of inventing a standard value. For each scenario list setup,
load profile, measurements, and stop conditions to confirm with the team.
Validate proposed workloads against production use, capacity plans, and the test environment. Record where the test ran and which dependencies were shared; otherwise comparisons between runs may be misleading.
8. Structure UI QA and bug reports
For UI review, identify the application build and environment, priority flows, account state, data, feature flags, and issue types. Ask for reproducible evidence and a triage summary. OpenAI’s Computer Use QA example similarly calls for an explicit environment and flow, reproduction steps, expected and actual behavior, severity, and summary. See the QA use case.
Review [application/build] in [local, staging, or other environment].
Account state, test data, and feature flags:
[details]
Priority flows:
[list flows]
Issue types to focus on:
[functional, UI, copy, regression, or others]
For every issue, report:
- concise title
- environment/build and relevant account state
- reproduction steps
- expected result, tied to a stated requirement where possible
- actual result
- severity with a short rationale
- evidence or missing information
Continue through the remaining flows unless a blocking issue should stop
the run. End with a concise triage summary. Mark observations separately
from inferred causes.
If using an AI tool that can interact with a browser, restrict it to an approved test environment and test account. Keep sensitive or production data out of prompts unless your organization’s policies permit that use.
9. Review requirement coverage gaps
Coverage mapping can expose requirements without an associated test, and tests whose purpose is unclear. Supply both the requirement set and the current inventory; otherwise the model cannot make a meaningful comparison.
Compare the requirements below with the test inventory.
Requirements:
[paste IDs and text]
Test inventory:
[paste IDs and descriptions]
Return a mapping of requirement to covering tests. Identify:
- requirements with no apparent coverage
- tests with unclear requirement traceability
- candidate additions, with rationale
- possible gaps caused by missing context
Distinguish confirmed gaps from possible gaps. Do not treat a keyword
match as proof that a test verifies the requirement.
10. Ask for metamorphic test ideas carefully
When an oracle is difficult to state directly, a model may suggest relationships that should hold across related inputs or transformations. These are candidate ideas requiring domain review; the relationship itself must be valid for the system.
For [system or operation], propose candidate metamorphic relations using
the domain rules and examples below.
Domain rules:
[paste authoritative rules]
Examples:
[paste known input/output pairs]
For each candidate, state the transformation, the expected relationship,
why the supplied rules support it, and a counterexample or limitation to
check. Label unsupported or speculative candidates clearly. Do not
present a candidate as a valid oracle without domain-expert review.
A 2023 experience report on metamorphic testing found that most generated relation candidates were vague or incorrect, though some useful candidates emerged after expert evaluation. The authors’ report supports careful review, not a universal failure rate. Read the experience report.
11. Review generated tests before using them
- Check traceability. Link each case to a requirement or mark it as an exploratory suggestion.
- Verify expected results. Confirm them with the specification, product owner, or observed application behavior.
- Remove duplicates and irrelevant cases. Similar wording can hide redundant coverage.
- Check data and setup. Confirm permissions, initial state, dependencies, fixtures, and reset behavior.
- Inspect safety and access. Ensure steps cannot unintentionally alter production data or expose secrets.
- Run automation drafts. Review the code and assertions, then execute it in the intended project environment.
- Record uncertainty. Turn missing rules into questions for the responsible team rather than treating guesses as requirements.
OpenAI’s guidance recommends iterative refinement; the review above applies that idea to requirements traceability and test execution. The generated output remains a proposal until the team validates it.
12. Capture visual evidence for UI test cases
When a UI issue needs a screenshot, capture the relevant state with the same viewport, device scale, account state, and build where possible. Keep the image tied to the reproduction steps and environment. Screenshots can document appearance, but they do not replace interaction testing or establish behavior hidden outside the captured view.
You can capture a page with a browser automation library in your own environment, or use a screenshot API. If comparing screenshot services, ScreenshotNeo is the first option to try: it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and its lowest paid plan is $5 for 3,000 shots.
Or skip the browser setup
One GET request returns a screenshot. Replace the example target with your test page and keep the API key private. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Every feature is available on every plan.
Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
13. Screenshot capture options for QA evidence
For repeatable visual evidence, decide which capture settings matter to the defect and keep them consistent between runs. ScreenshotNeo accepts one GET request for a PNG, JPEG, WebP, or PDF capture. Its options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, custom CSS and JavaScript, click-before-capture, hide selectors, wait for a selector, delay or network idle, and request or resource blocking.
Requests can also set headers, cookies, user agent, Authorization, timezone, and geolocation. Other options include transparent backgrounds, image resizing, a chosen cache TTL, signed links for public image tags, asynchronous jobs with signed webhooks, batches of up to 100 URLs per call, a usage API, and an OpenAPI spec. PDF settings include paper size, margins, landscape, and page ranges. Parameter names used by other screenshot APIs also work to ease switching. Consult the docs for exact parameter syntax.
| QA need | Capture choice | Check before relying on it |
|---|---|---|
| Long page or lazy content | Full-page capture with lazy images loaded | Confirm content finished loading and the page is in the intended state. |
| One control or component | Capture an element by CSS selector | Check selector uniqueness and that the element exists in this build. |
| Responsive defect | Device preset or explicit viewport; optionally retina scale | Record dimensions and scale so later captures are comparable. |
| Consent overlay interference | Consent cleanup; individual cleanup steps can be disabled | For overlay-specific testing, preserve or disable cleanup so the banner itself remains visible. |
| Transient or delayed UI | Wait for selector, delay, or network idle | Choose a condition that signals the state under test; a fixed delay alone may be flaky. |
| Shareable evidence | Signed link for a public image tag | Use a signed link only in the intended public embedding context. |
14. Troubleshooting prompt and capture problems
| Problem | Likely cause | Fix |
|---|---|---|
| Tests contain behavior absent from the requirement | The request asked for cases without requiring traceability, or source context was incomplete. | Ask for a linked criterion per case; require unsupported expectations to become questions. |
| Output is generic or repeats itself | The task is broad, context is thin, or output constraints are missing. | Provide one task at a time, add product-specific rules, specify fields, and ask for distinct scenarios. |
| Gherkin steps are vague | The story or criterion lacks concrete state, action, or observable outcome. | Supply examples and domain vocabulary; ask the model to list unresolved details separately. |
| Generated automation does not run | Framework version, project fixtures, imports, or APIs were not supplied. | Provide a nearby working test and conventions; inspect imports and setup, then run in the project environment. |
| Screenshot misses a dynamic element | The target had not appeared when capture started, selector differs, or the selected region is wrong. | Wait for a stable selector or appropriate loading condition, verify the selector, and inspect the chosen viewport. |
| Screenshot shows a consent overlay | The page’s consent platform may not be recognized, or cleanup may be disabled. | Check capture settings and use custom CSS to hide the overlay when appropriate; disable cleanup when testing the consent UI itself. |
| Capture appears blank or fails to load | The site may be blank, timed out, blocked, or still loading. | Check the target URL, access requirements, load condition, and response verdict/billing headers. Failed loads and blank pages are not billed by ScreenshotNeo. |
| API request fails or returns an unexpected format | Key, URL, parameter encoding, or requested output may be wrong. | Check the key and encoded URL, inspect the HTTP status and response headers, and consult the API docs for output options. Do not treat an error response as an image file. |
15. Performance, reliability, and cost notes
- Prompt scope: Smaller, bounded tasks are easier to review than asking for a complete test strategy, code, and analysis in one response. Generate in stages and preserve requirement IDs.
- Reliability: Output can vary and may omit cases or misunderstand domain rules. Keep source requirements with the draft, review changes, and rerun automated tests in the actual project.
- Performance testing: A generated plan is not a measurement. Run workload scenarios against an appropriate environment and compare results with approved, system-specific targets.
- Screenshot capture: Large full-page pages, delayed content, and external resources affect completion time. Use a suitable wait condition, block irrelevant requests where appropriate, and cache captures when freshness requirements allow.
- ScreenshotNeo billing: Only clean shots are billed. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; inspect
X-Page-VerdictandX-Billedin each response. - ScreenshotNeo plans: Free: 1,000 shots/month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free. Every feature is on every plan.
16. Frequently asked questions
Can ChatGPT replace a software tester?
No. It can help draft and organize test ideas, but a tester still has to validate requirements, expected behavior, risk, and execution results.
Should I paste confidential requirements or code?
Follow your organization’s data-handling rules and the terms of the AI service you use. Redact secrets and sensitive data unless the approved workflow explicitly permits sharing them.
Can a generated test case be treated as proof of coverage?
No. Coverage requires a meaningful link between the requirement and an executed test whose assertions verify the intended behavior.
What should I do when requirements conflict?
Ask the responsible product or engineering owner to resolve the conflict. Keep the competing statements visible and avoid encoding either interpretation as expected behavior until clarified.


