Human-AI Collaboration in Software Testing
Learn how developers can use AI to brainstorm software tests, verify suggestions, and keep humans in control of quality and coverage.
Humans and AI can work together in software testing by having people define intended behavior and risk areas, using AI to suggest test scenarios, and then reviewing, running, and maintaining the tests as a team. AI can broaden brainstorming, but a generated test is only useful when its expected result is correct and the test exercises behavior that matters.
How can humans and AI work together in software testing? Treat AI as a collaborator in test design, not as the final authority on correctness. A practical loop is: specify behavior, ask for candidate cases, inspect each case and its oracle, run the tests, and retain only tests that improve meaningful coverage.
What human–AI collaboration means for testing
Test generation is often described as if a model autonomously produces quality tests. In practice, people make important decisions before and after generation: which behavior matters, what context to provide, which suggestions to accept, and whether the expected outcome is actually correct. Those choices shape the quality and cost of the result.
A 2026 empirical study by Billy Shi and Per Ola Kristensson examined human–LLM interaction during test-case brainstorming. Its first study compared interaction with an LLM and web search with 16 participants. The second, with 24 participants, examined preemptive prompting, buffered responses, and guided input. These were studies of a bounded brainstorming task, not an evaluation of end-to-end production QA.
In the first study, the article reports participants spent 126% more time interacting with LLMs than with Google search. This is interaction time in that study, not total task time or a general estimate of the cost of AI-assisted testing. In the second study, preemptive prompting improved test quality by 33% and creativity by 35% on average and reduced user idle time by up to 49% in the studied task. These results are evidence about those participants, tasks, and interaction strategies; they do not guarantee the same gains for a different team or system. Read the ACM article.
A practical workflow for AI-assisted test development
The workflow below is a practical synthesis for development teams. It is not a workflow validated or prescribed by the studies described above.
- Define the behavior and risks. Write down the contract the code should satisfy, important inputs, state changes, error behavior, and risks such as authorization failures or data loss. Separate known requirements from assumptions.
- Give AI bounded context. Provide the relevant function signature, requirements, constraints, and existing test conventions. Ask for candidate scenarios first, or request test code if you can review it effectively. Do not include secrets or unrelated proprietary data.
- Ask for varied cases. Request normal cases, boundary values, invalid input, state transitions, failure paths, and interactions between relevant conditions. Ask the assistant to state its assumptions and expected result for each case.
- Review each proposal before adopting it. Check that the case follows the specification, that its expected result is justified, and that it is not a duplicate. Reject plausible-looking tests that assert an invented requirement.
- Run tests and inspect failures. A passing test shows that the current implementation agrees with the assertion; it does not, by itself, prove that the assertion expresses the intended behavior. Investigate both implementation failures and failures caused by incorrect test assumptions.
- Keep the test suite maintainable. Use the same review, naming, isolation, and cleanup standards as for human-written tests. Revise or remove tests that are flaky, redundant, or tied to an incidental implementation detail.
Example: turn a behavior contract into candidate cases
Suppose a function accepts a percentage discount from 0 through 100. Before asking an AI assistant to draft tests, state whether the endpoints are inclusive, what happens for non-numeric input, and whether values outside the range raise an exception or are normalized. Without those decisions, a generated test can silently encode the wrong contract.
A useful prompt asks for a table of inputs, expected outcomes, and the requirement each case covers. Review that table before requesting framework-specific code. This creates a clear checkpoint between brainstorming and adding assertions to the suite.
How to evaluate a collaboration approach
Compare AI-assisted testing methods on more than how many tests they generate. The dimensions below combine concerns examined in the human–LLM study with verification considerations teams need to assess in practice.
| Dimension | Questions to ask |
|---|---|
| Test quality | Do suggestions represent valid behavior and exercise meaningful branches or failure conditions? |
| Time and attention | How much time goes to writing prompts, waiting, switching context, reviewing, and repairing suggestions? |
| Breadth and creativity | Does the assistant surface useful cases the tester had not considered, rather than just restating obvious examples? |
| Human control and acceptability | Can the tester choose when AI contributes, steer it, and understand what assumptions it made? |
| Verification burden | Can a reviewer trace each test to a requirement and independently validate its expected result? |
The ACM study directly considered test quality, creativity, attention, and interaction design. Verification burden is a practical review axis; the sources cited here do not establish a broad benchmark of verification effort across testing tools.
Interaction patterns: preemptive, buffered, and guided
The 2026 study investigated three modified interaction strategies. They are useful design ideas to consider, but the reported results belong to the study’s test-case brainstorming task.
- Preemptive prompting: anticipate useful next steps and provide assistance before the user has to request every detail. The study reports promising gains for this approach on its measured quality, creativity, and idle-time outcomes.
- Buffered responses: manage how responses arrive so the user can continue working rather than repeatedly waiting on a conversational turn. The study investigated this strategy; do not assume it will reduce time in every workflow.
- Guided input: help the user provide relevant information through structured guidance. This may make constraints and missing details easier to surface, while still leaving decisions with the tester.
Choose an interaction pattern based on the team’s work. If reviewers need to approve each case, prioritize controllable suggestions and clear assumptions. If generating a wide list is useful, ask for ideas in batches and then review them together. Track whether these choices improve your own workflow instead of treating a study result as a universal setting.
Human review: what to check in an AI-generated test
- Requirement fit: Can you point to a specification, contract, or agreed behavior that supports the assertion?
- Oracle correctness: Is the expected result correct, including error type, side effects, and state after failure?
- Input validity: Does the test use realistic values and valid setup for the system under test?
- Coverage value: Does it exercise a distinct behavior, boundary, branch, or risk?
- Isolation: Does it depend on order, network state, a clock, random data, or external services without controlling them?
- Maintainability: Is the test readable and resilient to harmless implementation changes?
- Security and privacy: Did the prompt expose credentials, personal data, or code that should not be sent to the chosen AI service?
Do not equate more generated tests with better assurance. A large set of duplicate assertions, brittle snapshots, or tests with guessed expected values can increase maintenance cost while leaving important behavior untested.
What the available evidence does and does not show
The ACM article reports two empirical studies focused on user interaction during test-case brainstorming. Its results support examining how prompts, response timing, and user guidance affect the collaboration. The authors also discuss mixed initiative, acceptability, and user appropriation. The findings should not be generalized into claims that AI universally makes QA faster, that generated tests are correct, or that human review can be removed.
NIST’s 2025 GenAI pilot plan describes measurement and evaluation of AI-generated unit tests for elementary Python code. The plan signals that test effectiveness needs evaluation; it is not a completed benchmark or evidence that AI-generated tests are dependable. Read the NIST plan.
Using screenshots in visual testing workflows
For web applications, screenshots can help reviewers compare a rendered page across changes, viewport sizes, or themes. They are evidence about visual output, not a replacement for assertions about application behavior. A team can use a browser automation setup to capture a page, then review or compare the resulting image as part of its own test process.
DIY browser capture with Playwright
This runnable JavaScript example opens a page and saves a full-page PNG. Install Playwright and its Chromium browser first:
npm install playwright
npx playwright install chromium
Save as capture.mjs and run with node capture.mjs https://example.com:
import { chromium } from 'playwright';
const target = process.argv[2];
if (!target) throw new Error('Pass a URL, for example https://example.com');
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(target, { waitUntil: 'networkidle', timeout: 30000 });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
For pages with analytics or long-lived network connections, networkidle may never occur. Prefer waiting for a specific stable element or use domcontentloaded followed by an explicit selector wait. Keep viewport, browser version, fonts, locale, and test data consistent if screenshots are compared over time.
When an API capture is more suitable
A screenshot API can avoid maintaining browser binaries and capture infrastructure for simple page captures. ScreenshotNeo is a website screenshot API and MCP server made by Yorker Media. It accepts one GET request with a URL and can return PNG, JPEG, WebP, or PDF. Its cookie-consent handling accepts banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. See ScreenshotNeo and its API documentation.
Or skip the browser setup
Use this one-call request to capture a page as WebP. Replace the example URL with the page you need and provide your API key.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents, including Claude, Cursor, and any MCP client, take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. See the docs for request options, then sign up free for 1,000 screenshots a month, no card required.
Reliability, performance, and cost considerations
- Keep capture conditions stable. Fix viewport dimensions and wait for meaningful page readiness. Animations, dynamic data, consent state, fonts, and delayed images can make visual results vary.
- Use targeted waits. Waiting for all network activity can be slow or unreliable on pages with analytics and streaming requests. Wait for a specific selector when possible.
- Control external dependencies. Tests that rely on live websites can fail because the site, network, or third-party assets changed. Use controlled fixtures or stable test environments for repeatable checks.
- Account for review time. AI interaction may add prompting and verification work. The study’s reported interaction-time finding is specific to its experiment, so measure the full workflow in your own context.
- Estimate API use against actual needs. ScreenshotNeo plans are Free: 1,000 shots/month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free. Only clean shots are billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses indicate the page verdict and billing status in
X-Page-VerdictandX-Billedheaders.
Troubleshooting
| Problem | Likely cause | What to do |
|---|---|---|
| Generated test passes but misses a bug | The assertion may encode an incomplete requirement or test only a happy path. | Trace the test to the behavior contract; add boundary, error, and state-transition cases based on explicit requirements. |
| Generated test has an unexpected expected value | The assistant inferred behavior that was not specified. | Reject or correct the oracle using the specification or a domain decision. Do not keep a test just because it passes. |
| Many suggestions repeat existing tests | The prompt lacks suite context or asks only for a large volume. | Ask for distinct behaviors and a rationale per case; compare suggestions with existing coverage before implementation. |
| Visual screenshot test is flaky | Dynamic content, animations, fonts, viewport differences, or timing changed between runs. | Stabilize inputs and viewport, disable animation in the test environment, and wait for a meaningful page element before capturing. |
| Playwright navigation times out at network idle | The page keeps network connections open or continuously fetches data. | Wait for domcontentloaded and a specific selector, or use a bounded delay only when the page has no better readiness signal. |
| AI suggestions take longer to review than to write manually | Prompting, context switching, or oracle verification outweighs the brainstorming benefit. | Limit requests to risk areas where breadth helps, provide clearer constraints, and compare total review and rework time with a manual baseline. |
FAQ
Can AI replace software testers?
The evidence here does not support that conclusion. People still need to define expected behavior, validate test oracles, and decide whether the resulting coverage is useful.
Should AI write test code or only suggest scenarios?
Use the form that your reviewers can validate. Scenario lists make assumptions easier to inspect before code is added; code generation can save transcription when requirements and conventions are clear.
Does more test coverage mean the AI-generated suite is better?
Not necessarily. Coverage can rise while tests remain redundant or assert the wrong behavior. Review the requirement and expected outcome behind each test.
What does NIST’s pilot establish?
The cited NIST page describes a plan to evaluate AI-generated tests for elementary Python code. It does not report completed results establishing reliability.


