AI Testing Limitations: Why Human Testers Still Matter
AI can generate tests, but test output alone cannot establish that software is correct or safe in context. Learn where human judgment strengthens evaluation.
AI can generate candidate tests and help evaluate software, but the presence of generated tests does not show that a system has been tested adequately. Human testers still matter because someone must question what “correct” means, probe failures, and judge whether results make sense in the context where people will use the software.
That does not mean humans always outperform AI, or that a person must inspect every generated test. It means test generation, controlled evaluation, adversarial testing, and field evaluation answer different questions. Strong evaluation combines methods and makes their limits clear.
1. What AI testing can and cannot establish
AI can produce test code, suggest scenarios, and help execute structured evaluations. Each of those capabilities can be measured. NIST’s Code Pilot, for example, evaluates AI-generated unit tests for elementary-level Python code. That is evidence about a specific task and scope; it is not proof that generated tests cover every language, application, or production risk. NIST GenAI: Code Challenge (Pilot)
A test suite is useful only in relation to the risks and requirements it covers. Generated tests may run successfully while missing an important user journey, asserting an assumption that is wrong, or failing to represent a deployment condition. Passing tests show that the software met those assertions under those conditions. They do not, by themselves, show the requirements were complete or the conditions representative.
| Evidence | What it can show | What it does not establish alone |
|---|---|---|
| Generated test code | A model can propose executable checks for a defined task. | That the tests are relevant, complete, or aligned with user needs. |
| Passing test run | The tested build met the encoded assertions in that run. | That untested cases, environments, or outcomes are safe. |
| Benchmark result | Performance on the benchmark’s tasks and setup. | Performance in every language, product, or deployment context. |
| Field evaluation | Evidence about real interactions, interpretation, and effects. | A universal guarantee of correctness or safety. |
2. Why expected results are hard to define
Ordinary software tests often compare an observed result with a defined expected result. For AI-based systems, the expected answer may be difficult to specify. Systems can be complex, based on large datasets, poorly specified, or nondeterministic. ISO/IEC identifies the test-oracle problem: difficulty determining expected results and therefore deciding whether a test passed or failed. ISO/IEC TR 29119-11:2020
Consider an assistant that summarizes a policy. A test can check that the output contains required facts and omits forbidden disclosures. But deciding whether a summary is sufficiently clear, appropriately cautious, or useful to a particular audience may require criteria beyond exact string matching. For a creative or open-ended task, there may be several acceptable answers rather than one canonical output.
Human testers help by making assumptions visible: What user need is being served? Which mistakes matter most? What evidence counts as an acceptable answer? Their judgment should be recorded as explicit criteria and supported by evidence where possible. Human review is not automatically objective or consistent; teams can improve it with rubrics, examples, multiple reviewers for high-impact judgments, and documented disagreements.
3. Pre-deployment tests can miss deployment context
A model or feature may behave differently when it meets real users, varied inputs, connected tools, organizational policies, and changing conditions. NIST’s Generative AI Profile warns that available pre-deployment testing and evaluation approaches may be inadequate, applied nonsystematically, or fail to reflect deployment contexts. It also describes field testing as a way to examine how people interact with, consume, use, and make sense of AI-generated information, including the actions and effects that follow. NIST AI 600-1, Generative Artificial Intelligence Profile
Field evaluation can reveal problems that a fixed test set may not expose: users may misunderstand a confident-sounding answer, use an output in an unexpected workflow, or take consequential action based on it. These findings require attention to the product and the setting, not just the model’s isolated response.
Plan field work with defined questions, appropriate safeguards, and a process for responding to harmful or misleading outcomes. Record the deployment conditions and the limits of the findings. A small or short study can uncover useful issues, but it cannot establish that every user or setting will behave the same way.
4. Use complementary evaluation modes
NIST’s ARIA program distinguishes model testing, red-teaming, and field testing. It aims to measure technical and contextual robustness, going beyond a single emphasis on performance or accuracy. These modes provide different evidence and should be treated as complementary evaluation approaches, not replacements for all other software testing. NIST Assessing Risks and Impacts of AI (ARIA)
| Mode | Main question | Setting and evidence | Human contribution |
|---|---|---|---|
| Model testing | How does the system perform on defined capabilities and tasks? | Controlled tasks and measures; results are bounded by the test set and setup. | Choose meaningful tasks, inspect failures, and assess whether measures fit the intended use. |
| Red-teaming | How can the system fail under adversarial or challenging inputs? | Structured attempts to expose weaknesses; findings depend on the probes and access available. | Develop plausible misuse and edge-case probes, then judge severity and realistic consequences. |
| Field testing | How do people use and interpret the system in ordinary settings? | Interaction and outcome evidence in a deployment or representative context. | Observe workflows, identify misunderstandings, and interpret effects that raw scores may miss. |
For a product release, connect each mode to the risks it can address. A capability benchmark may be appropriate for a narrow model behavior; red-teaming may explore abuse paths; field testing may examine human interpretation and workflow effects. Traditional unit, integration, regression, accessibility, security, and operational testing still matter where relevant.
5. A practical evaluation workflow
- Define the use and stakes. Document who will use the feature, for what task, in which setting, and what harm or failure would matter.
- Write acceptance criteria. Separate objectively checkable requirements from judgments that need a rubric or human interpretation. Include unacceptable outcomes and escalation conditions.
- Generate candidate tests. Use AI to propose cases or test code if useful. Treat these as suggestions; review whether each case is valid, independent, and tied to a requirement.
- Run deterministic checks. Use unit, integration, regression, schema, policy, and other applicable tests. Control inputs and versions so failures can be reproduced.
- Evaluate variable outputs deliberately. Specify acceptable ranges or rubric dimensions where exact matching is unsuitable. Track model and configuration versions, and avoid treating one successful sample as stable behavior.
- Probe failures and misuse. Include boundary cases, malformed inputs, prompt injection or other relevant adversarial conditions, privacy concerns, and dependency failures according to the system’s design.
- Test with representative people and contexts. Observe how users interpret outputs and what they do next. Obtain appropriate consent and protect sensitive information.
- Record evidence and limits. Keep test inputs, expected criteria, outputs, environment, reviewer rationale, defects, and unresolved risks. State what the evaluation did not cover.
- Monitor after release. Define signals, reporting paths, rollback or mitigation actions, and a schedule for reevaluation as models, prompts, data, and use patterns change.
6. How to evaluate AI-generated tests
Review the tests as software artifacts and as claims about expected behavior. For each generated test, ask:
- Which requirement or risk does it address?
- Does the assertion check the intended behavior, or merely repeat the implementation?
- Would the test fail if the relevant defect were introduced?
- Does it depend on unstable state, timing, external services, or brittle output wording?
- Are boundary conditions, error handling, and negative cases represented?
- Can another developer understand and maintain it?
Useful quality signals include requirement traceability, mutation or fault detection where suitable, reproducibility, meaningful assertions, and maintainability. No single score proves test adequacy. NIST’s Code Pilot is a useful example of evaluating generated unit tests within a defined elementary Python scope; do not generalize its results to all systems. NIST’s broader GenAI evaluation program also includes questions about code reliability and human studies comparing human performance with AI system performance. NIST Evaluating Generative AI Technologies
7. Capture website behavior as one part of evaluation
When the system under test is a website or web application, screenshots can help preserve visual evidence for a test case, compare a page before and after a change, or document the state a reviewer needs to interpret. A screenshot does not establish functional correctness, accessibility, or the meaning of dynamic content. Pair visual evidence with assertions, logs, and user-oriented evaluation where those are relevant.
Browser automation provides control over navigation, viewport, waits, and interactions. For example, Playwright can capture a page after a selector appears:
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.locator('main').waitFor({ state: 'visible' });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
Install Playwright with npm install playwright and its browser with npx playwright install chromium. Use a representative test URL and avoid capturing real personal data in artifacts.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. A single GET request returns a PNG, JPEG, WebP, or PDF. For a test artifact in WebP format, use cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
9. Performance, reliability, and cost
- Keep the evaluation proportional to risk. Run fast deterministic checks on every relevant change; reserve slower human studies and deeper adversarial work for the risks and release decisions they inform.
- Make failures reproducible. Pin dependencies and model or service versions where possible, preserve inputs and configuration, and separate flaky infrastructure failures from product failures.
- Account for nondeterminism. Repeated runs or samples may be needed to understand variability. Report the sampling method and avoid presenting a single run as a stable estimate.
- Budget reviewer effort. Human review is most useful when reviewers have clear criteria, relevant expertise, and a way to escalate ambiguity. Sample and prioritize according to risk rather than assuming every output needs identical inspection.
- Track total evaluation cost. Include compute, external services, test maintenance, participant time, reviewer time, and the cost of investigating false alarms or missed defects.
- Preserve useful artifacts safely. Screenshots, prompts, and output logs can contain sensitive data. Limit access and retention to what the evaluation requires.
10. Troubleshooting evaluation problems
| Symptom | Likely cause | Practical fix |
|---|---|---|
| Generated tests pass, but users still find serious defects. | The suite does not cover the relevant requirement, workflow, or context. | Trace tests to risks and user journeys; add field evidence and missing scenarios. |
| A test fails intermittently. | Nondeterministic output, timing, shared state, or unstable external dependency. | Capture inputs and versions, isolate dependencies, use explicit waits, and define tolerances only where appropriate. |
| Reviewers disagree on whether an answer passes. | Acceptance criteria are vague or there is no agreed test oracle. | Create a rubric with examples, document borderline cases, and use adjudication for high-impact decisions. |
| A benchmark score improves while product quality appears unchanged. | The benchmark may not represent the target task, users, or deployment conditions. | Check benchmark relevance and add evaluations tied to actual use and consequences. |
| Screenshot artifacts show incomplete or inconsistent pages. | Capture occurred before key content loaded, or the page depends on dynamic state. | Wait for a stable selector or application-ready signal, use consistent viewport and state, and record capture configuration. |
| Field findings cannot be reproduced. | Context, participant task, model version, or input conditions were not recorded. | Preserve a privacy-conscious record of the scenario and configuration, and distinguish reproducibility limits from absence of a problem. |
11. Frequently asked questions
Does AI make software testing unnecessary?
No. AI can assist with test generation and evaluation, while teams still need to decide what matters, examine evidence, and assess real-use effects.
Should a human inspect every AI-generated test?
That depends on risk and the generation workflow. Evaluate the process, sample and review according to impact, and ensure critical requirements have trustworthy coverage.
Can a high accuracy score prove an AI system is safe?
No single score establishes safety. The result depends on the measure, task set, threshold, and context; complementary evidence is needed for broader claims.
Are human judgments reliable enough to use?
They can provide important contextual evidence when criteria, reviewer qualifications, and disagreement handling are documented. Human judgment also has limits and should be evaluated.


