ScreenshotNeo

BlogComparisons

Traditional Testing vs. AI Testing: Key Differences

Traditional tests check specified behavior. AI testing also evaluates data, variable outputs, and risk. Here’s how to combine both approaches.

By the ScreenshotNeo team4 October 20269 min read

Traditional testing checks whether software behaves as specified for selected inputs. Testing an AI-based system includes those familiar checks, but also evaluates whether its data, model, and outputs perform acceptably across relevant users, conditions, and risks. Since AI outputs may vary and may not have one uniquely correct answer, teams need explicit acceptance criteria and evaluation methods.

“AI testing” can mean two different things: testing an AI-based product, or using generative AI to help test other software. This guide focuses on the first meaning unless it says otherwise. ISTQB treats them separately: CT-AI covers testing AI-based systems, while CT-GenAI covers applying generative AI in testing activities. ISTQB’s CT-AI certification page

1. Key differences at a glance

Area Traditional software testing Testing AI-based systems
Expected behavior Requirements often describe specific outputs or actions for an input. Several outputs may be acceptable. Define measurable criteria or a repeatable evaluation procedure.
Inputs Choose cases to exercise requirements, code paths, boundaries, and integrations. Test inputs and their relevance, quality, and coverage as part of the system, alongside code and interfaces.
Assessment Assertions can compare an observed result with an exact expected value. Use task-appropriate metrics and judgments. A generative system should be assessed against task and risk criteria, not assumed to have one canonical answer.
Repeatability Under controlled conditions, reruns of deterministic tests are generally expected to produce the same result. Some systems are non-deterministic or change when data, models, or configuration change. Track versions and re-evaluate material changes.
Lifecycle Unit, integration, system, acceptance, performance, and security testing remain useful. Add relevant testing across input data, models, and machine-learning development activities.
Risk Established test-management and risk-based approaches guide quality and security work. Choose evaluation objectives and scenarios in light of intended use and possible negative impacts.

ISO calls the difficulty of deciding what result should pass or fail the test-oracle problem. AI systems may be complex, data-intensive, poorly specified, or non-deterministic, making expected results harder to define. ISO/IEC TR 29119-11:2020

2. What stays the same

AI features run inside software products. Their APIs, user interfaces, permissions, integrations, deployment settings, and ordinary code still need applicable functional, regression, performance, and security testing. AI testing adds evaluation of the AI component and its interaction with the product; it does not replace normal software verification.

ISO/IEC TS 42119-2:2025 explains how the ISO/IEC/IEEE 29119 software-testing series applies to AI systems and components, using a risk-based approach to select suitable practices. It describes familiar test processes and techniques in the AI context. ISO/IEC TS 42119-2:2025

3. A practical workflow for testing an AI feature

  1. Describe the intended use. State who will use the feature, for what task, and under which conditions. Include important user groups, languages, input types, and foreseeable misuse where relevant.
  2. Define acceptable and unacceptable outcomes. Write criteria before choosing a score. Specify what must happen, what variation is allowed, and what constitutes a consequential failure. For example, a support-answer feature may allow different wording but require that answers cite approved material and avoid disclosing another user’s data.
  3. Build representative evaluation data. Include ordinary inputs, boundary cases, incomplete or ambiguous inputs, and relevant conditions. Review whether the set represents intended users and use cases. Keep evaluation data separate from training data where needed to make the evaluation meaningful.
  4. Keep conventional tests around the feature. Test request validation, authentication, authorization, error handling, integrations, UI behavior, and deployment configuration as applicable. Use deterministic assertions for properties that should be exact.
  5. Evaluate model behavior against the criteria. Choose task metrics and human review or other application-specific judgment as appropriate. Assess additional characteristics such as robustness, safety, bias, reliability, or impact when they matter to the intended use.
  6. Probe risks deliberately. Test foreseeable edge cases and failure conditions, including malformed inputs, out-of-scope requests, and conditions relevant to the product’s harms. Match scenarios to the application rather than adopting an unrelated universal checklist.
  7. Record what was evaluated. Preserve the model and data versions, configuration, evaluation-set version, criteria, and results needed to interpret a run.
  8. Re-evaluate changes. Repeat suitable checks after changes to the model, data, prompts, configuration, dependencies, or operating conditions. Watch for shifts in input characteristics that can reduce performance.

4. Make evaluation reproducible

For a variable-output system, a useful test record describes the conditions and the decision procedure, not only one sample answer. Store the input set or its controlled version, model identifier, relevant configuration, evaluator or metric version, and aggregate and failure-level results. If a specific response must be stable, define that requirement explicitly and test it separately from quality criteria that permit variation.

Here is a small Python pattern for evaluating a deterministic per-case score against an agreed threshold. It is illustrative: the threshold and scoring function must come from the application’s acceptance criteria, not from a universal AI-testing rule.

from dataclasses import dataclass

@dataclass
class CaseResult:
    case_id: str
    score: float

MIN_SCORE = 0.90
results = [
    CaseResult("ordinary", 0.97),
    CaseResult("ambiguous", 0.88),
]

failures = [r for r in results if r.score < MIN_SCORE]
for result in results:
    print(f"{result.case_id}: {result.score:.2f}")

if failures:
    raise SystemExit(f"Evaluation failed: {len(failures)} case(s) below threshold")
print("Evaluation passed")

This example only shows threshold handling; it does not define how to score a model or establish that 0.90 is appropriate. Real evaluations should select metrics and thresholds based on the task, users, and risks.

5. Where visual testing fits

For an AI-enabled web product, screenshot comparison can help check the surrounding interface: whether a result panel renders, error states are visible, or a layout change breaks the user flow. A screenshot cannot establish that the model’s answer is accurate, fair, safe, or useful. Treat visual checks as one part of conventional product testing, with separate evaluation for model behavior.

For repeatable visual checks, keep viewport, browser conditions, test account state, and page data controlled. Dynamic content may require hiding variable regions or waiting for the relevant element before capture. Compare screenshots alongside functional assertions so a visually similar page does not conceal a broken interaction.

6. Or skip the browser setup

For capturing a page used in a UI check, ScreenshotNeo is a website screenshot API and MCP server. It returns PNG, JPEG, WebP, or PDF from one GET request. Its clean-shot flow accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

See the ScreenshotNeo API documentation for parameters and response details. The API also supports full-page and selector captures, viewport and device presets, dark mode, custom CSS and JavaScript, waits, request blocking, cookies and headers, caching, async jobs, bulk capture, and signed links. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents.

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. AI agents can take screenshots through the MCP server. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for free ScreenshotNeo screenshots.

7. Performance, reliability, and cost considerations

  • Performance: AI evaluation can include data preparation, model execution, and human review. Run fast, deterministic software checks early; reserve broader or expensive evaluations for appropriate stages. Select test-set size and review effort to match the decision and risk.
  • Reliability: A passing result applies to the tested data, model, configuration, and criteria. Preserve those versions and repeat relevant evaluations after changes. A single score does not prove behavior for every user or input.
  • Cost: The sources cited here do not establish a universal cost comparison between traditional and AI testing. Account for the work your evaluation actually requires, including data curation, model runs, infrastructure, and expert review. Avoid paying for activity that does not answer a defined acceptance or risk question.
  • Visual capture: ScreenshotNeo offers 1,000 shots per month free, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free; every feature is on every plan.

8. Common problems and fixes

Problem Likely cause Practical fix
There is no clear pass/fail answer. Expected behavior or acceptable variation was not specified. Define the task, acceptance criteria, allowed variation, and failure conditions before selecting an evaluation method.
The score looks good, but users report failures. The test set or metric may not reflect intended users, conditions, or important failure modes. Review representative coverage and add targeted scenarios. Use a complementary evaluation lens when the risk requires it.
Results change between runs. The system may be non-deterministic, or a model, dataset, configuration, or dependency may have changed. Record versions and run conditions. Separate requirements for exact repeatability from criteria that allow output variation.
Previously acceptable behavior degrades over time. Input data or operating conditions may have shifted, or a model or configuration changed. Monitor relevant performance and input conditions; investigate shifts and re-evaluate after material changes.
A screenshot assertion fails intermittently. Capture may occur before content settles, or dynamic page content and viewport conditions vary. Wait for the target state, stabilize test data and viewport, and mask only genuinely variable regions.
A visual check passes while the AI output is wrong. Screenshot comparison tests presentation, not semantic quality. Add separate behavioral and task-quality evaluation with criteria appropriate to the AI feature.

9. Standards and guidance

  • ISO/IEC TR 29119-11:2020 introduces AI-system testing challenges including the test-oracle problem. It is a 2020 technical report listed by ISO as under review; do not treat it as the newest ISO work.
  • ISO/IEC TS 42119-2:2025 describes applying established software-testing processes to AI systems and selecting techniques with a risk-based approach.
  • ISTQB CT-AI v2.0 focuses on testing AI-based systems, including input-data, model, and ML-development testing. The certification page states CTFL is a prerequisite. ISTQB distinguishes it from CT-GenAI, which covers using generative AI in testing.
  • NIST TEVV-Athlon is an initial public draft framework for customizing assessments to AI system goals and context. As of October 4, 2026, its public comment period is scheduled to close October 6, 2026; check the page for current status.
  • NIST AI Risk Management Framework resources provide related material for organizations working on AI risk management and evaluation.

10. Frequently asked questions

Does AI testing require a different test team?

Not inherently. The work may involve software testers plus people who understand the model, data, intended use, and relevant domain risks. Match expertise to the system and evaluation questions.

Is a benchmark score enough to approve an AI feature?

No single benchmark is established as sufficient for every application. A score answers only the question represented by its data and metric; acceptance should also reflect intended use and material risks.

Can generative AI write traditional software tests?

That is the other meaning of “AI testing”: using AI to assist testing work. It is distinct from evaluating an AI-based system. Generated tests still need review and should be checked against requirements.

Should every AI output be judged by a human?

The cited guidance does not prescribe one evaluator for every application. Select an evaluation procedure that fits the task and risk; some properties can be checked automatically, while others may need domain judgment.

Conclusion

Traditional testing asks whether software meets specified behavior. AI testing retains those checks and adds focused evaluation of data, variable outputs, model lifecycle changes, and application-specific risks. Write acceptance criteria first, use evaluation methods that match the intended use, and keep the test conditions traceable so results remain interpretable.