Human Intelligence and AI in Software Testing
Learn where AI can help with software testing, how to test AI-based systems, and what human testers still need to judge and verify.
AI in software testing describes two related but different activities: using AI to help test conventional software, and testing software that itself contains AI. In the first, AI may help draft tests, analyze code, or maintain automation. In the second, the test target includes data and model behavior, which can be probabilistic and non-deterministic. In both cases, generated results need human review against requirements, observed behavior, risk, and privacy constraints.
AI can expand or speed up parts of testing, but it does not remove the need for people to decide what matters, whether evidence is adequate, and whether a release is acceptable. Available industry research describes opportunities, but does not establish one universal division of work or a general productivity gain.
1. What does “AI in software testing” mean?
There are two meanings to keep separate:
- Testing with AI: applying generative AI or machine-learning tools to help test ordinary software, such as proposing test cases or summarizing failures.
- Testing AI-based systems: evaluating a product whose behavior depends on machine-learning models, generative AI, or data-driven components.
ISTQB reflects this distinction in separate learning tracks: CT-AI v2.0 focuses on testing AI-based systems, while CT-GenAI focuses on using generative AI in the test process. They solve different training needs.
2. How AI can assist testing conventional software
A 2025 mapping study groups proposed or reported applications into areas such as test-case and script generation, requirements analysis, code and root-cause analysis, UI testing, test prioritization, defect prediction, execution, and maintenance. These are possible application areas, not a guarantee that a tool is mature or useful for every team.
| Task | Possible AI contribution | Human check |
|---|---|---|
| Requirements analysis | Suggest scenarios, boundary conditions, and ambiguities from supplied requirements. | Confirm the interpretation with product rules, stakeholders, and existing behavior. |
| Test design | Draft positive, negative, and edge-case ideas or turn examples into test skeletons. | Check coverage, oracle quality, and whether each case tests a meaningful risk. |
| Automation scripts | Generate or explain code for unit, API, or browser tests. | Run it, review selectors and assertions, and ensure it fails for the intended defect. |
| Failure analysis | Summarize logs, cluster similar failures, or suggest likely causes. | Trace the explanation to evidence; distinguish product defects from environment noise. |
| Prioritization | Rank tests using change information, history, or risk signals. | Check that important but rare or newly introduced risks are not hidden by historical data. |
| Maintenance | Suggest updates when interfaces or test scripts change. | Verify that a “healed” test still checks the original requirement and has not weakened an assertion. |
| Visual checks | Help inspect or compare page captures and identify visible differences. | Decide whether a difference is a regression, an intentional change, or rendering noise. |
The tool’s output is a proposal. Treat generated tests as untrusted until they run against the actual system and have assertions that reflect an explicit expected result. A plausible test can still encode an incorrect requirement, miss an important state, or pass without checking anything useful.
3. A practical human–AI workflow
- Frame the risk. A tester or engineer identifies the requirement, affected users, failure impact, and evidence needed for confidence.
- Give bounded context. Provide the relevant specification, API contract, code excerpt, or sanitized failure data. Remove secrets and personal data; follow organizational privacy and security rules.
- Ask for proposals. Request test ideas, data partitions, boundary cases, automation skeletons, or a failure summary. State the expected behavior and constraints explicitly.
- Review before execution. Check each suggestion for unsupported assumptions, missing cases, unsafe data, and weak assertions.
- Run in the real environment. Execute against the system version and configuration that matter. Preserve logs, inputs, and results so failures can be reproduced.
- Evaluate and adapt. Remove duplicate or low-value tests, add cases for uncovered risks, and keep tests that provide useful evidence.
- Make the release decision with accountable owners. A person reviews unresolved failures and determines whether evidence is sufficient for the risk involved.
This workflow is practical guidance inferred from the sources’ emphasis on evaluating generated results and considering hallucinations, reasoning errors, bias, privacy, and security. It is not a measured universal allocation of tasks.
4. How to test AI-based systems
For AI-based products, testing must account for more than conventional code paths. ISTQB’s CT-AI v2.0 outline describes probabilistic behavior, non-determinism, and reliance on data as characteristics that complicate exact repeatability. It organizes coverage around input data testing, model testing, and ML development testing, and includes generative AI and large language models.
Input data testing
- Check schema, types, ranges, missing values, malformed inputs, and preprocessing assumptions.
- Test representative, boundary, and unusual inputs, including cases where data quality is poor.
- Consider whether the test data reflects relevant user groups and deployment contexts; identify gaps rather than treating a sample as universally representative.
- Protect sensitive information in datasets, prompts, logs, and test artifacts.
Model testing
- Define acceptance criteria before looking at results. Choose metrics that connect to intended use and failure cost.
- Evaluate functional performance with appropriate data partitions and record model version, configuration, and evaluation conditions.
- Test expected behavior, edge cases, and harmful or disallowed behavior relevant to the product.
- For generative systems, examine variability across repeated runs where relevant, unsupported claims, instruction handling, and failure behavior. A single favorable answer does not establish reliable behavior.
- Use domain-specific review when correctness or safety depends on specialist judgment.
ML development and release testing
- Track data and model lineage so a result can be tied to the inputs, code, and model version that produced it.
- Re-evaluate after changes to training data, preprocessing, prompts, model versions, dependencies, or deployment configuration.
- Test the product integration too: permissions, fallbacks, timeouts, error handling, and the user experience around uncertain outputs.
- Monitor relevant behavior after release and define what change or failure triggers investigation or rollback.
Do not force a deterministic pass/fail expectation onto behavior that is intentionally variable. Define acceptable ranges and evidence appropriate to the system. Conversely, “the model is probabilistic” should not excuse failures: product owners still need clear acceptance criteria and explicit treatment of unacceptable outcomes.
5. What human testers contribute
Human judgment is especially important where work depends on context: deciding which user harm matters, interpreting ambiguous requirements, spotting a plausible but unsupported explanation, and deciding whether a failure warrants blocking a release. People also need to review generated scripts for privacy and security implications and check that automation has not silently weakened coverage.
The AI-T ontology paper describes a conceptual framework intended to support human testers, guide intelligent agents in generating or reusing tests, and aid mixed human–agent testing teams. That is a framework for organizing knowledge and collaboration, not evidence that a particular agent performs well.
A useful responsibility boundary is: let tools propose and process; keep people responsible for risk framing, priority choices, review, interpretation, and release accountability. The exact division depends on the system and team, and the cited research does not establish one allocation that works everywhere.
6. What the evidence says—and does not say
Karhu, Kasurinen, and Smolander’s 2025 secondary study mapped industry-context work published from 2020 onward. It identified possible uses such as test generation, code analysis, and intelligent test automation, while reporting that industry implementations and observed benefits in the mapped evidence were limited. The study is a mapping of available research, not a controlled estimate of how much AI improves speed or quality.
So treat claims of guaranteed savings, fewer defects, or tester replacement cautiously. The reviewed sources do not provide a broadly generalizable causal estimate for how much human–AI testing improves quality or speed. Measure outcomes in your own context, including false alarms, missed defects, review effort, maintenance cost, and whether the tests catch failures that matter.
7. Example: capture a page for a visual test
A screenshot can be a useful artifact in UI checks, bug reports, and visual comparisons. It does not determine whether a page is correct; a person or a well-defined assertion still needs to judge the expected result. Here is a small DIY example using Playwright in Node.js to capture a page. Install Playwright with npm install playwright and install its browser with npx playwright install chromium.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
try {
const response = await page.goto('https://example.com', {
waitUntil: 'networkidle',
timeout: 30000
});
if (!response || !response.ok()) {
throw new Error(`Page load failed: ${response ? response.status() : 'no response'}`);
}
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
})();
For a visual regression check, capture a known baseline under controlled browser, viewport, font, and data conditions, then compare later captures using a documented tolerance. Review changed regions; dynamic timestamps, ads, animations, and personalized content can create noise. Do not use a raw pixel difference as the sole release criterion.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; its API documentation covers the parameters. For example, this Python call saves a WebP capture:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Cookie banners are accepted and removed before the shot, along with known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month, with no card required.
9. Options, reliability, and cost when using screenshots
For browser-based capture, make the viewport, device scale, page state, and wait condition explicit. Full-page captures can expose content below the fold; element captures can focus on a component. Control animation and dynamic data where possible, and record capture conditions with the artifact. If the page is blank or still loading, diagnose navigation and readiness before interpreting the image.
ScreenshotNeo supports full-page and selector captures, dark mode, device presets and custom viewport, retina scale, custom CSS and JavaScript, click-before-capture, selector or delay or network-idle waits, request and resource blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent background, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage API, and PDF options. It also accepts parameter names used by other screenshot APIs. Use only the controls your capture requires; extra waits and very large pages can increase capture time. Cache only when a fresh page state is not required.
ScreenshotNeo plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan. Use the usage API and billing headers to reconcile consumption and distinguish clean captures from non-billable outcomes.
10. Troubleshooting
| Problem | Likely cause | What to do |
|---|---|---|
| Generated tests pass but bugs remain | Assertions encode the wrong expectation, cases are shallow, or important states were omitted. | Trace each test to a requirement and risk; add explicit boundary and failure cases; review what the test would fail on. |
| AI analysis sounds confident but is wrong | The model inferred context that was not provided or produced a reasoning error. | Require references to logs or code evidence; verify the claim directly; do not use unsupported analysis as a release decision. |
| AI-generated browser script is flaky | It may rely on unstable selectors, timing assumptions, or changing page content. | Use stable selectors, explicit readiness conditions, controlled data, and repeatable environments; review any self-healing selector changes. |
| AI-system results vary between runs | Variation may be expected from probabilistic or non-deterministic behavior, or may indicate configuration drift. | Record model and configuration versions, repeat relevant cases, define acceptable behavior ranges, and investigate changes against the acceptance criteria. |
| Model evaluation looks strong but users see failures | Evaluation data may not represent the deployed inputs or usage context. | Review data coverage and preprocessing, collect privacy-safe failure examples, and add representative scenarios to evaluation. |
| Screenshot is blank or incomplete | Navigation failed, the page needs more readiness time, or content is loaded only after interaction. | Check the HTTP/navigation result, wait for a relevant selector or state, and trigger required interactions before capture. |
| Screenshot differs from the baseline everywhere | Viewport, device scale, fonts, browser, content, or theme changed. | Normalize those conditions and regenerate a baseline only after confirming the change is intended. |
| Screenshot request is not billed as expected | The result may be a cache hit or a failed/non-clean page verdict. | Inspect the X-Page-Verdict and X-Billed response headers; consult the API docs for request and verdict details. |
11. Frequently asked questions
Will AI replace software testers?
The cited sources do not establish that conclusion. AI can assist with particular tasks, while context, risk judgment, review, and accountability remain necessary. How work changes depends on the product and the quality of the tools and process.
What should a human tester check in AI-generated test cases?
Check that each case matches a real requirement, covers a meaningful risk, uses safe data, has a useful oracle, and fails when the relevant behavior is broken.
Which ISTQB path is relevant?
CT-AI is for testing AI-based systems. CT-GenAI is for applying generative AI in testing work. Both pages list CTFL as a prerequisite; check ISTQB for current syllabus and exam arrangements.
Can screenshot automation decide whether a visual change is a defect?
It can capture or compare evidence, but whether a difference violates intended behavior depends on requirements and context. Define tolerances and review meaningful changes.


