Accessibility Testing False Positives: Causes and How to Reduce Them
Learn why accessibility scanners raise false positives, how to verify findings in context, and how to reduce noise without hiding real barriers.
An accessibility scanner finding is a signal to investigate, not a verdict. A false positive is a reported issue that manual review determines is not actually an issue. A false assurance is the opposite problem: a scan passes a simplistic check while a real accessibility problem remains.
To reduce false positives, reproduce the reported interface state, inspect the affected element in context, check the relevant WCAG criterion and its exceptions, then document the decision. Keep automated checks in your workflow, but pair them with manual review and assistive-technology testing. A clean scan alone cannot prove that a site is accessible.
1. What counts as an accessibility false positive?
The UK Department for Education defines false positives as issues flagged by testing tools that are not issues after manual review. That definition matters: a report is not false merely because a developer disagrees with it or finds the warning inconvenient. The element, page context, applicable criterion, and any relevant exception must support the decision. See the Department for Education’s false positives guidance.
Automated tools are good at repeatable checks against patterns they can inspect. They generally cannot infer every editorial meaning, user intention, or exception from markup alone. A finding can therefore be technically accurate about a pattern while still requiring human judgment about whether that pattern is a failure in context.
| Term | What happened | What to do |
|---|---|---|
| True positive | The reported condition is a real accessibility problem. | Fix it and verify the affected state again. |
| False positive | Review shows the reported condition does not violate the relevant requirement in context. | Record the rationale and manage any rule adjustment transparently. |
| False assurance | A check passes, but a meaningful accessibility problem remains. | Use manual and assistive-technology checks to find what automation did not assess. |
| Needs review | The evidence is incomplete or the applicable requirement is unclear. | Keep it open until someone evaluates the relevant criterion and user outcome. |
2. Why do accessibility checkers report false positives?
Missing meaning and editorial context
A rule can detect that an image has alternative text without deciding whether the words communicate the image’s relevant information. It may also flag link text without determining whether the link’s purpose is clear from its surrounding context. Review what a user needs from the element, not just whether an attribute or string exists.
Exceptions in the criterion
A contrast checker may flag a logo’s colors even where the relevant contrast criterion exempts logos or branding. Confirm which criterion the tool is testing and whether an exception applies before closing the finding. Do not generalize one exemption to nearby text or controls that have different requirements.
The scan saw a different interface state
A scan covers the rendered content and state it evaluates. If a menu, dialog, disclosure panel, or other region was inactive or not rendered, its contents may not have been checked. The axe-core API documentation instructs users to expose inactive or non-rendered regions before analysis. Activate important interactions and evaluate each resulting state.
Ruleset, configuration, and content mismatch
A tool’s rules, exclusions, severity labels, and supported content formats affect what its results mean. For Section 508 evaluations, Section508.gov’s testing overview advises evaluating how a vendor defines and quantifies rules against the standards and expectations in use. Consider whether the ruleset is versioned, whether exclusions are visible, whether the source format is handled faithfully, and whether the findings give useful context.
3. How can a scan pass while the page still has an accessibility problem?
A passing result only says that the checks that ran did not report a failure. A simplistic presence check can accept an alternative-text attribute without judging whether its value is useful. The Department for Education gives examples such as classroom art labeled alt="car" or alt="image123": the attribute exists, but it does not provide a meaningful description. Conversely, a decorative image can appropriately use alt="".
Automation can also miss content that was not loaded, authenticated pages it could not reach, interaction states it did not activate, or user experiences that require judgment. The Department for Education says tools identify around 30% to 40% of issues; this describes the share of issues tools identify, not a false-positive rate and not a guarantee for a particular tool. Do not treat an empty report as proof of conformance.
4. A practical workflow for triaging findings
- Capture enough evidence. Record the rule identifier, affected element, page or flow, scan configuration and version, and interface state at scan time. This is a practical reporting approach consistent with defining scope, evaluating samples, and reporting results in W3C’s WCAG Evaluation Methodology (WCAG-EM) 2.0.
- Reproduce the state. Load the same page and repeat the interaction. If the reported content is inside a menu or dialog, expose it before running the analysis. If the state depends on authentication or data, arrange for the test to reach that state.
- Read the rule and criterion. Identify the rule’s actual claim and the relevant WCAG criterion. Check whether an exception applies. Do not decide based on the warning label alone.
- Inspect the element in context. For alternative text, assess whether the text is accurate and useful for the image’s purpose; for links, assess whether the purpose is understandable in context. Check the rendered result as well as the underlying markup.
- Check the user outcome. Manually evaluate whether someone can understand and operate the feature. Use assistive technology where the interaction or content calls for it.
- Classify and document. Mark a finding false positive only when review supports that conclusion. Otherwise fix it or leave it as requiring further evaluation. Record the rationale, evidence, owner, and any configuration change.
- Look for blind spots. Test relevant pages and states that the scan did not reach. Include manual and assistive-technology checks in the evaluation plan.
- Report scope and limits. Say what pages, states, samples, and methods were evaluated, along with unresolved findings and exclusions. WCAG-EM 2.0 provides an informative, technology-agnostic method for scoping, sampling, evaluating, and reporting; it is not a new normative WCAG requirement.
5. How to reduce false positives without hiding real issues
- Make scans reproducible. Fix the page state, test account, viewport, and relevant data so a reviewer can see the same condition.
- Expose dynamic content deliberately. Include automated steps for opening menus, dialogs, and other important interactive regions before scanning them.
- Keep rules and versions visible. Track the scanner and ruleset version, configuration, exclusions, and severity mapping with the result.
- Use narrow, reviewed exceptions. If a repeated finding is genuinely inapplicable, document the reason and limit the exclusion to the affected rule or element where the tool permits. Revisit exclusions when the page or ruleset changes.
- Preserve the original signal. Keep the raw report or finding history so a configuration change does not silently erase recurring results.
- Improve finding context. Retain selectors or element references, page URLs, rule identifiers, and remediation details. A reviewer should be able to locate the exact element and reproduce the state.
- Pair fast checks with human evaluation. Use automation for repeatable checks and manual or assistive-technology evaluation for context, interaction, and user experience.
There is a real balance to manage. Section508.gov explains that automated tools cannot apply human subjectivity, so they can produce excessive false positives or, when configured to eliminate them, test only a small portion of requirements. The goal is a defensible balance of useful precision and meaningful coverage, not a blank report at any cost.
6. Choosing and configuring an evaluation tool
There is no comparable, shared-corpus false-positive-rate benchmark in the reviewed authoritative sources for ranking tools. Avoid claims that a particular scanner has the lowest false-positive rate unless a relevant, comparable evaluation supports that claim. Instead, assess tools against your content and workflow:
| Evaluation question | Why it matters |
|---|---|
| Which standards and ruleset versions does it implement? | Findings are only interpretable against known requirements and a traceable ruleset. |
| Does it handle your content in its native format? | Format fidelity affects which checks can run and how faithfully they represent the delivered experience. |
| Can it reach authenticated and dynamic states? | A scan cannot assess content it cannot access or interactions it never exposes. |
| Are customization and exclusions transparent and version controlled? | Teams need to understand what was suppressed and reproduce results over time. |
| Are severity and remediation details useful? | Reviewers need enough context to reproduce, prioritize, and resolve findings. |
| Does it fit acceptance tests and CI? | Repeatable checks are easier to maintain when they run at stable points in the development workflow. |
7. Example: run axe-core against a rendered page with Node.js
This example scans the page currently loaded in Playwright. It first opens a menu so its rendered content is included. Replace the URL and menu selector with your application’s route and control. It prints violations for triage; a violation still needs contextual review, and a clean result is not a conformance verdict.
npm install --save-dev playwright @axe-core/playwright
npx playwright install chromium
// Save as accessibility-scan.mjs
import { chromium } from 'playwright';
import AxeBuilder from '@axe-core/playwright';
const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'domcontentloaded' });
// Adapt this to expose the state you need to evaluate.
const menuButton = page.locator('[aria-haspopup="menu"]').first();
if (await menuButton.count()) {
await menuButton.click();
}
const results = await new AxeBuilder({ page }).analyze();
console.log(JSON.stringify({
url,
violations: results.violations.map(({ id, impact, description, help, nodes }) => ({
id,
impact,
description,
help,
elements: nodes.map(({ target, failureSummary }) => ({ target, failureSummary }))
})),
passes: results.passes.length,
incomplete: results.incomplete.length
}, null, 2));
} finally {
await browser.close();
}
Run it with node accessibility-scan.mjs https://your-site.example/page. The menu selector is only an example; use a stable selector for the interaction in your application. Add explicit steps for each meaningful state, and keep incomplete results available for manual review rather than treating them as failures or successes automatically.
8. Or skip the browser setup
For a visual reference image of a page during review, you can use ScreenshotNeo, a website screenshot API and MCP server for developers. A screenshot can help preserve visual evidence for a review, but it does not run accessibility checks or establish conformance. The one-call API returns an image or PDF; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot, and each of those steps can be turned off. Bot checks, blank pages, failed loads, and cache hits are never billed; response headers identify the page verdict and billing status. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Use screenshots as supporting visual evidence alongside the accessibility evaluation methods above.
Sign up for 1,000 free screenshots a month, with no card required.
9. Troubleshooting common scan disputes
| Symptom | Likely cause | What to do |
|---|---|---|
| A finding disappears when you rerun the scan. | The page state, timing, data, or loaded content changed. | Record the state and inputs, make the interaction reproducible, and rerun against the same conditions. |
| A menu or dialog has no findings because it was closed. | The scan evaluated only the rendered state. | Open the region before analysis and scan each important state. |
| A logo is flagged for contrast. | The checker may not account for the applicable logo exception. | Verify the exact criterion and exception, and document why it applies to that logo. |
| An image passes because it has an alt attribute. | The check may verify presence without judging whether the text is accurate or useful. | Review the image’s purpose and the alternative text; use an empty value only when the image is decorative. |
| The report is clean, but a user still cannot complete a task. | The scan may have missed an interaction, context, or assistive-technology issue. | Reproduce the task manually, test relevant states and assistive software, and add a regression check where practical. |
| Reports differ between environments. | Different ruleset versions, configuration, content, authentication, or exclusions may be in use. | Compare versions and settings, then align the scan conditions and document intentional differences. |
| A team wants to suppress a recurring warning globally. | A broad exclusion can hide valid findings elsewhere. | Review the affected elements and limit any justified exclusion; keep it visible, version controlled, and periodically reviewed. |
10. Performance, reliability, and cost considerations
Automated checks are useful as repeatable early checks, including in a development workflow, but expanding scans to every state and route adds work and requires reliable setup. Choose representative pages and high-value states, then make the chosen sample explicit. Use smaller, focused checks during routine development and broader evaluation at appropriate review points; do not claim that either scope alone covers every requirement.
For reliable results, control scan versions and configuration, ensure the test can reach the intended content, and capture the page state alongside the result. A flaky or inaccessible test setup creates noisy reports that are difficult to distinguish from real changes. Track exclusions so reducing alert volume does not quietly reduce coverage.
Section508.gov cautions that tuning to eliminate false positives can reduce the requirements tested. The reviewed sources do not provide a common benchmark for scanner false-positive rates or comparative pricing, so tool cost and accuracy should be assessed against your own content, supported formats, ruleset needs, and review capacity. Budget time for human evaluation; automation does not replace contextual judgment.
11. FAQ
Can an automated accessibility scan prove my site is accessible?
No. A scan reports on the checks it ran and the states it reached. Combine automated results with manual review and assistive-technology testing, and report the scope and limits of the evaluation.
Should every scanner finding be fixed?
Every finding should be reviewed. Fix genuine issues; document a false positive only when the applicable criterion and page context support that decision.
Is a tool’s detection coverage its false-positive rate?
No. They measure different things. The Department for Education’s 30% to 40% figure describes issues identified by tools, not the share of tool findings that are false positives.
Can I compare scanners by a single false-positive percentage?
The reviewed authoritative sources do not provide a comparable tool-by-tool rate under a shared corpus and method. Compare rules, supported content, state access, configuration transparency, guidance, and workflow fit instead.
Does WCAG-EM 2.0 make a scan a formal conformance evaluation?
No. WCAG-EM 2.0 is informative guidance describing an evaluation process. It is not a new normative requirement or a replacement for WCAG.
Sources
- UK Department for Education, False positives.
- U.S. Section 508 program, Overview of Testing Methods for 508 Conformance.
- W3C Accessibility Guidelines Working Group, WCAG Evaluation Methodology (WCAG-EM) 2.0.
- Deque Systems / axe-core maintainers, axe-core API documentation.
- UK Department for Work and Pensions, How to do accessibility testing.


