How to Combine Accessibility Testing Methods for Better WCAG Coverage
Combine automated checks, manual evaluation, assistive technology, and user input in a documented WCAG review. Learn how to scope, sample, report, and avoid overstating coverage.
For better WCAG coverage, combine automated checks with manual inspection, assistive technology testing, and input from people with disabilities. Define the product scope, WCAG target, and supported browsers and assistive technologies first; then explore important views and workflows, select and document a representative sample, investigate tool findings, and report the limits of what was evaluated. A scan or sampled review does not prove that every part of a product is free of accessibility barriers.
1. Define the evaluation
Write down what the evaluation covers before choosing tools or test cases. State:
- Product and boundaries: Identify the website, application, or other digital product, including relevant platforms and areas that are in scope or excluded.
- Purpose: For example, checking a release candidate, assessing a key workflow, or documenting conformance.
- WCAG version and target level: Choose the version and conformance level against which you will evaluate. Level AA is generally accepted and recommended in the WCAG-EM 2.0 methodology, but the target for a particular evaluation must be explicit. W3C WCAG-EM 2.0
- Accessibility support baseline: Name the browsers, operating systems, assistive technologies, and other user agents included in the evaluation. A result applies to the combinations you tested; it does not automatically establish behavior in every environment.
WCAG-EM 2.0 is a W3C Group Note that supports WCAG evaluation. It does not add requirements to WCAG. Its evaluation method now covers apps and other digital products as well as websites. Read the WCAG-EM overview.
2. Explore the product before choosing tests
Map the important views, content types, technologies, and functionality. Include complete processes where they matter, such as account creation, sign-in, search, form submission, or checkout. Record variations that may change accessibility behavior: for example, validation errors, empty and populated states, expanded menus, dialogs, and success or failure messages.
This exploration helps you choose tests that represent how the product works, rather than selecting pages at random without context. Capture how to reach each important state, including settings, inputs, and actions required to reproduce it.
3. Select and document a useful sample
Evaluate every view and state when practical. When that is not feasible, choose a structured sample covering distinct and important views, content, and functionality. Add random sampling where appropriate to discover issues outside the known high-priority areas. WCAG-EM describes sampling as part of a manageable evaluation, not as evidence about unexamined parts of a product. See the WCAG-EM sampling method.
For each sample, document its URL or location, how to reach it, the state and data used, and the interactions to perform. Include the reason for selecting it. A compact sample plan might look like this:
| Sample | Why it is included | States and interactions |
|---|---|---|
| Sign-in | Core entry workflow; has form controls and error handling | Empty submit, invalid credentials, successful sign-in |
| Search results | Repeated content and changing result counts | No results, several results, keyboard navigation |
| Product detail | Distinct layout and primary action | Expand details, choose an option, submit action |
| Checkout | Multi-step process with validation | Required-field errors, correction, completion |
Adapt the examples to the product. If the same template drives many pages, sample distinct templates and meaningful variations, then record what was not evaluated. Do not describe a sample as full-product coverage.
4. Use automated checks as one layer
Automated and semi-automated tools can make evaluation more efficient by identifying issues that can be checked mechanically. Run them on the selected views and states, then inspect findings in context. A reported issue may require human judgment to decide whether it is a real barrier, while a clean result cannot settle questions that require understanding content, purpose, or interaction.
Choose tools based on the criteria that matter for your evaluation:
- Criteria coverage: Which checks are automated, and which need human evaluation?
- Context and interaction: Can the tool assess the actual content, state, and workflow, or does a person need to judge them?
- Repeatability: Can the same configuration and sample be run again and compared?
- Environment support: Does the method fit the browsers, assistive technologies, and platforms in your declared baseline?
- Time and expertise: What setup and interpretation do results require?
- Real user participation: Does the evaluation include feedback from people with disabilities?
Section 508 guidance discusses automated, manual, and hybrid testing, and recommends assessing a tool’s rule methods and accuracy against expectations. Keep the tool name, version, configuration, date, and method with your results. Section 508 overview of testing methods. W3C also describes evaluation tools as a way to improve efficiency, not as a replacement for evaluation. W3C evaluation tools guidance.
5. Manually inspect content and interaction
For every sampled view, evaluate the applicable WCAG requirements against the chosen support baseline. Manual review is needed wherever a check depends on meaning, context, or a complete interaction. Include checks such as:
- Whether text alternatives communicate the purpose of meaningful images, and whether decorative images are handled appropriately.
- Whether headings, labels, instructions, and link names make sense in context.
- Whether color is the only way information or status is conveyed, and whether text remains legible at the tested presentation settings.
- Whether all functionality can be reached and operated by keyboard, with a visible focus indicator and a sensible focus order.
- Whether dialogs, menus, validation errors, and dynamic updates can be understood and operated without losing a user’s place.
- Whether forms identify errors and provide enough information to correct them.
- Whether pages and controls remain usable at the tested zoom, viewport, and input settings.
These are prompts for evaluation, not a substitute for checking the requirements in your selected WCAG version and level. A checklist helps make work repeatable, but the evaluator still has to understand the interface and judge the actual experience.
6. Test with assistive technologies and involve users
Test representative workflows with the assistive technologies in your support baseline. For example, check whether a screen reader can identify and operate controls, follow changes in page state, and understand form errors. Combine that with keyboard-only use and any other input modes relevant to the product. Record the exact browser, operating system, assistive technology, and workflow so findings can be reproduced.
Where practical, involve people with disabilities in evaluating real tasks. Their experience can reveal friction and barriers that a conformance checklist alone may not convey. User input complements systematic evaluation; it does not replace it. WCAG-EM recommends considering user involvement as part of evaluation. WCAG-EM 2.0.
7. Combine the methods into a repeatable workflow
- Set scope and target. Record the product boundaries, evaluation purpose, WCAG version and level, and accessibility-support baseline.
- Explore key views and workflows. Inventory distinct templates, content, states, and complete processes.
- Choose the sample. Cover important and distinct areas; add random samples if useful. Record reachability steps and why each sample was selected.
- Run automated checks. Save tool name, version, configuration, date, and output. Treat output as findings to investigate, not as a verdict.
- Evaluate manually. Check applicable requirements and investigate automated findings in their actual content and interaction context.
- Use assistive technology and user input. Test end-to-end tasks with the declared baseline and, where practical, include people with disabilities.
- Report results and limits. Identify what was evaluated, what was not, and which conclusions the sample supports.
- Track fixes and retest. Re-run relevant checks after changes and record whether the original barrier is resolved and whether the change affected related states.
8. Report findings without overstating coverage
A useful report lets another person understand and reproduce the evaluation. Include:
- Product identity, scope, purpose, and evaluation dates.
- WCAG version and target level.
- Sample selection method, evaluated views and workflows, and excluded areas.
- Browsers, platforms, assistive technologies, and relevant configuration.
- Tools, versions, settings, and the procedures used.
- Findings with the affected view and state, steps to reproduce, observed behavior, applicable requirement, and retest status.
- Any score calculation and its method, if you choose to include a score.
WCAG-EM alone usually does not produce a whole-product conformance claim. Any public evaluation statement has conditions and should identify the evaluated product and target. Be clear that sampled results apply to the sampled material and tested environment. W3C WCAG-EM 2.0 reporting guidance.
9. Treat aggregate scores carefully
A single score can hide which barriers remain, where they occur, and who is affected. W3C’s WCAG-EM 2.0 methodology says “there is currently no single metric that is known to address the required reliability, accuracy, and practicality.” If a score helps track work, publish how it was calculated and keep the underlying findings available so readers can understand and repeat the evaluation. W3C WCAG-EM 2.0.
10. Capture reproducible visual evidence
Screenshots can help show the visual state associated with a finding, such as clipped content, a missing focus indicator, or an error message. They are supporting evidence only: a static image cannot establish keyboard operability, screen-reader output, or whether a workflow is usable. Record the URL, viewport, state, and steps alongside any image so another evaluator can reproduce the context.
For a documented browser-based process, capture the relevant state, save the image with the finding, and retain the interactive test steps separately. If you use screenshots in reports, ensure the captured content does not expose personal or sensitive data.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It can capture a page as PNG, JPEG, WebP, or PDF; screenshots can help preserve visual evidence for an accessibility finding, while the accessibility evaluation itself still requires the methods above. The API accepts a URL in one GET request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners are accepted and removed before the shot; 60+ known consent platforms, newsletter popups, and chat widgets can be removed, and each step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - 1,000 screenshots per month are free with no card. Paid plans start at $5 for 3,000 screenshots; all features are on every plan.
Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.
Troubleshooting accessibility evaluations
| Problem | Likely cause | What to do |
|---|---|---|
| The automated scan reports no issues, but users encounter a barrier. | The relevant requirement needs judgment or interaction testing, or the affected state was not scanned. | Reproduce the task manually, check the applicable WCAG requirement, and add the missing state or workflow to the evaluation sample. |
| A tool reports many issues that do not appear to be failures. | A rule may need contextual interpretation, or the tool configuration does not match the evaluated page. | Verify each finding against the rendered state and requirement; document the tool rule, configuration, and disposition. |
| A finding cannot be reproduced by another evaluator. | The report omits navigation steps, account state, inputs, environment, or timing. | Record exact steps, test data that can safely be shared, browser and assistive technology versions, and the initial state. |
| The sample is large but misses a critical workflow. | Sampling focused on page count rather than distinct functionality and complete tasks. | Map product workflows first, then ensure the sample covers critical task paths and meaningful states. |
| Screen-reader output differs across tests. | Tests used different browser or assistive technology combinations, or dynamic state was inconsistent. | Declare a support baseline and repeat the same steps in each chosen combination, recording differences rather than merging them. |
| A score improved, but important barriers remain. | An aggregate score hides severity, context, or untested product areas. | Review issue-level findings and disclose the score formula, sample, and limits. |
| A public statement sounds broader than the evidence. | The report implies a whole-product conclusion from a sample or a limited environment. | Name the evaluated product, target, sample, environments, and exclusions, and narrow the claim to what was examined. |
Performance, reliability, and cost
The time and cost of an evaluation depend on product size and complexity, the number of workflows and states sampled, the accessibility expertise available, and the environments included. Tools can make some checks more efficient, but reviewing findings, manual evaluation, assistive technology testing, and user participation still require planning and time. A smaller, well-documented sample can be useful for a focused evaluation, but it cannot support claims about unexamined areas.
For ongoing work, integrate repeatable automated checks into development and release workflows where appropriate, then schedule manual and assistive technology review for meaningful changes and representative end-to-end workflows. Keep the sample and baseline stable when comparing results; record intentional changes so a difference in coverage is not mistaken for a change in accessibility. No universal percentage of WCAG coverage by automation is established by the cited sources, so avoid treating one as a planning guarantee.
Frequently asked questions
Can automated tools certify a site as WCAG conformant?
A tool result alone is not a complete evaluation. Combine tool findings with manual judgment and interaction testing, and state the scope and limits of the evaluation.
Should every page be tested?
Evaluate every view when practical. Otherwise, use a documented sample of distinct views and workflows, and do not imply that untested areas were evaluated.
Does WCAG-EM add requirements to WCAG?
No. WCAG-EM is a W3C methodology that supports evaluation; it does not add WCAG requirements.
Is input from people with disabilities a replacement for conformance evaluation?
No. It adds evidence about real use and complements systematic evaluation against the selected WCAG target.


