How to Build an Accessibility Testing Strategy
Build a repeatable accessibility evaluation process across planning, design, development, release, and maintenance—with clear scope, representative sampling, and human review.
A useful accessibility testing strategy combines ongoing developer checks, automated tools, structured conformance evaluation, expert manual review, and input from people with disabilities. Start by defining the product, user journeys, content, and WCAG conformance target you need to evaluate. Then map the product, select a representative sample when full coverage is impractical, investigate findings with human judgment, assign remediation, and retest.
Accessibility evaluation is a lifecycle practice, not a final-release scan. The W3C advises evaluating early and throughout development so issues can be found while design and implementation are still changing. The specific plan depends on your product, team, release process, and applicable requirements.
1. Integrate evaluation throughout the product lifecycle
Include accessibility work in planning, design reviews, development, content production, quality assurance, release decisions, and recurring maintenance. Each stage catches different problems: a design review can identify a confusing interaction before it is built, while implementation checks can uncover markup and keyboard issues in the actual product.
- Planning: identify the product surfaces, user journeys, target conformance level, and people responsible for evaluation.
- Design: review interaction patterns, content structure, focus behavior, and states before implementation.
- Development: run relevant automated checks and perform focused manual checks as components and features change.
- Quality assurance: evaluate representative workflows with the implemented product and assistive technologies relevant to your users and platform.
- Release and maintenance: record what was evaluated, retest fixes, and revisit areas affected by changes.
W3C describes accessibility as something to integrate “from the beginning and throughout the project lifecycle.” This is practical guidance for organizing work; it does not create additional WCAG requirements. W3C: Evaluating Web Accessibility
2. Define the scope and goal
Write down what the evaluation is meant to answer and what it covers. A plan for internal improvement may differ from an evaluation supporting procurement, a release decision, ongoing monitoring, or an external conformance report.
Specify:
- Product and version: name the website, application, service, or other digital product and the version or release under review.
- Surfaces and user journeys: include the relevant platforms, authenticated areas, key tasks, and important states.
- Content: identify formats such as HTML, documents, or other content types that users need.
- Evaluation goal: state whether the work is exploratory, a structured conformance evaluation, a release check, or ongoing monitoring.
- Target: identify the WCAG version and conformance level chosen for the evaluation, if applicable.
- Constraints and exclusions: document any platform, access, language, or sampling limits.
The applicable target may depend on a contract, policy, or jurisdiction. The target cannot be determined for every organization from general methodology guidance, so confirm the requirements that apply to your situation. Keep the distinction clear between a WCAG requirement, a test technique, a team practice, and a usability activity.
WCAG-EM is a methodology for evaluating conformance with WCAG; it does not define extra WCAG requirements or replace the standard. The current W3C methodology, WCAG-EM 2.0, extends the earlier website-focused method to apps and other digital products. W3C: WCAG-EM Overview · WCAG-EM 2.0 methodology
3. Explore the product before choosing tests
Build an inventory before selecting pages or screens. A list of URLs alone may miss important functionality, states, technologies, or restricted areas.
For each in-scope product, map:
- Common views, screens, and repeated templates.
- Essential user tasks and the steps required to complete them.
- Repeated components, such as navigation, dialogs, forms, and media controls.
- Interactive states, including validation errors, expanded content, loading, empty, and success states.
- Content types and technologies used to deliver the experience.
- Authenticated or password-protected areas, when they are in scope.
- Platform-specific experiences that could change interaction or accessibility behavior.
This exploration helps prevent a review from covering only the easiest or most visible part of the product. It also gives your team the context to select suitable tools and a meaningful sample.
4. Choose a representative sample
Evaluate the whole product when that is feasible for the goal. When it is not, document a deliberate sample that reflects the product’s important differences. There is no universal number of pages or screens that guarantees adequate coverage.
Build the sample around:
- Common views and high-use or essential user journeys.
- Different page, screen, and content types.
- Distinct interaction patterns and important states.
- Technologies and formats the product relies on.
- Areas with known issues or meaningful implementation differences.
Consider how consistent the implementation is, what prior automated and manual evaluations found, and how much confidence the evaluation needs to provide. WCAG-EM 2 notes that a higher level of confidence often calls for a larger sample. If shared components are implemented consistently, representative coverage may be more informative than selecting many near-identical views; if implementations vary, expand the sample to reflect that variation.
Record how the sample was selected and what it leaves out. The findings apply to the evaluated scope and sample; they should not be presented as proof that unexamined parts are conformant.
5. Combine automated checks and human evaluation
Automated tools can surface potential problems quickly and support a reviewer’s work. They cannot check every accessibility aspect or determine accessibility by themselves. Human judgment is required, and tools can produce false or misleading results. Treat a scan result as evidence to investigate, not as a complete accessibility verdict. W3C: Selecting Web Accessibility Evaluation Tools
A practical evaluation can include these complementary activities:
| Activity | What it contributes | Limit to keep in mind |
|---|---|---|
| Automated checks | Find some machine-detectable issues and help teams check frequently. | Cannot evaluate every accessibility aspect or certify conformance on their own. |
| Manual expert review | Lets an evaluator inspect context, behavior, and issues that need interpretation. | Depends on suitable expertise, scope, and evaluation time. |
| Assistive technology checks | Can reveal how relevant product flows work with the technologies being evaluated. | The setup and coverage should fit the product and its users; a limited setup does not represent every experience. |
| Evaluation with people with disabilities | Adds insight into actual use and barriers that checklists or automated scans may miss. | It complements standards-based evaluation; it does not by itself establish WCAG conformance. |
| Structured conformance evaluation | Organizes scope, exploration, sampling, evaluation, and reporting against a chosen WCAG target. | Its conclusions are bounded by the declared scope, sample, methods, and evaluator knowledge. |
Choose tools for the role they will play. Compare their purpose, product and format coverage, supported standards, scan scope, access to restricted areas, workflow, reporting, cost, platform needs, language support, and accessibility of the evaluation tool itself. Tool coverage varies across websites, applications, HTML, EPUB, ARIA, CSS, SVG, PDF, and other formats. Some teams will need more than one tool, and a tool’s listings and capabilities can change over time. W3C’s guidance is to choose for the team’s process, product complexity, and evaluation needs—not to assume one tool fits every case.
6. Match the work to evaluator expertise
Effective evaluation calls for people who can interpret the applicable accessibility standards, inspect accessible design and implementation, use relevant assistive technologies, and understand how people with disabilities interact with digital products.
Involve people with disabilities where possible. Their participation can help the team understand real-world experiences and barriers that a checklist or automated result may not reveal. Plan this work respectfully: make the product and tasks accessible to participants, provide appropriate context, and record what was evaluated. User evaluation is valuable evidence about experience, but it does not replace a structured conformance evaluation.
7. Record findings so the team can act
A finding should give the team enough information to reproduce, understand, and address the issue. For each finding, record:
- Product, version, platform, and relevant location or user journey.
- Evaluation scope, sample, and method used.
- The relevant WCAG criterion or a clear description of the observed barrier, where applicable.
- Evidence and reproduction steps, including the state or interaction involved.
- Whether the result is confirmed, needs investigation, or could not be evaluated.
- An owner, remediation status, and retest outcome.
Also report the parts that were not evaluated, the sampling approach, tool limitations, and any access or setup constraints. WCAG-EM includes a report tool to help structure evaluator input; it does not run the checks for you. A report is useful when another person can understand what was examined and where its conclusions stop.
8. Prioritize fixes and retest
Assign findings to people who can address them, connect fixes to the affected components or journeys, and retest corrected issues. When the same issue appears in a shared component, check the other places that use it. Add recurring checks to the workflow so fixes are less likely to be lost in later changes.
Your organization can define a severity or release-gate process that fits its users and product. The sources cited here do not prescribe one universal ranking formula or release gate, so document your chosen process as a team decision rather than a W3C mandate.
9. Keep the strategy useful over time
Review the plan when the product, technology, content, evaluation goal, or team workflow changes. A strategy becomes stale if it repeatedly scans only familiar surfaces while new features and states go unexamined.
At a regular review, ask:
- Do the scope and target still reflect the product and evaluation goal?
- Do the selected views and journeys still represent the current experience?
- Are new formats, platforms, restricted areas, or interaction patterns included?
- Do evaluators have the right knowledge and access?
- Do reports result in owned fixes and retesting?
- Are chosen tools still appropriate for the product and team workflow?
Or skip the browser setup
If your strategy includes capturing pages for visual review, examples, or records, ScreenshotNeo can return a screenshot or PDF from one request. Its website screenshot API and MCP server can accept cookie or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether it was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
For example, this cURL request saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
See the ScreenshotNeo API documentation for request options. These examples capture a rendered page; a screenshot is not an accessibility evaluation and cannot establish conformance.
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
Performance, reliability, and cost considerations
Set evaluation frequency and scope to match the product’s release cadence and risk. Frequent automated checks can help surface issues as changes land, while representative manual evaluation and user involvement require planned evaluator time. No single scan frequency or sample size fits every product; record the reasoning behind yours.
For reliability, preserve the product version, access conditions, sample, tools, and methods used so findings can be understood and retested. Treat tool output as a lead that may need confirmation. If a scan cannot access a protected area or a particular state, document that gap and arrange an appropriate evaluation path instead of treating the missing result as a pass.
Budget for tool licenses or services, evaluator expertise, access and setup, participant involvement where used, remediation, and retesting. Compare tools on total fit and coverage rather than price alone. The research sources do not establish universal costs or a guaranteed savings figure.
Troubleshooting common strategy problems
| Problem | Likely cause | What to do |
|---|---|---|
| A scan reports no issues, but users still encounter barriers. | Automated checks cannot examine every accessibility aspect. | Use manual evaluation and relevant user evaluation alongside tools; investigate the specific journey and state. |
| The report claims the whole product conforms based on a few pages. | The sample and its limits were not stated clearly. | Define the scope and target, document sample selection and exclusions, and qualify conclusions to match the evidence. |
| Important functionality is missing from the evaluation. | The team selected pages before mapping tasks, states, content, and technologies. | Explore the product first, then update the sample to include essential workflows and distinct interaction patterns. |
| A tool cannot inspect a page or flow. | The product may require authentication, a particular state, or a platform-specific setup the tool does not cover. | Check tool scope and access needs. Arrange the required access or use a suitable manual evaluation method, then document any remaining gap. |
| Two tools return conflicting results. | Tools can use different checks or provide results that need interpretation. | Review the underlying content and method, reproduce the issue, and use informed human judgment to determine the finding. |
| Findings recur after they were fixed. | Fixes were not retested across reused components or later changes. | Retest the correction, check other uses of shared components, and add an appropriate recurring check to the workflow. |
| The team cannot agree on a sample size. | A fixed page count is being treated as universal. | Base the sample on product consistency, key functionality, content and technology diversity, prior findings, and the confidence needed; document the rationale. |
| A screenshot is being used as proof of accessibility. | Visual evidence is being confused with a full accessibility evaluation. | Use screenshots only for the review task they support. Evaluate interaction, content, and relevant criteria with appropriate methods and expertise. |
Frequently asked questions
Does WCAG-EM create additional accessibility requirements?
No. WCAG-EM is a methodology for evaluating conformance with WCAG; it does not add requirements to WCAG.
Can an automated accessibility score prove conformance?
No. Automated tools can help find potential issues, but human judgment and evaluation methods suited to the product are required.
Does testing with people with disabilities replace a conformance evaluation?
No. It contributes evidence about real product experience and complements standards-based evaluation.
How many screens should an evaluation include?
There is no universal count. Choose coverage based on the product’s variation, essential tasks, content and technologies, prior findings, and the confidence the evaluation needs.
What is the role of WCAG-EM 2.0?
It is W3C’s current evaluation methodology, extending the earlier website-focused method to apps and other digital products. Consult the methodology for the process and sampling guidance.


