How to Test and Monitor Website Accessibility
Build a repeatable accessibility workflow with automated checks, manual evaluation, WCAG-EM, and ongoing monitoring—without mistaking a scan for conformance.
Test website accessibility throughout development, then repeat checks as pages, content, and components change. Combine automated tools with keyboard and other manual checks; use expert evaluation for consequential or complex services. For a formal WCAG conformance evaluation, follow W3C’s WCAG-EM methodology. A scan or score alone cannot establish that a site is accessible or conforms to WCAG.
1. Start with a repeatable evaluation plan
Decide what you are evaluating, which methods you will use, who will review findings, and how issues will be tracked. Include representative templates and important user journeys, not just the homepage.
- Set the scope. List the site or application, key page templates, critical workflows, and any restricted pages that need credentials.
- Choose applicable criteria. Identify the WCAG version and conformance level relevant to your work. Record that scope; do not infer it from a tool’s default settings.
- Pick complementary methods. Plan automated checks, manual evaluation, and expert review where appropriate.
- Record and triage findings. Track the page or component, observed barrier, reproduction steps, severity for users, owner, and fix status.
- Recheck fixes and changed areas. Keep a regression check in the development process and evaluate again when content, templates, components, or code change.
W3C WAI recommends evaluating early and throughout development so issues can be found when they are easier to address. See the W3C evaluation overview.
2. Add automated checks to development
Automated accessibility tools can identify some detectable issues, such as certain missing names or invalid relationships. They cannot determine whether every interaction is understandable, whether alternative text conveys the right meaning, or whether a complete user journey works with assistive technology. Different tools find different issues, so treat results as leads for evaluation rather than a complete verdict.
For a JavaScript project, axe-core can be run in acceptance tests. The example below uses Playwright and the axe Playwright integration. Install the packages using your project’s package manager, then save this as a test file in a project where Playwright is configured:
import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';
test('home page has no automatically detected serious or critical violations', async ({ page }) => {
await page.goto('http://127.0.0.1:3000/');
const results = await new AxeBuilder({ page })
.withTags(['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa'])
.analyze();
const seriousOrCritical = results.violations.filter(
violation => ['serious', 'critical'].includes(violation.impact ?? '')
);
expect(seriousOrCritical, JSON.stringify(seriousOrCritical, null, 2)).toEqual([]);
});
This is a sample automated gate, not a conformance test. The selected tags and impact threshold are project choices; review all findings, including lower-impact ones, and adjust the scope to match the product. A passing test means this run reported no violations in the rules and page state it checked. It does not mean the page or site is accessible.
Run checks against representative states: menus open and closed, validation errors, dialogs, tabs, and other interactive states. Automated analysis of only the initial page state may miss issues that appear after interaction.
3. Follow scans with manual checks
Use manual evaluation to examine issues that tools cannot reliably judge from markup alone. A practical pass includes:
- Keyboard: operate the page without a pointer. Check that controls can be reached and activated, focus is visible, and focus order makes sense.
- Structure: inspect headings, landmarks, labels, instructions, and relationships between controls and their messages.
- Zoom and reflow: enlarge content and check that information and actions remain available without avoidable loss or overlap.
- Content and context: judge whether link text, instructions, error messages, and text alternatives make sense in context.
- Assistive technology: test representative workflows with the screen reader and platform combinations relevant to your audience.
- Dynamic behavior: check dialogs, menus, updates, validation, and route changes for sensible focus handling and announcements.
Browser-based tools can give quick page-level feedback. DWP guidance names axe DevTools and WAVE as options; WAVE also describes page evaluation and offerings for collecting data across many pages. These are examples, not endorsements. Select a tool based on its method, scope, coverage, access requirements, and reporting needs. See the W3C evaluation tools directory.
For consequential or complex services, combine automated results with manual checks and a professional audit, as recommended in UK government accessibility guidance and DEFRA guidance.
4. Evaluate the whole site and formal conformance
A one-page checker answers a narrow question about the page and state it analyzed. Site-wide evaluation needs a deliberate sample or broader crawl, attention to shared templates, and access to relevant restricted content. A representative sample should cover distinct layouts, content types, and important workflows; repeated instances of the same template may be less informative than a page with a unique interaction.
Choose tools by method (automated detection, manual guidance, or simulation), scope (one page, a sample, full site or app), ability to reach password-protected content, guideline coverage, platform support, and output. W3C’s directory records differing combinations of these capabilities. It lists axe Monitor as an enterprise monitoring and reporting platform and axe DevTools among evaluation tools.
For a formal WCAG conformance evaluation, use the W3C Website Accessibility Conformance Evaluation Methodology (WCAG-EM). It provides a structured methodology for evaluating conformance and a report generator. Document scope, sample, process, findings, and limitations so readers understand what was evaluated. A scanner score by itself is not a WCAG-EM report or a conformance finding.
5. Monitor accessibility after launch
Monitoring makes repeated checks practical, especially when a site is too large to revisit page by page. It complements development and manual evaluation; it does not replace them. Repeat checks when templates, components, content, or code change, and include checks in release workflows where feasible.
There is no universal scan cadence established by the sources here. Set one based on release frequency, site size, and risk. A frequently changing service may benefit from checks on each release plus scheduled coverage of broader page samples. A smaller, slower-changing site may choose periodic scans and checks after meaningful content or template updates.
Assign an owner to review alerts, confirm whether reported issues affect users, prioritize fixes, and verify resolutions. Avoid treating issue counts or scores as the objective: focus on user barriers and whether important tasks can be completed.
6. Troubleshooting common accessibility testing problems
| Symptom | Likely cause | What to do |
|---|---|---|
| The scan passes, but users still report barriers. | Automated rules cover only detectable conditions, or the tested state omitted an interaction. | Reproduce the user journey, test keyboard and assistive technology behavior, and manually evaluate content and context. |
| The test fails on a page that works in a browser. | The test may run before content settles, use a different route or viewport, or expose a real issue only in that state. | Check the captured page state and test setup, wait for the relevant content, then verify the finding manually before changing code. |
| Only the homepage is covered. | The workflow has no page inventory or sampling strategy. | List templates, unique components, key journeys, and restricted areas; add representative pages to the evaluation plan. |
| Authenticated pages are missing from monitoring. | The checker cannot access the session or protected route. | Choose a workflow that supports authenticated content, configure access safely, and confirm which pages it actually reached. |
| Reports contain many repeated findings. | A shared component or template may repeat the same problem. | Group findings by root cause, fix shared code where appropriate, then verify representative instances and affected workflows. |
| A score changes between scans. | Page content, tool rules, crawl scope, or tested state may have changed. | Compare the exact pages, states, tool configuration, and findings; do not interpret the score change alone as a change in conformance. |
| A third-party widget fails checks. | The widget may have accessibility barriers outside your code. | Document the impact, check available configuration or alternatives, and include the issue in remediation and procurement decisions. |
7. Performance, reliability, and cost
Keep fast automated checks close to code changes, and reserve broader crawls and expert evaluation for workflows that can accommodate their wider scope. Large sites, authenticated routes, dynamic pages, and third-party content can increase setup and review effort. The sources do not establish a universal scan speed, detection rate, or cost comparison, so assess tools against your own scope and reporting needs.
For reliable comparisons over time, keep the tested URL set, credentials, browser configuration, rules, and page states consistent. Investigate missed pages and intermittent loading before treating a clean report as meaningful. Budget for human review and remediation as well as tool access: a tool subscription does not itself resolve barriers or establish conformance.
Or skip the browser setup
For screenshots used in visual review or issue reports, ScreenshotNeo is a website screenshot API and MCP server. It captures a URL as PNG, JPEG, WebP, or PDF; it is useful for visual evidence, but a screenshot does not test accessibility or prove WCAG conformance. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo accepts cookie or consent banners and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with response headers identifying the page verdict and billing. Its MCP server provides screenshot, page-info, and PDF tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, no card required.
FAQ
Can automated accessibility testing find every issue?
No. Automated checks find some detectable issues. Manual evaluation is needed to judge content, interaction, and user journeys.
Does a high accessibility score mean the site conforms to WCAG?
No. A score is an output from a particular tool and scope. Formal conformance work requires an appropriate evaluation method and documented scope.
How often should I monitor a website?
Choose a cadence based on how often the site changes, its size, and the impact of barriers. Recheck significant changes and keep checks in the release process where practical.
Should I test every page?
Cover templates, unique components, important journeys, and restricted areas. Use whole-site monitoring when page-by-page checking is impractical, while retaining representative manual evaluation.


