ScreenshotNeo

BlogGuides

How to Build Accessibility Testing into Your Team’s Workflow

Build accessibility checks into planning, code review, CI, manual QA, and release follow-up. Learn what automation catches, what people must evaluate, and how to close the loop.

By the ScreenshotNeo team4 October 202611 min read

Build accessibility evaluation into the product lifecycle: set a scope and target during planning, include checks close to code changes and in CI, schedule manual evaluation of real interactions, involve people with disabilities, and record findings until fixes have been rechecked. Automated scans help catch some detectable issues and regressions; they cannot establish that a product is accessible or prove conformance.

This workflow applies to websites, web apps, mobile apps, and other digital products. Adapt the tools and test coverage to your platform, product risks, and established conformance target. For a structured evaluation, W3C’s [WCAG Evaluation Methodology (WCAG-EM)](https://www.w3.org/WAI/test-evaluate/conformance/wcag-em/) describes five stages: define scope, explore the product, select a sample, evaluate it, and report findings. WCAG-EM supports WCAG; it does not add requirements to the standard.

1. Set the evaluation scope and target

Make the scope concrete before choosing scanners or test gates. Agree on:

  • Product and platform: website, web app, mobile app, kiosk, or another digital product.
  • Surfaces: pages or views, authenticated areas, embedded content, and supported technologies.
  • Important flows: sign-up, search, checkout, account management, or other tasks users need to complete.
  • States: validation errors, dialogs, menus, loading, empty and success states, expanded content, and responsive layouts.
  • Target: the WCAG version and conformance level the organization has actually selected, plus any applicable jurisdictional or contractual requirements.
  • Ownership: who evaluates, triages, fixes, approves exceptions, and verifies corrections.

Do not claim a conformance commitment until the target and applicable requirements are established. Standards, legal obligations, and procurement terms may differ by product and jurisdiction; get qualified advice where needed.

For web teams, make an inventory of templates and unique interaction patterns, not just URLs. A dozen pages that share one template may have less distinct coverage than a single complex checkout flow.

2. Map views, functionality, and user journeys

Explore the product the way users encounter it. Identify key views, content types, designs, technologies, and functions. Include states that appear only after interaction or after an error. A static screenshot or a scan of the public landing page will not cover a keyboard-operated menu, a form error, or a screen-reader announcement after a save.

When complete coverage is not practical, choose a documented representative sample. Include distinct templates and components, high-value user flows, and unusual or high-risk interactions. WCAG-EM also describes structured and random sampling. State what was sampled, what was excluded, and why; a sample is not an exhaustive audit.

3. Put automated checks near code changes

Use automation where it can give quick feedback: editor or lint checks, component or unit tests, browser-based end-to-end checks, and CI or pull-request builds. Pick a layer that fits the stack and test the rendered states your users actually reach. A code-level rule can catch some patterns before runtime; a browser test can inspect a rendered page after a flow has opened a dialog or displayed validation errors.

One browser-test pattern is to run an accessibility rule engine against a rendered page after navigating to the state under test. For example, with Playwright and the axe Playwright integration installed in the project, a test can look like this:

import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';

test('checkout has no detected critical or serious violations', async ({ page }) => {
  await page.goto('http://localhost:3000/checkout');
  await page.getByLabel('Email').fill('reader@example.com');
  await page.getByRole('button', { name: 'Continue' }).click();

  // Scan the state reached by the flow, not only the initial page.
  const results = await new AxeBuilder({ page }).analyze();
  const serious = results.violations.filter(violation =>
    ['critical', 'serious'].includes(violation.impact ?? '')
  );

  expect(serious, JSON.stringify(serious, null, 2)).toEqual([]);
});

This example illustrates one integration pattern, not a complete audit or a universal blocking policy. Configure the test runner and packages according to their current documentation and your project’s supported versions. Add separate tests for other important views and states; a scanner only evaluates the DOM and state it receives.

Decide how CI findings affect a build

A pull-request check can report results, fail on newly detected findings, or use another agreed policy. Choose deliberately:

  • For a new product or a cleaned-up baseline, blocking on agreed severity levels may prevent regressions.
  • For a mature product with existing findings, distinguish new issues from the baseline so teams can improve without hiding new regressions in a large backlog.
  • Document exclusions and suppressions with a reason, owner, and review date. Do not suppress a finding merely to make the build green.
  • Keep manual evaluation in the release workflow even when automated checks pass.

Microsoft’s sample repository demonstrates automated axe checks in CI and pull-request builds and describes configuring builds to fail on results. That is an implementation example, not a W3C requirement or a policy every team should copy unchanged.

4. Add structured manual evaluation

Automation has limited reach. W3C says no tool alone can determine whether a site meets accessibility standards; knowledgeable human evaluation is required. Microsoft likewise notes that interactive barriers may not be found by automated tools. Reserve time for people to evaluate operation and meaning, not just scan output.

Keyboard and interaction checks

  • Complete the key flow using a keyboard only. Check Tab and Shift+Tab order, visible focus, and whether focus gets trapped or lost.
  • Use the expected keyboard interaction for controls, menus, dialogs, and composite widgets. Confirm a user can open, operate, and leave each interaction.
  • Trigger invalid, loading, success, and empty states. Check whether focus and status changes make sense and whether errors are discoverable and actionable.
  • Check that controls have understandable names and that instructions do not rely only on color, position, or pointer gestures.

Display and reflow checks

  • Resize the viewport and zoom the page. Look for content that becomes obscured, clipped, overlapped, or requires avoidable horizontal scrolling.
  • Inspect text spacing, focus visibility, and layout changes at the display sizes your product supports.
  • Where relevant, evaluate high-contrast or forced-colors presentation and user display preferences.

Assistive technology checks

  • Test representative workflows with screen readers and the browser or operating-system combinations your users rely on.
  • Where relevant to the product, include voice recognition and other assistive technologies.
  • Check that reading order, control names, instructions, errors, state changes, and announcements match what a user needs to act.

These are examples of manual coverage, not an exhaustive WCAG checklist. Use the selected WCAG target and qualified evaluators to define criterion-by-criterion coverage. People with disabilities and assistive technology users should inform evaluation; an individual’s experience is valuable input, but no one participant represents every disabled user.

5. Include accessibility in planning, design, and development

Evaluation is more effective when it begins before release. Include accessibility questions in refinement and design review: what is the keyboard interaction, how will errors be presented, what name will a control expose, and what happens when content is enlarged? Review component patterns and content as they are created, then verify them in the integrated product.

Give designers, developers, QA, and product owners a shared route for raising issues. A short checklist at design handoff or pull-request review can prompt teams to consider semantics, focus behavior, contrast, labels, and error handling. Treat that checklist as a prompt for evaluation rather than a substitute for testing.

6. Record findings and verify fixes

A useful evaluation record lets another person understand what was checked and what remains. Record:

  • scope, target, product version, and evaluation date;
  • views, flows, states, and sample selection, including exclusions;
  • tools, browsers, assistive technologies, and manual steps used;
  • findings with reproducible steps, expected and actual behavior, impact, and affected users or tasks where known;
  • owner, priority, status, remediation, and verification result.

Assign each actionable finding and track it to a fix or a documented decision. Re-run the relevant automated test and repeat the manual step that exposed the issue. Then check nearby components or flows that may share the same implementation.

The W3C WCAG-EM Report Tool can structure a report from supplied evaluation results. It formats the information you provide; it does not run the checks or independently verify the findings.

7. Repeat throughout the lifecycle

Make accessibility evaluation recurring: during design and implementation, in pull requests, in scheduled manual QA, and after significant changes. Add a focused check when shared components, navigation, authentication, or other critical flows change. A final review or periodic monitoring can add assurance, but deferring all evaluation to a release gate makes problems harder to find and fix.

When What to do Record
Planning Set scope, target, critical journeys, owners, and evaluation approach. Target and product inventory.
Design Review interaction patterns, content, states, and display behavior. Design decisions and open questions.
Development Use code-level checks and test components and representative states. Findings and relevant automated output.
Pull request / CI Run browser checks against changed or critical flows; apply the team’s documented gate. Build result, new findings, reviewed suppressions.
Manual QA Evaluate keyboard use, zoom/display changes, assistive technology, and user tasks. Steps, environment, observations, and gaps.
Release and follow-up Review unresolved issues, verify fixes, and schedule further evaluation. Status, owners, sample limitations, and next review.

Choosing tools and support

Compare options against the product and workflow rather than treating a scanner score as a measure of accessibility. Consider:

  • platform coverage: web, mobile, desktop, or other products;
  • where checks run: editor or linter, code tests, browser inspection, or CI;
  • test type: automated, guided or semi-automated, and manual;
  • fit with the framework, build system, and CI environment;
  • how results support triage, remediation, and repeatable reports;
  • the rules and standards covered, and the gaps that still need human evaluation;
  • whether the plan includes assistive technology and input from people with disabilities.

W3C’s evaluation methodology is tool-independent. Deque documents examples of CI integration and different testing layers, but those examples do not establish that one vendor is best or that a particular product is required. Teams can also seek accessibility expertise, audit, consulting, or training when internal experience is not enough. Choose support based on the skills and evaluation scope the team needs.

Using screenshots as a visual review aid

Visual captures can help reviewers compare page appearance across viewports or preserve evidence of a rendered state. They cannot establish semantic correctness, keyboard operability, screen-reader behavior, or conformance. Use them alongside the interactive and assistive-technology evaluation above. If your team needs reliable page captures for review records, [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server; it complements accessibility evaluation but does not replace it.

Or skip the browser setup

For a screenshot used in visual review, ScreenshotNeo takes one GET request and returns an image. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
  • Cookie banners are accepted and removed before the shot; newsletter popups and chat widgets are removed too. Each step can be turned off.
  • Bot checks, blank pages, and failed loads are never billed; response headers identify the page verdict and billing outcome.
  • An MCP server provides screenshot tools for AI agents, including take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

Performance, reliability, and cost

Accessibility checks add work to a test pipeline, so put fast feedback close to changes and reserve deeper manual evaluation for planned QA and high-impact flows. Keep automated tests focused on useful views and states; avoid rescanning the same unchanged surface in every layer without a reason. A screenshot can preserve visual evidence, but image capture has separate network, rendering, and service costs and does not replace a browser interaction test.

Reliability depends on repeatable setup: use stable test data, wait for the state under test, record the browser and assistive technology used, and distinguish product defects from test-environment failures. Review flaky tests rather than silently retrying or suppressing them. For cost, account for staff time to triage and fix findings as well as any tool or service fees. The dossier provides no comparative pricing or performance data for accessibility tools, so evaluate those against your own scope and workflow.

Troubleshooting common workflow problems

Problem Likely cause What to do
CI passes, but keyboard users cannot complete a flow. The automated rules did not evaluate keyboard operation or the relevant interaction state. Add a manual keyboard test and a browser test that reaches the state; verify focus movement and exit behavior.
A scanner reports no violations, but a screen reader experience is confusing. Automated checks cannot judge every reading order, announcement, or task-level issue. Reproduce the task with the relevant assistive technology and record the observed behavior for a knowledgeable evaluator.
Every pull request fails on a large legacy backlog. The gate treats all existing findings as newly introduced. Establish and review a baseline, then apply a transparent policy to new findings and track existing ones to owners.
Developers suppress findings to unblock builds. The gate lacks an accountable exception process or results are noisy. Require a reason, owner, and review date for exclusions; validate whether the finding is a true issue before suppressing it.
A page scan misses a menu, dialog, or validation error. The test scanned only the initial page or did not trigger the relevant state. Drive the flow to each important state before scanning; add a specific test for each distinct interaction.
Teams disagree about whether a sample represents the product. Sampling criteria and exclusions were not documented. Record templates, flows, and states represented, along with exclusions and rationale; describe the limits of the evaluation.
A report is mistaken for a conformance result. Report generation was confused with evaluation. Keep the supplied evidence, scope, steps, and evaluator conclusions with the report. A report tool structures input; it does not perform the checks.

FAQ

Does an accessibility scan certify a site as accessible?

No. A passing scan only describes the checks and page states the tool evaluated. Conformance requires evaluation against the chosen target, including knowledgeable human review.

Should every accessibility finding block a pull request?

That depends on the project’s baseline and release policy. Define which findings block, how existing issues are tracked, and how exceptions are reviewed; apply the policy consistently.

Is WCAG-EM another accessibility standard?

No. It is a methodology for evaluating conformance to WCAG, not an additional set of WCAG requirements.

Can a small team use this workflow without a dedicated accessibility specialist?

A small team can establish scope, run automated checks, and schedule manual evaluation. For evaluation requiring expertise the team does not have, seek qualified accessibility support and involve people with disabilities.

Sources