How to Test a Website’s Accessibility States with Screenshot Comparisons
Test meaningful interface states with screenshots, automated accessibility checks, accessibility-tree assertions, and manual keyboard review. Learn what each method can and cannot prove.
Test accessibility state by state. For each meaningful state, reach it through realistic browser interaction, wait until it settles, run an automated accessibility scan, assert important accessibility-tree structure, and compare a screenshot against a baseline created in the same rendering environment. Then review keyboard operation and focus manually. Screenshot comparisons show rendered pixels; they do not prove that controls are exposed with the right semantics or can be operated by people.
This guide uses Playwright Test, @axe-core/playwright, and Playwright’s ARIA snapshot assertions. The example is a runnable pattern to adapt to your application. It tests a page with a navigation menu, a dialog, and a form; replace the locators and expected structure with your own.
1. Choose the states worth testing
A page is not a single accessibility test target. List states users can actually encounter, based on the product’s interactions and risk. A practical starting set is:
- Default page, including its landmarks, headings, and primary controls.
- Keyboard focus on important controls, including visible focus indicators.
- Expanded navigation or another disclosure control.
- Open dialog, popover, or other modal interface.
- Form after invalid submission, with its error messages and any summary.
- Successful submission or confirmation state.
- Loading, empty, and failure states, where those states change what users can do.
Do not aim for an assumed exhaustive list of every possible combination. Cover each meaningful interaction and state transition, and add scenarios where a failure would block an important user task. Record how each state is reached so the test exercises the real interaction rather than setting hidden implementation state.
2. Set up Playwright and axe
Install Playwright Test and the axe integration in an existing Node.js project:
npm install --save-dev @playwright/test @axe-core/playwright
npx playwright install
Playwright’s installation and configuration can differ by project. Follow its official getting-started guide if your project does not already have a Playwright configuration. The test below assumes the application is available at http://127.0.0.1:3000 and that its browser-test server is started separately.
3. Write a state-based test
Create tests/accessibility-states.spec.js. Change the accessible names, routes, expected headings, and test data to match your application. The test scans each rendered state after the interaction that reveals it, checks selected accessibility-tree structure, and captures visual baselines.
const { test, expect } = require('@playwright/test');
const AxeBuilder = require('@axe-core/playwright').default;
async function expectNoAxeViolations(page, options = {}) {
let builder = new AxeBuilder({ page });
// Optional: limit a scan to a relevant region, for example:
// builder = builder.include('#main-content');
// Optional: select rule tags supported by axe-core:
// builder = builder.withTags(['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa']);
if (options.include) builder = builder.include(options.include);
if (options.tags) builder = builder.withTags(options.tags);
const results = await builder.analyze();
expect(results.violations, JSON.stringify(results.violations, null, 2)).toEqual([]);
}
test.beforeEach(async ({ page }) => {
await page.goto('/accessibility-demo');
await expect(page.getByRole('main')).toBeVisible();
});
test('default page has expected structure and appearance', async ({ page }) => {
await expectNoAxeViolations(page);
await expect(page).toMatchAriaSnapshot(`
- main:
- heading "Account settings" [level=1]
- button "Open help"
- button "Save changes"
`);
await expect(page).toHaveScreenshot('account-default.png');
});
test('navigation menu is accessible when expanded', async ({ page }) => {
const menuButton = page.getByRole('button', { name: 'Open navigation' });
await menuButton.click();
await expect(page.getByRole('navigation', { name: 'Main' })).toBeVisible();
await expect(page.getByRole('link', { name: 'Billing' })).toBeVisible();
await expectNoAxeViolations(page);
await expect(page).toMatchAriaSnapshot(`
- navigation "Main":
- link "Overview"
- link "Billing"
`);
await expect(page).toHaveScreenshot('account-navigation-expanded.png');
});
test('help dialog has accessible structure and appearance', async ({ page }) => {
await page.getByRole('button', { name: 'Open help' }).click();
const dialog = page.getByRole('dialog', { name: 'Help with account settings' });
await expect(dialog).toBeVisible();
await expectNoAxeViolations(page);
await expect(dialog).toMatchAriaSnapshot(`
- dialog "Help with account settings":
- heading "Account settings help" [level=2]
- button "Close help"
`);
await expect(page).toHaveScreenshot('account-help-dialog.png');
});
test('invalid form submission exposes errors', async ({ page }) => {
await page.getByRole('button', { name: 'Save changes' }).click();
await expect(page.getByText('Enter a display name')).toBeVisible();
await expect(page.getByLabel('Display name')).toHaveAttribute('aria-invalid', 'true');
await expectNoAxeViolations(page);
await expect(page).toMatchAriaSnapshot(`
- main:
- heading "Account settings" [level=1]
- textbox "Display name" [invalid]:
- text: Enter a display name
- button "Save changes"
`);
await expect(page).toHaveScreenshot('account-invalid-form.png');
});
The illustrative ARIA snapshots are intentionally focused. Match the snapshot syntax and locator names to your installed Playwright version and your application’s actual accessible tree. Assertions for text, state, and accessible names are often more maintainable when they target only the important parts of a large page.
4. Run the tests and review baselines
- Start the application in the same way your team starts it for browser tests.
- Run
npx playwright test tests/accessibility-states.spec.js. - On the first run, Playwright creates screenshot references for the configured browser and host environment. Review those images before accepting them as expected appearance.
- On later runs, inspect visual diffs and determine whether a change is intentional. Update a baseline only after review, for example with
npx playwright test tests/accessibility-states.spec.js --update-snapshots. - Keep the browser, operating system or container image, fonts, rendering settings, and execution mode consistent when generating and comparing references.
Playwright’s visual comparison documentation explains screenshot assertions and their environment sensitivity. A baseline is a reviewed reference, not an accessibility specification. A diff can flag a change worth investigating, but a pixel match cannot show that a control has a useful name or works from the keyboard.
5. Add explicit keyboard checks
Include keyboard operations in your test plan and perform human review as well. For example, a browser test can check that Tab reaches a control and that its focus indicator is visually apparent in a focused screenshot:
test('save button can be reached by keyboard and shows focus', async ({ page }) => {
await page.keyboard.press('Tab');
await page.keyboard.press('Tab');
const saveButton = page.getByRole('button', { name: 'Save changes' });
await expect(saveButton).toBeFocused();
await expect(page).toHaveScreenshot('account-save-button-focused.png');
});
The number of Tab presses depends on the page’s focus order. Prefer a deliberate test setup that starts from a known focus position and assert the intended control, rather than assuming a fixed count across unrelated pages. Also check that controls respond to the keyboard, focus moves into and out of dialogs appropriately, and focus remains visible as content changes.
Review interactions that appear on hover. The W3C WAI preliminary checks call out functionality available only on mouse hover and unavailable through keyboard focus. A screenshot of the hover state does not demonstrate that a keyboard user can reach the same content. See WAI Easy Checks for preliminary checks, including visible keyboard focus and hover-only functionality.
What each method tells you
| Method | What it can establish | Limit |
|---|---|---|
| Screenshot comparison | Whether rendered pixels differ from a saved reference in the tested environment. | It does not establish semantics, accessible names, keyboard access, or usability. Rendering changes can cause diffs. |
| Automated accessibility scan | Whether the configured rules found detectable violations in the rendered state that was scanned. | It finds only some problems and depends on the state reached and rules selected. Passing is not proof of conformance. |
| ARIA snapshot assertion | Whether selected accessibility-tree structure matches an expected template. | It is a structural check, not a complete usability assessment. Playwright’s snapshot matching is order-sensitive. |
| Manual keyboard, assistive technology, and inclusive user review | Whether people can reach, understand, and operate the interface in context. | It requires planned human review and coverage suited to the product and its risks. |
Playwright’s accessibility-testing guide recommends running scans after the interaction that reveals content and explains that automated testing finds only some accessibility problems. Its ARIA snapshots documentation describes accessibility-tree assertions and order-sensitive matching. Combine these forms of evidence instead of treating any one as a complete verdict.
Reporting results clearly
Keep findings distinct so a reviewer knows what was checked and what a failure means. For each state, record:
- State and route: the user action that revealed it and any relevant test data.
- Visual comparison: whether pixels changed, which region changed, and whether the change was accepted.
- Automated scan: rules or tags configured and the violations reported.
- Accessibility tree: the expected structure and any mismatch.
- Human review: keyboard and assistive technology observations, with follow-up issues.
This makes it less likely that a passing screenshot comparison will be read as proof of accessibility, or that a passing axe scan will be mistaken for a complete review.
Reliability and performance considerations
- Wait for the state, not an arbitrary pause. Assert that the expected menu, dialog, or error is visible before scanning and capturing. Use a delay only when the product behavior genuinely depends on elapsed time.
- Stabilize data and animation. Use deterministic test data and consider reduced motion or disabling known animations in the test environment where appropriate. Do not hide a meaningful state just to make a screenshot stable.
- Keep the environment consistent. Browser version, operating system, fonts, device scale factor, and rendering mode can affect pixels. Use the same environment for baselines and comparisons.
- Scope large pages carefully. Axe scans can target a relevant region with
include(), but page-wide scans may catch issues outside that region. Choose scope deliberately and document it. - Test the interaction path. Directly loading a URL that happens to render a dialog may miss a broken trigger, focus transition, or keyboard behavior.
- Manage baselines as reviewed artifacts. Store and update them with code review. Avoid mass updates that accept unexplained differences.
- Balance coverage and runtime. Each state adds browser work. Prioritize high-value states and share setup where safe, while keeping tests isolated enough that one state cannot leak into another.
These practices improve repeatability; they do not make a screenshot or automated scan a substitute for manual evaluation. WAI maintains a Web Accessibility Evaluation Tools List for finding evaluation tools, but the appropriate checks depend on the interface and review goals.
Troubleshooting
The screenshot fails even though the page looks unchanged
Cause: The test environment differs, content is dynamic, a font or image loaded at a different time, or an animation was captured mid-transition. Fix: compare the environment and diff, wait for specific content and fonts where relevant, stabilize test data, and handle animation deliberately. Update the baseline only after deciding that the rendered change is expected.
The accessibility scan reports violations only in an open menu or dialog
Cause: The scan is finally covering content that was absent from the default state, or the interaction revealed a real issue. Fix: keep the state-specific test, inspect the reported node and rule, and correct or document the issue according to your review process. Do not dismiss a finding merely because the default-page scan passed.
The scan passes, but keyboard users cannot reach an action
Cause: Automated rules do not detect every operability problem; a mouse-only trigger may not be represented as a detectable violation. Fix: add keyboard interaction checks and human review. Check whether the same content or action is available through keyboard focus and operation.
An ARIA snapshot fails after a seemingly harmless change
Cause: The accessible tree changed, or its order differs from the stored snapshot. Fix: inspect the current tree and decide whether the structure change is intended. Keep snapshots focused on meaningful landmarks and controls to reduce noise, but retain assertions for the semantics that matter.
The expected locator is missing after clicking
Cause: The trigger or accessible name differs, the interaction did not succeed, or the test queried before the state appeared. Fix: use role-based locators that match the actual accessible name, assert the trigger’s expected state, then wait for the revealed locator to become visible. Avoid using a timeout as a replacement for a state assertion.
A focus screenshot looks correct but focus is not actually on the intended control
Cause: A visual outline may be present for another reason, or the test’s assumed Tab order is wrong. Fix: assert toBeFocused() on the intended locator and test the actual keyboard path. Review the focus indicator visually as a separate check.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. A screenshot can help capture a state after your test or review flow reaches it; it does not replace the axe scan, accessibility-tree assertions, or human checks above. See the ScreenshotNeo docs for API options.
For example, capture a URL as WebP with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
FAQ
Does a passing screenshot comparison mean the page is accessible?
No. It means the rendered output matched its reference within the comparison settings. It says nothing conclusive about accessible names, keyboard operation, or assistive technology use.
Should every possible state get its own screenshot?
Capture states that represent meaningful interactions and risks. A practical state plan is based on the product’s actual flows, not every theoretical combination.
Can axe replace manual accessibility review?
No. An automated scan can find some detectable issues in the rendered state and rules it checks. Manual assessment and inclusive user testing remain part of a complete evaluation.


