ScreenshotNeo

BlogGuides

How to Plan Testing for a Design System

Build a risk-based test plan for reusable components, from acceptance criteria and automated checks to manual accessibility reviews and real service testing.

By the ScreenshotNeo team4 October 202611 min read

A useful design system test plan defines what each component promises, checks its documented states and user-facing behavior, and then verifies the assembled service separately. Combine focused code tests, task-based checks, automated accessibility rules, visual review, and manual accessibility testing. Passing tests for a component library do not certify every product that uses it.

This guide gives you a practical plan, a reusable test matrix, and a way to decide which findings block a merge. The exact browser and assistive-technology coverage should fit your product, audience, supported platforms, applicable law, and target accessibility standard.

1. Define the contract and risk

Before choosing tools, agree on what the system supports and what “working” means. For every component, document its purpose, public API, behavior, states, keyboard interaction, semantic requirements, responsive behavior, and known limits. This becomes the acceptance contract against which implementation and changes can be reviewed.

  • Scope: components, patterns, tokens, examples, supported browsers, viewports, input methods, and assistive technologies.
  • States: default, hover, focus, active, disabled, loading, empty, error, validation, expanded or collapsed, and any component-specific variants.
  • Behavior: keyboard operation, state transitions, form submission, validation, navigation, and expected focus placement.
  • Accessibility: semantics, names and instructions, focus visibility, contrast, status announcements, and applicable acceptance criteria.
  • Responsive expectations: content wrapping, reflow, overflow, touch targets, and behavior at supported viewport sizes.
  • Known limitations: constraints, workarounds, browser-specific issues, and cases where a consuming service must add behavior.

Name the accessibility standard and version, the jurisdiction or policy that applies, and when the requirement takes effect. Requirements vary and change; confirm current legal obligations and platform support rather than treating a compliance label as a complete test plan. GOV.UK’s accessibility strategy says legal or regulatory requirements take precedence over its general timing for adopting a newer standard. Its Service Manual describes GOV.UK Frontend as meeting WCAG 2.2 AA; that statement applies to that system and should not be generalized to another library or to services using it.

Prioritize risks by impact and reach. A keyboard trap or misleading error message in a shared component can affect many services, so it deserves early attention. Record the expected behavior, severity if it fails, and who can make or approve the fix.

2. Test behavior at more than one level

Use a testing pyramid: many quick, focused checks near the component, plus a smaller number of broader task tests. Each layer answers a different question.

Layer What it answers Good candidates Tradeoff
Unit Does isolated logic behave correctly? Formatting, state helpers, validation rules, conditional rendering Fast and easy to diagnose, but does not prove that a user can complete a task.
Component or integration Does the rendered component respond correctly to its inputs and interactions? Accordion expand/collapse, tab selection, error display, disabled behavior Covers markup and events more closely; still may miss problems in the full service.
Feature or task Can a person complete a meaningful flow? Choose a tab and find its content; enter invalid data and recover; submit a form Slower and harder to debug, so use for representative tasks rather than every permutation.
Service-level Does the consuming product work with its own content and composition? Real routes, app logic, CSS overrides, integrated forms and navigation Necessary to catch integration barriers; requires product-specific setup and data.

Cover documented variants and meaningful interactive states, not just one default render. Include edge cases such as empty, unusually long, localized, and validation content where relevant. Check supported responsive layouts and keyboard paths. Keep tests aligned with public behavior so a refactor can change internals without rewriting assertions that do not matter to users.

Make documentation examples executable where practical. GOV.UK’s accessibility strategy reports that, by May 2023, its process ran JavaScript in examples and checked every example for each component rather than only the first. That is a useful coverage pattern to consider, not a requirement that every system copy GOV.UK’s exact implementation.

3. Automate repeatable checks

Run fast, reliable checks locally and in continuous integration (CI), and make their scope visible. A practical CI sequence is:

  1. Install pinned dependencies and build the component library and examples.
  2. Run unit and component or integration tests for the changed code.
  3. Validate rendered example markup where suitable.
  4. Run automated accessibility checks on meaningful examples and states.
  5. Capture visual comparisons at agreed viewports and review any changed output.
  6. Publish reports and assign failures or approved exceptions to an owner.

Automated accessibility rules can flag some problems, but a clean report is not proof of conformance or usable interaction. The GOV.UK Design System strategy attributes to a 2017 GDS study the finding that automated tools found about 30% of issues in that study. That is a qualified historical finding, not a universal detection rate for every product or tool. Use automated results for repeatable detection and triage, then cover questions that need human judgment through manual review and user research.

GOV.UK describes using jest-axe and @axe-core/puppeteer for checks against examples, and its developer documentation describes a wrapper that can raise JavaScript errors and fail CI. Tool versions and repository workflows can change, so consult the project’s current documentation before adopting a particular setup. For every automated check, document what it covers, what it cannot determine, and any excluded rule or state with a reason and owner.

4. Review visual changes deliberately

Visual regression checks compare rendered output and flag differences across selected component states and viewports. They can help catch unintended changes to layout, spacing, typography, color, borders, focus indicators, and content wrapping. They do not decide whether a change is wrong: an intentional redesign also produces a difference.

  1. Choose representative browsers or rendering environments and viewport sizes.
  2. Capture important documented states, including focus and error states where relevant.
  3. Review diffs when a baseline changes; check whether the change is intentional and consistent with the component contract.
  4. Record who can approve a baseline update and what happens when reviewers disagree.
  5. Decide whether a visual check is informational, requires human approval, or blocks merging for a specific high-risk change.

GOV.UK’s developer documentation describes Percy screenshots running on each pull request, while noting that its visual check is not a mandatory merge condition and a reviewer approves or rejects highlighted changes. This is one workflow example, not a universal policy or product endorsement. ScreenshotNeo can capture example pages for visual review: see the ScreenshotNeo API documentation for supported capture options.

5. Test accessibility and usability with people and assistive technology

Manual accessibility testing covers interaction and perception that automated rules cannot settle. Select combinations based on your users, supported platforms, and risk, and record the environment so a finding can be reproduced.

  • Operate the component using only a keyboard; verify logical order, visible focus, activation, and escape or dismissal behavior.
  • Inspect the rendered HTML and accessibility tree for appropriate structure, names, roles, and state.
  • Use screen readers and screen magnifiers on supported browser and operating-system combinations.
  • Check relevant high-contrast or display modes, zoom, reflow, and speech input.
  • Review whether instructions, labels, errors, status changes, and focus movement make sense during an actual task.
  • Involve disabled participants in research when the complexity, sensitivity, or uncertainty of the feature warrants it.

Keep a testing record with browser, operating system, assistive technology, versions where useful, input method, component state, steps, expected and actual behavior, and evidence. GOV.UK’s strategy describes recording browser and assistive-technology combinations in a testing template. Choose your own matrix to match your audience and support commitments; there is no single combination list that proves usability for everyone.

6. Test services that consume the system separately

Once library checks pass, test real products that use the components. Content, composition, app logic, custom styles, JavaScript enhancements, and surrounding HTML can introduce barriers even when the shared component works correctly. The GOV.UK Service Manual puts the distinction plainly: “Using the GOV.UK Design System in a service does not immediately make that service accessible.”

Test the assembled interface and representative end-to-end tasks in the service’s own context. Include design and prototype reviews before production as well as checks of the implemented code. When a service reports a problem, identify whether it belongs in the shared component, its documentation, or the service’s integration so that the right team can act.

7. Decide how findings affect a merge

Set failure policy before a release is under pressure. Fast, deterministic checks such as unit tests can usually block a merge when they fail. A visual difference often needs human review because it may be intentional. Manual accessibility findings need clear severity, evidence, ownership, and a decision path; do not silently ignore a disputed report or turn every uncertain automated flag into a blocker.

Check or finding Possible merge policy Decision to document
Deterministic unit or behavior failure Block when a documented contract is broken Who owns the fix and whether the expected behavior needs correction
Automated accessibility violation Block confirmed, high-impact violations; route uncertain findings for review Who confirms impact and records an exemption, if justified
Visual diff Require review for changed baselines; block only where policy says the change must be approved first Who approves the change and updates the baseline
Manual accessibility or usability issue Prioritize by severity and risk; define escalation for release-critical barriers Evidence, affected users and tasks, owner, due date, and rationale for any accepted risk

Store reports, exemptions, test context, and decisions somewhere maintainers can find them alongside normal development work. Review the plan when the public API, browser support, standards, audience, or risk changes. Exceptions should name the covered scope, reason, approver, and a date or condition for review.

Reusable design system test matrix

Start with this compact matrix, then add rows for your components and real risks. A matrix is a planning and accountability aid; it is not evidence that an untested combination is safe.

Component or task State and risk Acceptance criterion Method Environment Owner and frequency Failure policy or exception
Accordion Expanded and collapsed; keyboard access State is operable and conveyed; focus remains understandable Component interaction, automated rules, keyboard and screen-reader review Supported browser and selected assistive technology Component maintainer; CI plus scheduled manual review Block contract failures; record any scoped exception and rationale
Form field Empty, invalid, long error; recovery Instructions and error are associated and understandable Unit, task test, automated check, manual task review Representative browser, viewport, and input method Form owner; CI and before release Escalate confirmed barriers by severity
Navigation task Small viewport and keyboard use People can reach and identify destinations Feature test, visual review, manual keyboard check Supported viewport and browser Pattern maintainer; pull request and release Review visual changes; block broken task behavior

Using screenshots in the visual review loop

A screenshot is useful evidence for appearance at a particular URL, viewport, and state. It cannot establish keyboard behavior, semantic correctness, screen-reader usability, or service-wide accessibility. Treat images as one artifact in the plan, alongside code and human review.

For repeatable captures, keep the URL, viewport, device scale, wait condition, state setup, and any custom CSS or JavaScript consistent. Capture the same state before and after a change, and review diffs with context. If dynamic content changes between runs, make it deterministic where possible or exclude the unstable region deliberately and document why.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. For component documentation or a visual review capture, the simple request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the API docs for options such as viewport and device presets, full-page or selector capture, waiting for a selector or network idle, custom CSS and JavaScript, and signed async jobs. This capture is a visual artifact, not an accessibility test.

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, with page verdict and billing information in response headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month, no card required.

Troubleshooting a design system test plan

Tests pass, but the service is still inaccessible

Cause: library checks do not cover the service’s content, composition, CSS overrides, JavaScript, or complete tasks. Fix: add service-level task checks and manual reviews in the consuming product; assign the finding to the team that owns the source.

An automated accessibility scan reports no issues

Cause: automated rules cover only problems that can be evaluated from the rendered page and configured rules. Fix: retain keyboard, assistive-technology, visual, and task-based checks. Treat the scan as one layer, not a conformance certificate.

A visual test fails on every run

Cause: dynamic data, animation, fonts, timing, browser differences, or an inconsistent viewport can create unstable images. Fix: stabilize test data and capture timing, wait for required content, use consistent rendering settings, and isolate or document genuinely variable regions. Do not approve a baseline blindly.

The suite is too slow to run on every change

Cause: expensive end-to-end and cross-platform checks are being used for cases that can be covered by fast local tests. Fix: keep focused unit and component checks in the quick feedback path; run a smaller representative task suite in CI and schedule broader manual or platform coverage at risk-appropriate intervals.

A documented example fails while the component test passes

Cause: the example may use different props, content, JavaScript, or composition than the isolated test. Fix: render and check documentation examples, including their scripts and meaningful variants, and repair either the example or the component contract.

Reviewers disagree about whether a finding is a blocker

Cause: severity, evidence, ownership, or merge policy was left implicit. Fix: reproduce the issue, record affected tasks and users, apply the agreed risk criteria, and have the named adjudicator document the decision and any time-limited exception.

FAQ

Does a passing design system test suite prove every product using it is accessible?

No. It provides evidence about the library and tested examples. Each consuming service needs its own checks of content, integration, behavior, and user tasks.

Should every documented example be tested?

Test every example that represents supported behavior or a meaningful configuration, especially when it includes JavaScript or distinct semantics. Remove or clearly label examples that are illustrative rather than supported.

Can automated accessibility tools replace manual testing?

No. They are useful for repeatable rule checks, but do not establish that interaction, content, or tasks work for people using assistive technology.

Should visual regression failures block every pull request?

That depends on risk and review capacity. Define whether a diff requires approval, which changes block, and who may update baselines.