How to Test Design Systems
Test design systems with a repeatable mix of component, interaction, visual, accessibility, and end-to-end checks. Learn what to cover and how to run it in CI.
Test a design system by checking its components in representative states, then combine render, interaction, visual, accessibility, and integration checks in continuous integration (CI). No single test type proves that a component is correct: a screenshot can reveal a visual change but cannot verify keyboard behavior, while an isolated interaction test may miss an application integration problem.
Use component stories as reusable examples of the public contract: supported props, variants, responsive modes, meaningful empty and populated states, errors, and key interactions. Run fast checks on changes to shared components, review visual differences, and reserve end-to-end (E2E) tests for risks that need the full application stack.
1. Define what the design system promises
Start with the component contract, not an exhaustive grid of every possible prop combination. For each component, list the options consumers are meant to use and the user situations where getting the result wrong would matter.
- Inputs: supported props, tokens, slots, and content types.
- Variants: size, emphasis, intent, icon placement, and other documented choices.
- States: default, disabled, loading, focused, selected, invalid, empty, populated, and error states where applicable.
- Layouts: supported breakpoints and narrow or wide content.
- Interactions: typing, selection, opening and closing, submitting, dismissing, and keyboard navigation.
- System-level promises: localization, zoom, design-token use, and parity between design files and code.
Prioritize representative combinations from the public API and consequential situations. For example, test a disabled destructive button if that is a supported combination; do not generate a test for every arbitrary mix of unrelated options simply to increase a count.
2. Use component stories as test cases
A story renders one example of a component in a known state. Treat stories as executable fixtures, not only as documentation. A basic story can act as a render smoke test: if it throws while rendering, the check fails. A collection of focused stories can then supply inputs for interaction, accessibility, and visual checks.
Keep examples understandable and stable. Give stories representative names, provide deterministic data, and mock network or time-dependent inputs when needed. The Storybook testing guide describes component tests as browser-rendered tests that simulate user interactions while focusing on a UI unit; see the [official component testing documentation](https://storybook.js.org/docs/writing-tests/component-testing).
3. Test behavior at the component boundary
For stateful components, test what a user can do and what the component promises in response. In Storybook, a story’s play function can set up state, interact with the rendered component, and assert outcomes. Mock dependencies or network responses when the component needs them so the case stays isolated and repeatable.
Useful behavior cases include:
- Typing into a field updates its value and validation feedback appears at the intended time.
- Opening a dialog exposes its contents; closing it works through the close control and expected keyboard action.
- Submitting a form calls the expected callback and handles success or error states.
- Selecting an item updates the visible selection and the accessible state.
- Keyboard navigation reaches controls in the expected order and does not trap focus unexpectedly.
Assert observable outcomes and the documented component contract. Avoid assertions tied to internal implementation details, such as private state variable names, unless that detail itself is part of the contract.
4. Add visual regression checks
Visual checks compare a rendered story with an accepted baseline image. They are useful for catching unintended changes to spacing, typography, colors, borders, alignment, and responsive layout. Review changes when a component, style, or token changes; an intentional redesign should update the baseline after review.
Visual testing is complementary to behavior testing. A screenshot cannot establish that a control responds to a keyboard, that validation data is correct, or that a screen reader receives useful semantics. Storybook documents visual testing for stories and cross-browser workflows, including its hosted service Chromatic, in its [visual testing guide](https://storybook.js.org/docs/writing-tests/visual-testing).
Choose visual coverage based on risk: prioritize commonly used components and variants, recently changed tokens, and layouts where a small difference could hide or obstruct content. Include browser coverage that matches the design system’s support policy. Pixel differences can result from fonts, browser versions, or rendering environments, so keep those inputs controlled and inspect diffs instead of accepting every change automatically.
5. Check accessibility with automation and review
Run automated accessibility checks against rendered stories. Storybook’s accessibility addon audits the DOM for common issues using rules informed by WCAG and other accepted practices. Its [accessibility testing documentation](https://storybook.js.org/docs/writing-tests/accessibility-testing) says axe-core can detect up to 57% of WCAG issues; this is an attributed detection estimate, not a measure of all accessibility barriers.
Review reported violations and also manually inspect behavior that automated rules cannot settle. Check keyboard operation, accessible names and semantics, contrast, zoom, and reduced-motion behavior where relevant. An “incomplete” automated result is a prompt for human inspection. A clean automated report is useful evidence, but it does not prove that a component works for every person or assistive technology.
6. Verify design-system promises beyond component code
Some system guarantees need checks outside an isolated story. Adapt these examples to the support matrix and conventions the design system actually publishes:
- Responsive behavior: inspect each supported breakpoint and confirm content remains usable.
- Zoom: at 400% browser zoom, check that content remains available without overlap or forced horizontal scrolling where the system promises reflow.
- Localization: change the language and confirm default text updates; inspect longer translated strings for clipping.
- Design-to-code parity: compare code props and options with the corresponding Figma component.
- Token use: inspect whether styles use existing design tokens in both code and design files.
These are concrete examples from the [CMS.gov Design System component maturity guidance](https://designsystem.digital.gov/documentation/maturity-model/). Its breakpoints and maturity criteria are specific to that system, so treat them as prompts and tailor the exact acceptance criteria to your own promises.
7. Use end-to-end tests for integration risks
Keep most component contract checks isolated for quick feedback. Add E2E tests when a failure depends on the product stack or a realistic workflow across multiple components—for example, completing checkout through a form, dialog, validation message, and confirmation page.
Storybook stories can be reused in browser automation tools such as Playwright or Cypress when that helps exercise a scenario. Keep the distinction clear: an isolated story catches component-level failures quickly; a full application test checks that integrated behavior works in its real context. Avoid duplicating every story as an E2E test, since broad E2E suites require more setup and maintenance.
8. Run the right checks in CI
Run relevant checks on pull requests that change shared components, styles, tokens, or stories. A practical sequence is:
- Build or render changed stories and fail on rendering errors.
- Run interaction tests and automated accessibility checks for selected stories.
- Capture visual snapshots and publish differences for review.
- Run targeted E2E tests for affected integration paths.
- Use coverage reports to locate important untested branches or interactions, then decide whether those gaps represent risk.
Make failures easy to diagnose: report the story or scenario, preserve useful browser output, and show visual diffs where available. Keep fixtures deterministic so unrelated network failures or changing content do not obscure a real regression. Storybook cautions against treating 100% coverage as a universal target; use coverage to identify meaningful gaps rather than as the definition of quality.
9. Choose methods by the failures they catch
| Method | Best at catching | Environment | Review needed? |
|---|---|---|---|
| Story render smoke check | Render exceptions and missing dependencies | Isolated story | Usually only when it fails |
| Interaction test | State and user-action regressions | Browser story with controlled data | Review assertions and failures |
| Visual regression | Unintended appearance changes | Browser snapshot and baseline | Yes, for changed images |
| Automated accessibility audit | Common DOM and rule-based accessibility violations | Rendered DOM | Yes, especially incomplete results |
| Manual accessibility review | Usability and assistive-technology issues automation cannot establish | Human review with relevant tools | Yes |
| End-to-end test | Integration failures across components and application services | Running product or realistic app environment | Review test failures and scenario coverage |
Consider feedback speed, browser coverage, maintenance cost, and whether a result needs human interpretation. A balanced portfolio uses complementary methods rather than expecting any one tool to prove overall quality. Storybook also notes that component tests can become expensive to maintain when applied indiscriminately; see its [testing overview](https://storybook.js.org/docs/writing-tests).
Or skip the browser setup
If you need a screenshot of a rendered story or page for visual review, [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server from Yorker Media. A single request can capture a page as PNG, JPEG, WebP, or PDF. Use it for rendered appearance evidence; keep interaction and accessibility checks in your test suite, since an image does not verify behavior or accessibility.
For a basic screenshot, adapt the target URL to a publicly accessible story URL. See the [ScreenshotNeo API docs](https://screenshotneo.com/docs/) for options and authentication details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo can accept cookie or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, with no card required.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| A story fails before an interaction runs | The component throws during render, a required provider is missing, or fixture data is incomplete. | Run the story by itself, inspect browser output, and add the provider or deterministic fixture that the component contract requires. |
| An interaction test passes locally but fails in CI | Timing depends on animation, network activity, or asynchronous state; the assertion may run too early. | Mock external responses, wait for the user-visible result, and avoid fixed sleeps where a state-based wait is possible. |
| Visual snapshots change across runs | Fonts, browser versions, animations, timestamps, or remote content differ between environments. | Pin the rendering environment where possible, disable or stabilize motion, and use fixed data and fonts. Review the diff before updating a baseline. |
| A visual diff appears after a token update | The token changed multiple components or a snapshot environment shifted. | Inspect representative stories and affected breakpoints; accept only changes that match the intended system update. |
| Accessibility audit reports “incomplete” | The rule requires context or human judgment. | Manually verify the flagged case, including keyboard use and semantics, and record the review outcome. |
| Coverage is high but users still find defects | Coverage measures executed code, not whether important states, expectations, or assistive-technology needs were checked. | Review the component contract and risk gaps; add targeted behavior, visual, manual accessibility, or integration checks. |
| E2E suite is slow or brittle | Too many isolated component details are being tested through a full application, or tests depend on unstable services. | Move component-level assertions to isolated stories and retain E2E coverage for workflows that require integration. |
Performance, reliability, and cost
Fast feedback comes from keeping the common path focused: run render, interaction, and accessibility checks for relevant stories, then use visual review and a smaller set of integration scenarios for high-risk changes. Broader browser matrices and full application journeys add confidence but also consume CI time and need more environment maintenance.
For reliability, make network data, dates, fonts, and animations predictable. Separate failures in the test environment from regressions in the component, and make snapshots reviewable. A passing automated check should be interpreted according to what it actually observes: rendered pixels, a DOM rule, an interaction assertion, or an integrated workflow.
There is no universal cost or performance figure for this testing portfolio. Estimate from your own component count, browser support matrix, change frequency, CI capacity, and baseline review effort. Avoid paying the maintenance cost of duplicating every low-risk component state in every testing layer.
FAQ
Do all components need visual snapshots?
No. Prioritize shared, high-impact components, key variants, and areas with meaningful visual risk. A focused snapshot set is easier to review and maintain.
Can an accessibility addon certify a design system as accessible?
No. It can find common rule-based issues in rendered output, but manual and assistive-technology review is still needed for context and usability.
Should design systems have end-to-end tests?
Yes, for selected risks that cross component and application boundaries. Keep most component contract checks isolated for faster, more focused feedback.
Is 100% test coverage the goal?
Not by itself. Coverage can reveal overlooked paths, but risk-based representative scenarios are a better guide to whether the system’s promises are being checked.


