ScreenshotNeo

BlogHow-to

How to Test Mermaid Diagrams with Visual Regression Testing

Catch invalid Mermaid syntax and unintended visual changes with a repeatable workflow using Mermaid, Playwright, and CI.

By the ScreenshotNeo team4 October 20267 min read

Test Mermaid diagrams in two separate steps: parse each definition to catch syntax errors, then compare a screenshot of the rendered diagram with an approved baseline. Use Mermaid CLI when the generated artifact is what matters; use a browser screenshot test when the production page, CSS, fonts, theme, and viewport are part of the result.

1. Decide what you are testing

A visual regression test should represent the output you need to keep stable. Choose one of these targets:

  • Generated artifact: Test an SVG, PNG, or PDF produced from a Mermaid source file. This is a focused way to check diagram rendering.
  • Rendered application page: Test the page that initializes Mermaid in your site. This also exercises browser rendering, page styles, theme, fonts, and layout around the diagram.

If readers encounter the diagram inside your documentation or app, prefer a browser test for the final visual check. A separately generated image cannot catch every integration issue. Mermaid documents both browser rendering and CLI rendering; see the Mermaid usage documentation and Mermaid CLI README.

2. Validate Mermaid syntax first

Syntax validation is fast and answers a different question from screenshot comparison. Mermaid’s parse API checks whether a definition is valid and returns its diagram type; invalid input throws unless errors are suppressed. It does not prove that the resulting diagram looks right.

In a project that already imports Mermaid, add a focused test such as:

import mermaid from 'mermaid';

const definition = `
flowchart LR
  A[Source] --> B[Rendered diagram]
`;

try {
  const diagramType = await mermaid.parse(definition);
  console.log(`Valid Mermaid definition: ${diagramType}`);
} catch (error) {
  console.error('Invalid Mermaid definition:', error);
  process.exitCode = 1;
}

Adjust module loading to the test runner and Mermaid version already used by your project. Keep parser errors visible in CI: a screenshot failure is harder to diagnose if invalid syntax is first reported as an empty or missing diagram.

3. Render a source file with Mermaid CLI

For a standalone diagram file, Mermaid CLI can emit SVG, PNG, or PDF. Install and pin the CLI version through your project’s chosen package-management workflow, then render with the documented command pattern:

mmdc -i architecture.mmd -o architecture.svg

For a raster artifact, change the output extension to .png; for a PDF, use .pdf. CLI options include theme and background configuration. Consult the CLI README for the complete options supported by the version you pin. Mermaid CLI can also process Markdown containing Mermaid blocks and produce Markdown that references generated SVG files.

Pin Mermaid and renderer configuration so an upgrade is a deliberate change. A renderer upgrade can alter output even when the diagram source stays the same; review generated changes instead of silently refreshing baselines.

4. Compare the diagram in its real browser page

Use Playwright’s screenshot assertion when the browser presentation is the regression target. The following test assumes the app is available at the configured base URL and renders an SVG inside .mermaid; adapt the route and selector to your integration.

import { test, expect } from '@playwright/test';

test('architecture diagram stays visually stable', async ({ page }) => {
  await page.goto('/docs/architecture');

  const diagram = page.locator('.mermaid svg');
  await expect(diagram).toBeVisible();
  await expect(diagram).toHaveScreenshot('architecture-diagram.png');
});

Install and configure Playwright Test according to its snapshot documentation. The first run creates a missing baseline. Inspect that image, then commit it as the expected result. In subsequent runs, Playwright compares the captured image with the baseline. To accept an intentional visual change, review the diff and regenerate snapshots with npx playwright test --update-snapshots.

Wait for the rendered diagram to be ready

toBeVisible() checks visibility, but a particular app may need additional readiness conditions. Wait for the Mermaid SVG or another app-specific signal that confirms rendering is complete. If the diagram loads fonts or data asynchronously, wait for those resources or for a stable page state before capturing. Avoid relying on a fixed delay unless the app has a specific timing requirement; arbitrary waits can be slow and still flaky.

Capture only the relevant area

An element screenshot keeps unrelated page content out of the comparison. Use a page screenshot instead if layout around the diagram is part of the contract. In either case, fix the viewport and keep dynamic content out of the capture. Playwright supports screenshot assertions for locators and pages, along with comparison options such as maxDiffPixels.

5. Make the comparison reproducible

Browser screenshots can vary with operating system, browser version, settings, hardware, power source, and headless mode. Generate baselines and compare them in the same environment where practical. Use a fixed browser project, viewport, installed fonts, and headless configuration in CI.

Choose test cases based on the modes your product supports. There is no universal matrix; add cases when the corresponding difference is visible to users.

Dimension Include a case when
Theme Users can view diagrams in more than one theme or theme affects diagram colors.
Viewport Resizing can change wrapping, fit, clipping, or legibility.
Browser or OS Cross-browser or cross-platform presentation is a supported requirement.
Fonts Your product supplies fonts or font loading can change text metrics.

Keep the matrix focused on supported user experiences: every extra environment creates more baselines to maintain and review.

6. Review visual changes and set tolerances carefully

When a screenshot assertion fails, inspect the actual image, expected image, and diff. Decide whether the difference is a regression or an intentional design change before updating the baseline. Treat snapshot updates as reviewed code changes.

Playwright uses pixel comparison and supports options such as maxDiffPixels. A tolerance can absorb small rendering noise, but a large allowance can hide meaningful changes. Set it based on observed, reviewed behavior and keep it consistent with the visual importance of the diagram.

7. Run the checks in CI

  1. Run Mermaid parsing checks to catch invalid definitions.
  2. Render standalone artifacts with the pinned Mermaid CLI when those artifacts are shipped or generated.
  3. Run the browser screenshot assertion against the real route for diagrams whose appearance depends on the app.
  4. Upload or otherwise make the expected, actual, and diff images available in the normal CI review workflow.
  5. Require review of baseline updates and keep renderer, browser, fonts, and viewport settings consistent.

Do not treat a passing parse test as a visual check, or a passing screenshot test as proof that every Mermaid source file was validated. Keep each assertion matched to the property it verifies.

8. Choose the right testing route

Approach Best for What it does not cover alone
Mermaid parse Fast syntax validation. Rendering, layout, and visual correctness.
Mermaid CLI Generated SVG, PNG, PDF, or Markdown diagram artifacts. The full production browser integration unless that route is itself what you render.
Playwright screenshot assertion What a browser page or diagram element displays. Other environments or modes not included in the configured cases.

Mermaid’s project overview describes using Argos for pull request visual regression testing and Applitools in its release process; the CLI README also references Percy. These are examples, not requirements. Confirm current capabilities and terms with any service before choosing it.

9. Troubleshooting

Symptom Likely cause Fix
The parse check fails. Invalid Mermaid syntax or a definition that differs from the expected input. Print the exact definition, inspect the parser error, and fix the source before comparing screenshots.
The SVG locator never appears. The route, selector, Mermaid initialization, or render readiness assumption is wrong. Check the page in the test browser, match the selector to the actual markup, and wait for an app-specific render signal.
A baseline is missing. This is the first run or the snapshot name/project does not match an existing baseline. Verify the test name and browser project, then generate and review the initial baseline intentionally.
Snapshots differ on every run or machine. Environment, font, browser, viewport, or dynamic page content varies. Use a consistent runner and configuration, load the intended fonts, fix viewport size, and exclude unrelated dynamic regions.
The diff is noisy after a renderer update. Mermaid or browser rendering behavior changed. Review the change as a dependency upgrade, inspect representative diagrams, and update baselines only after approval.
The screenshot passes despite a visible issue. The capture targets the wrong element, readiness is premature, or the tolerance is too permissive. Capture the user-visible region, wait for final rendering, and tighten the threshold based on reviewed output.

10. Performance, reliability, and maintenance

Parsing definitions is generally the focused check to run early; browser screenshots require page navigation and rendering, and therefore add more work to a CI run. Keep browser coverage targeted to representative diagrams and supported presentation modes. Use CLI rendering for generated artifacts when that is the artifact you need to protect.

Reliability depends on repeatable inputs: fixed Mermaid and browser versions, consistent fonts and viewport, deterministic diagram data, and explicit readiness. Snapshot review also matters: automated comparison finds differences, but a human still decides whether a change is intended. No universal performance benchmark or threshold is prescribed by the cited tools.

Or skip the browser setup

If your goal is to capture a page containing a Mermaid diagram without setting up a browser capture pipeline, ScreenshotNeo provides a website screenshot API and MCP server. Its screenshot API can capture a URL in one GET request; it does not replace the syntax check or a checked-in visual baseline when regression comparison is required.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/docs/architecture -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; responses identify the page verdict and billing status. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does Mermaid parsing catch layout regressions?

No. Parsing validates Mermaid syntax. A rendered artifact or browser screenshot comparison is needed to check appearance.

Should I snapshot the SVG or the whole page?

Snapshot the SVG or its container when diagram appearance is the target. Snapshot the page when surrounding layout and integration are part of the user-visible contract.

Should every browser and theme have a baseline?

Only include modes your product supports and where output differences matter. Add cases deliberately because each one adds baseline maintenance.