How to test a web app’s modal dialogs with screenshot comparisons
Test modal dialogs with Playwright screenshot comparisons, accessible interaction checks, stable baselines, and practical fixes for flaky visual tests.
Use Playwright Test to open a modal through its normal trigger, check its visibility and accessible name, verify key keyboard and focus behavior, then compare a screenshot of either the dialog or the full viewport with a reviewed baseline. Element screenshots protect the panel; viewport screenshots also cover the backdrop, placement, clipping, and the page behind it. Screenshot comparisons do not prove that a dialog behaves accessibly, so pair them with semantic and keyboard assertions.
This guide uses TypeScript with Playwright Test. It covers baseline setup, complete examples, capture options, stability, troubleshooting, and an API alternative for capturing pages outside a browser test. Playwright APIs can change; consult the visual comparisons guide and page assertion reference when adapting the examples.
1. Decide what the test must protect
Write down the user-visible states and defects that matter before adding screenshots. A focused set of meaningful states is easier to review than many nearly identical captures.
- Open state: panel position, size, typography, controls, backdrop, and whether the underlying page is obscured as intended.
- Validation or error state: error copy, field styling, spacing, and any change in dialog height.
- Confirmation state: success content and available next actions.
- Responsive and content variations: include a narrow viewport or long content when they are part of the design contract.
- Interaction behavior: focus enters the dialog, remains contained while tabbing, and returns to the opener or another logical location after closing.
Use the smallest set that represents distinct requirements. If the test is only about panel typography, capture the dialog element. If it must catch a broken backdrop, incorrect centering, clipping, or layering against the page, capture the viewport.
2. Install Playwright Test and create a baseline
In a Node.js project, install the test runner and its browser binaries:
npm init playwright@latest
npx playwright install
Choose TypeScript when prompted, or add a test file such as tests/settings-dialog.spec.ts to an existing Playwright Test project. The sample assumes the application is available at http://localhost:3000 and has a button named “Open settings” that opens a dialog named “Settings.” Replace these with the real route and accessible names.
import { test, expect } from '@playwright/test';
test('settings dialog matches its expected appearance', async ({ page }) => {
await page.goto('http://localhost:3000/settings');
const opener = page.getByRole('button', { name: 'Open settings' });
await opener.click();
const dialog = page.getByRole('dialog', { name: 'Settings' });
await expect(dialog).toBeVisible();
await expect(dialog).toHaveScreenshot('settings-dialog.png');
});
Run it with npx playwright test tests/settings-dialog.spec.ts. On the first run, Playwright creates the expected screenshot. Review and commit that baseline with the test. Later runs compare the captured image with the expected image and report differences. Update a baseline only after reviewing the actual and diff images and deciding the visual change is intentional.
For teams, make the browser, operating system, fonts, viewport, and device scale consistent between baseline creation and comparison. Playwright notes that rendering can vary across environments and recommends generating and comparing screenshots in the same environment. See its visual comparisons guide.
3. Choose dialog-only or viewport capture
Capture only the dialog
Locator screenshot assertions focus on the selected element, which is useful when the modal panel itself is the contract under test. Keep the accessible locator narrow and specific.
const dialog = page.getByRole('dialog', { name: 'Settings' });
await expect(dialog).toBeVisible();
await expect(dialog).toHaveScreenshot('settings-dialog.png');
Capture the full viewport
Use a page screenshot when the relationship between the modal and the page matters: backdrop opacity, overlay coverage, stacking order, alignment, or content clipped at the viewport edge.
await expect(page).toHaveScreenshot('settings-dialog-viewport.png');
A page screenshot can include unrelated content that changes for reasons outside the dialog. Control or mask only irrelevant variation; do not mask the dialog or meaningful content whose appearance the test is meant to protect.
4. Pair visual checks with semantics and keyboard behavior
The visual assertion catches appearance changes. Explicit assertions describe the dialog’s name, controls, messages, and interaction expectations. The WAI-ARIA Authoring Practices Guide says focus moves inside when a dialog opens, stays within while users tab, Escape closes it, and focus ordinarily returns to the invoking element. Follow the behavior your application promises, including documented exceptions. See the WAI-ARIA modal dialog pattern and W3C’s native dialog technique.
import { test, expect } from '@playwright/test';
test('settings dialog looks and behaves as expected', async ({ page }) => {
await page.goto('http://localhost:3000/settings');
const opener = page.getByRole('button', { name: 'Open settings' });
await opener.click();
const dialog = page.getByRole('dialog', { name: 'Settings' });
const save = dialog.getByRole('button', { name: 'Save' });
const close = dialog.getByRole('button', { name: 'Close' });
await expect(dialog).toBeVisible();
await expect(dialog).toHaveScreenshot('settings-dialog.png');
await expect(save).toBeVisible();
await expect(close).toBeVisible();
// Adapt this expectation if the dialog intentionally focuses another element.
await expect(save).toBeFocused();
await page.keyboard.press('Shift+Tab');
await expect(close).toBeFocused();
await page.keyboard.press('Escape');
await expect(dialog).toBeHidden();
await expect(opener).toBeFocused();
});
Focus order depends on the application’s markup and intended initial focus. Adjust the assertions to match the documented interaction contract rather than assuming the Save button is always first. Add further Tab and Shift+Tab steps when needed to establish that focus does not escape the modal. If the product does not support Escape to close, test its documented close path instead.
Use focused role, name, visibility, and text assertions for requirements such as an error message or confirmation button. Playwright also supports accessibility snapshot assertions for comparing accessibility-tree structure. Those can complement visual checks, but focused assertions are often clearer than a broad tree snapshot when only a few semantics matter.
5. Stabilize screenshot comparisons
Playwright’s screenshot assertions wait for two consecutive screenshots to match before comparing the final capture with the expected image. This helps with short-lived rendering changes, but it cannot make unstable application data deterministic. See the assertion options and locator assertion options.
| Source of variation | Practical control |
|---|---|
| Different OS, browser, fonts, or browser settings | Generate and compare baselines in the same CI image and browser configuration. |
| Viewport or device scale changes | Set a fixed viewport and device scale in the Playwright project configuration. |
| Animations and transitions | Disable or finish animations for the test, or wait for the intended settled state. |
| Time, randomized data, remote content, or user-specific content | Use deterministic fixtures, fixed test data, and controlled network responses where practical. |
| Dynamic regions irrelevant to the requirement | Mask or hide only those regions with a short comment explaining why their pixels do not matter. |
| Small rendering noise | Use a narrow, explained threshold only when needed; review diffs instead of loosening tolerance until meaningful changes pass. |
Use a fixed viewport in configuration, for example:
import { defineConfig } from '@playwright/test';
export default defineConfig({
use: {
baseURL: 'http://localhost:3000',
viewport: { width: 1280, height: 800 },
deviceScaleFactor: 1,
},
});
Project configuration options and screenshot assertion options are documented in the test use options reference and the screenshot assertion references above. Keep any threshold local to the assertion and explain what rendering noise it accommodates.
6. Test distinct dialog states without multiplying noise
For a validation state, drive the dialog through its real controls and assert both the expected message and the screenshot. This illustrates the shape; adapt selectors and expected copy to the application.
test('shows the required-field error in the settings dialog', async ({ page }) => {
await page.goto('/settings');
await page.getByRole('button', { name: 'Open settings' }).click();
const dialog = page.getByRole('dialog', { name: 'Settings' });
await dialog.getByRole('button', { name: 'Save' }).click();
await expect(dialog.getByText('Choose a name')).toBeVisible();
await expect(dialog).toHaveScreenshot('settings-dialog-validation.png');
});
For responsive behavior, define a viewport in the relevant project or set it before navigation with page.setViewportSize(), then open the modal and assert the layout that matters. Avoid taking several captures at near-identical widths unless they represent separate supported breakpoints or known failure risks.
7. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| First run creates a baseline unexpectedly | No expected screenshot exists yet. | Review the generated image, confirm the test reached the intended state, then commit the baseline. |
| Screenshot differs on every machine | OS, browser build, fonts, device scale, or rendering environment differs. | Run baseline generation and comparison in the same pinned environment and configuration. |
| Screenshot is captured before the dialog settles | Animations, asynchronous content, or late layout shifts remain active. | Wait for a user-visible condition or stable test state; control animation and data variation. |
| Dialog locator times out | Trigger failed, accessible role/name differs, or the dialog is in another frame. | Check the trigger, inspect the rendered role and accessible name, and scope the locator to the correct frame if applicable. |
| Viewport diff includes unrelated page content | The chosen boundary is broader than the requirement. | Use a locator screenshot if only the panel matters, or stabilize the surrounding page when overlay behavior is part of the requirement. |
| Keyboard assertion fails while screenshot passes | Appearance is intact but focus behavior does not match the expected interaction. | Inspect initial focus, tab order, close behavior, and focus restoration separately; fix the interaction or correct an unsupported assumption in the test. |
| Baseline update hides a real regression | Expected images were refreshed without reviewing the diff. | Review actual and diff artifacts, tie the change to an intentional design update, and update only affected baselines. |
8. Performance, reliability, and maintenance
Every browser test has page-load, interaction, rendering, and image-comparison work. Keep the suite efficient by testing representative modal states, reusing normal test setup, and avoiding redundant screenshots that assert the same visual contract. A dialog-only capture can reduce unrelated page variation; a viewport capture is necessary when backdrop or positioning is part of the requirement.
Reliability depends on controlling the environment and the content under test. Store baselines with the test code, review diffs in code review, and regenerate them in the same environment used for comparison. Do not treat a passing screenshot as proof of accessibility, correct focus behavior, or successful interaction: keep the semantic and keyboard assertions.
Playwright is a browser-testing workflow that stores expected screenshots alongside tests. For a separate capture workflow, ScreenshotNeo is a website screenshot API and MCP server. Its API can capture a page as an image or PDF, but an API capture does not replace Playwright’s stored-baseline comparison and browser interaction assertions.
Or skip the browser setup
For a page capture outside your test runner, ScreenshotNeo takes a URL in one GET request and returns an image or PDF. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
Replace the example URL with a page you are authorized to capture. The Node.js example uses Bun’s file writer; in Node.js, save the returned bytes with writeFile from node:fs/promises. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Learn more at ScreenshotNeo and the documentation. Sign up for 1,000 free screenshots a month with no card.
Frequently asked questions
Should I compare the modal element or the whole page?
Capture the element to protect the panel itself. Capture the viewport when backdrop, placement, clipping, layering, or the background page is part of the visual requirement.
Does a screenshot comparison test accessibility?
No. Add assertions for accessible role and name, controls, focus movement, focus containment, and the supported close and return-focus behavior.
When should I update the expected screenshot?
After reviewing the actual and diff images and confirming the visual change is intentional. Keep the updated baseline with the code change that explains it.
Can I use accessibility snapshots instead of visual comparisons?
They answer a different question by comparing accessibility-tree structure. They complement screenshots; neither one alone covers appearance, semantics, and keyboard behavior.


