How to Test a Web App’s Modal Dialogs with Screenshot Regression Tests
Build reliable modal screenshot tests with Playwright: open dialogs through their normal controls, verify focus and keyboard behavior, and review stable visual baselines.
Use Playwright Test to open a modal through its normal user-facing control, verify its accessible name and keyboard behavior, then compare the meaningful open state with a reviewed screenshot baseline. A screenshot catches visual changes; it does not prove that the dialog is named correctly, traps focus, or can be dismissed. Test those behaviors explicitly alongside the visual assertion.
This guide uses TypeScript and Playwright Test. The same workflow applies to other web test runners: establish a deterministic state, capture a reference, compare later renders in a consistent environment, and review every proposed baseline change.
1. Set up Playwright and a stable test environment
Install Playwright Test and its browser binaries using the official Playwright installation guide. Keep the Playwright version, browser project, viewport, locale, color scheme, and test data stable between baseline creation and normal runs. Rendering can vary by operating system, browser version, settings, hardware, power source, and headless mode, so create and verify snapshots in the same environment.
A typical project configuration can pin the browser project and viewport used for this test:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
use: {
browserName: 'chromium',
viewport: { width: 1280, height: 800 },
locale: 'en-US',
colorScheme: 'light',
},
});
Use the browser and platform configuration that matches the environment where the snapshots will be reviewed and maintained. If your suite intentionally runs across several browsers or platforms, treat their baselines as distinct references rather than assuming a single image is portable.
2. Write a modal test with behavior and visual checkpoints
The example assumes a settings page with a button named “Open settings,” a dialog named “Settings,” and a close button. Adapt accessible names and routes to your application. The screenshot is taken after the dialog opens and its content is visible; the keyboard checks then exercise dismissal and focus restoration.
import { test, expect } from '@playwright/test';
test('settings dialog appearance and behavior', async ({ page }) => {
await page.goto('/settings');
const opener = page.getByRole('button', { name: 'Open settings' });
await opener.click();
const dialog = page.getByRole('dialog', { name: 'Settings' });
await expect(dialog).toBeVisible();
await expect(dialog).toHaveAttribute('aria-modal', 'true');
await expect(dialog).toContainText('Manage your preferences');
// Make the focus state intentional if it affects the dialog's appearance.
const firstControl = dialog.getByRole('textbox', { name: 'Display name' });
await expect(firstControl).toBeFocused();
await expect(dialog).toHaveScreenshot('settings-dialog.png', {
animations: 'disabled',
});
await page.keyboard.press('Escape');
await expect(dialog).toBeHidden();
await expect(opener).toBeFocused();
});
The initial focus target depends on the dialog’s purpose and contents. Assert the target your interface intends. If it does not focus a textbox, select an appropriate control or the dialog container and test that instead. Do not add a test expectation that contradicts the design simply to match this sample.
Playwright’s toHaveScreenshot() creates a reference on first use and compares later captures against it. Screenshot assertions wait until two consecutive screenshots match before comparing, which helps avoid capturing an intermediate render. Review the created image and commit it with the test code only after confirming it represents the intended state. See the official visual comparisons guide.
3. Verify focus, keyboard use, and dismissal
A modal’s keyboard and focus behavior is part of the test contract. The W3C WAI-ARIA Authoring Practices Guide says focus moves inside when a dialog opens, Tab and Shift+Tab stay within the dialog’s tab sequence and wrap at its ends, Escape closes it, and focus ordinarily returns to the invoking control. See the WAI-ARIA APG modal dialog pattern.
Test the first and last tabbable controls directly. This example assumes the dialog’s first and last controls are known and that normal tab order is intended. Add or adapt assertions if your dialog contains disabled controls, custom focus handling, or a workflow-specific focus destination.
test('settings dialog keeps keyboard focus inside', async ({ page }) => {
await page.goto('/settings');
const opener = page.getByRole('button', { name: 'Open settings' });
await opener.click();
const dialog = page.getByRole('dialog', { name: 'Settings' });
const first = dialog.getByRole('textbox', { name: 'Display name' });
const last = dialog.getByRole('button', { name: 'Save settings' });
await expect(dialog).toBeVisible();
await expect(first).toBeFocused();
await page.keyboard.press('Shift+Tab');
await expect(last).toBeFocused();
await page.keyboard.press('Tab');
await expect(first).toBeFocused();
await page.getByRole('button', { name: 'Close settings' }).click();
await expect(dialog).toBeHidden();
await expect(opener).toBeFocused();
});
Also verify the visible close control. Where background content must be inert, test that a meaningful background action cannot occur while the modal is open. A pixel snapshot cannot establish either of those behaviors. If your product flow deliberately sends focus somewhere other than the opener after dismissal, assert the documented destination.
4. Choose the right screenshot scope and state
Use a dialog locator screenshot when the visual contract is the dialog’s internal content and styling. Use a page screenshot when the backdrop, alignment, position, or relationship to surrounding content matters. A dialog-only capture can reduce unrelated page changes, but it may omit the overlay and placement that users see.
Snapshot meaningful states separately when they represent distinct UI outcomes: the initial open state, a validation error, or a confirmation state. Keep state setup explicit and deterministic so a baseline has a clear purpose. If focus styling is part of the appearance, put focus in the same place before each capture.
5. Stabilize captures without hiding defects
Prefer fixed test data, a known viewport, and controlled application state over broad masking. Disable animations for the assertion when motion creates transient frames. Playwright also supports screenshot controls including caret handling, clipping, full-page capture, masks, CSS or device scale, style injection, and acceptable pixel or color differences; use the screenshot assertion API reference for the supported options and current defaults.
For genuinely irrelevant dynamic content, mask only the smallest locator that contains it, or inject a narrowly scoped stylesheet. Keep assertions for meaningful text and state even when a visual region must be masked. A mask removes visual coverage in its pixels; it should not conceal content whose correctness matters to users.
await expect(dialog).toHaveScreenshot('settings-dialog.png', {
animations: 'disabled',
caret: 'hide',
mask: [page.getByTestId('volatile-clock')],
// Set maxDiffPixelRatio or color tolerances only when the
// team's reviewed visual policy calls for them.
});
The example intentionally leaves difference thresholds unset. A stricter comparison catches small changes but is more sensitive to rendering noise; a more tolerant comparison can miss defects. Tune thresholds against a stable environment and inspect actual diffs rather than using tolerance to make unexplained changes disappear.
6. Create, review, and update baselines
- Run the focused test in the pinned project and environment to create the initial reference screenshot.
- Open the generated image and check that the dialog is in the intended state, including focus styling, backdrop, and content.
- Commit the accepted snapshot with the test so reviewers can inspect it alongside code changes.
- On a later visual difference, inspect the diff and determine whether the change is intentional or a regression.
- Update snapshots only after review. Playwright documents
--update-snapshotsfor refreshing references; run the focused test with that option after deciding the new appearance is correct.
A passing screenshot assertion means the current capture is within the configured comparison policy for that baseline. It does not mean the modal meets accessibility requirements or that every relevant state was tested.
7. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| The snapshot differs on every machine | Baselines and runs use different operating systems, browser versions, headless modes, fonts, or rendering settings. | Generate and run snapshots in the same pinned environment; keep browser and project configuration consistent. |
| The screenshot catches a partial transition | An animation, delayed content, or asynchronous data is still changing. | Wait for the expected dialog content or state, use deterministic fixtures, and disable animations for the screenshot assertion where appropriate. |
| The dialog locator is not found | The dialog lacks the expected role or accessible name, the opener did not activate it, or the test runs before the page is ready. | Inspect the rendered accessibility tree, fix the dialog’s semantics/name, and use role-based locators that match the actual interface. |
| Focus assertion fails after open | The application does not move focus into the modal, or the test expects the wrong initial control. | Choose and implement an intentional initial focus target, then assert that target. Avoid relying on incidental browser focus. |
| Escape closes the wrong layer or nothing closes | The app has nested overlays, a custom dismissal rule, or a bug in its keyboard handler. | Define the intended dismissal behavior, test the active topmost dialog, and add an explicit close-control test. |
| Snapshot changes after a dependency update | The browser engine, fonts, CSS, or component rendering changed. | Review the image diff as a product change. Update the baseline only when the new appearance is intended. |
| Masking makes a test pass but a defect slips through | The mask covers too much or hides meaningful content. | Narrow the masked locator and assert important text/state separately; prefer deterministic data where possible. |
| A screenshot passes despite broken keyboard behavior | Visual comparison checks pixels, not operability. | Add role/name, focus, Tab and Shift+Tab, Escape, close-control, and focus-return assertions. |
8. Performance, reliability, and cost
Screenshot assertions render and compare images, so keep the capture area and number of states aligned with the visual risks you need to cover. A locator capture generally avoids comparing unrelated page pixels; a full-page or page capture is appropriate when the backdrop and placement matter. Do not trade away meaningful coverage merely to reduce runtime.
Reliability comes mainly from repeatable inputs and a stable rendering environment: deterministic data, fixed viewport and browser project, intentional focus state, and waiting for the actual UI state. Repeatedly increasing thresholds or masking large areas can reduce useful signal without fixing the source of variation.
For teams running browser tests, budget for the browser execution and CI resources used by the suite. Keep baselines under version control and review updates as code changes; the sources cited here do not establish a universal runtime or dollar cost, so measure your own suite and environment.
Or skip the browser setup
If you need a rendered page image without managing a browser capture service, ScreenshotNeo is a website screenshot API and MCP server. It takes a URL in one GET request and returns a PNG, JPEG, WebP, or PDF. For visual regression work, it can capture a public test page, but your Playwright assertions should still verify modal focus and keyboard behavior.
See the ScreenshotNeo API documentation. This cURL request saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo accepts the cookie or consent banner as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Those clean-page captures can help when you want a screenshot endpoint without maintaining browser capture code, while local regression assertions remain responsible for comparing reviewed modal states.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Should a modal screenshot test replace accessibility tests?
No. A screenshot records appearance. Test the dialog role and name, focus placement, keyboard navigation, dismissal, and background interaction directly.
Should I snapshot the whole page or only the dialog?
Snapshot the dialog for its internal visual contract. Include the page when backdrop, placement, or surrounding context is part of the behavior you need to preserve.
When should a snapshot baseline be updated?
After reviewing the diff and confirming that the new appearance is intended. Do not accept a changed baseline just to clear a failing run.


