How to Compare Screenshots When CSS Layout Shifts After Page Load
Wait for the intended page state, stabilize the browser, then compare a repeatable screenshot with a reviewed baseline. Learn how to diagnose layout shifts in Playwright.
To compare screenshots when a page shifts after load, capture it only after it reaches the application state your test is meant to check, keep the browser and viewport consistent, and compare that stable capture with a reviewed baseline. In Playwright Test, expect(page).toHaveScreenshot() retries until two consecutive screenshots match, then compares the capture with its stored expectation. That helps with short-lived visual instability; it does not prove the page is in the right state. Assert the meaningful state first.
A pixel difference tells you that rendered pixels changed. It does not tell you whether the change is a defect or what caused it. Inspect the changed region and diagnose late content, fonts, images, third-party widgets, test data, or environment drift before changing a baseline or raising a threshold. Playwright’s visual comparison guide and screenshot assertion API describe the behavior and controls discussed here.
1. Install Playwright and create a visual test
The following runnable example uses Playwright Test with TypeScript. Install the test runner and Chromium, create tests/layout.spec.ts, then run the test once to create the initial baseline.
npm init playwright@latest
npx playwright install chromium
Choose TypeScript when the initializer asks for a language, or use an existing Playwright Test project. Add this test:
import { test, expect } from '@playwright/test';
test('results page matches its visual baseline', async ({ page }) => {
await page.setViewportSize({ width: 1440, height: 900 });
await page.goto('http://127.0.0.1:3000/search?q=playwright');
// Wait for the state the test actually intends to capture.
await expect(page.getByRole('heading', { name: 'Search results' })).toBeVisible();
await expect(page.getByTestId('results')).toHaveAttribute('data-loaded', 'true');
// Playwright waits for two consecutive matching screenshots before comparison.
await expect(page).toHaveScreenshot('search-results.png', {
fullPage: true,
animations: 'disabled',
});
});
Replace the example URL, heading, and test ID with selectors and state signals from your application. The app-specific assertions matter: screenshot stability cannot distinguish a fully rendered results page from a stable loading skeleton.
On the first run, Playwright writes a reference image under the test’s snapshot directory. Review and commit that image. Later runs compare against it. Baseline files include browser and platform naming because rendering can vary by operating system, browser version, hardware, settings, and headless mode. Keep baseline generation and comparison in the same environment when possible. Playwright documents baseline generation and updates.
2. Wait for a real ready state, not an arbitrary delay
Choose a condition that expresses why the page is ready for this test. Useful signals include a result heading becoming visible, a component exposing a loaded state, a known request completing, or a loading indicator disappearing. A fixed sleep can hide a race on a slow run and waste time on a fast one; use a delay only when a specific timed behavior is itself under test.
// Preferred: wait for a meaningful state.
await expect(page.getByTestId('product-grid')).toBeVisible();
await expect(page.getByTestId('product-grid')).toHaveAttribute('aria-busy', 'false');
// For a known application request, register the wait before the action/navigation.
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/products') && response.status() === 200
);
await page.goto('http://127.0.0.1:3000/products');
await responsePromise;
await expect(page.getByTestId('product-grid')).toBeVisible();
A network response completing does not guarantee that the UI has rendered it; pair request waits with a visible or otherwise meaningful application assertion. Likewise, “network idle” can be a poor readiness signal for pages with polling, analytics, or persistent connections. Use a fixed delay or network-idle wait only when it matches the behavior you want to verify, rather than as a universal page-ready rule.
Playwright’s screenshot assertion waits for two consecutive screenshots to match before comparing the last one with the baseline. This handles transient visual changes between captures, but delayed application work can start later. Keep the semantic state assertion even when using the built-in retry.
3. Keep capture conditions repeatable
Capture the same state under the same conditions when creating and comparing a baseline. Browser rendering can differ across host operating systems, browser builds, hardware, power conditions, settings, and headless mode. If the product must support multiple rendering targets, maintain appropriate baselines for separate projects instead of mixing them.
- Browser and OS: use the same Playwright browser project and CI image for baseline and comparison.
- Viewport and device scale: set dimensions explicitly and keep device scale consistent. A different breakpoint or pixel density can alter layout or image size.
- Fonts and assets: ensure required fonts load in both environments. Font substitution changes glyph widths and can move entire sections.
- Data and locale: seed stable content, dates, and user state; set locale and timezone where they affect visible output.
- Browser state: control authentication, cookies, local storage, and consent state. Start from a known state for each test.
Playwright’s screenshot assertion defaults to disabling animations and hiding the text caret. With animations disabled, finite animations are fast-forwarded and infinite animations are canceled for the capture, then resumed. If the animation itself is the feature under test, allow or test it deliberately instead of suppressing it. For other supported controls, see the API options.
4. Choose the right capture scope and control volatile content
Use a viewport screenshot when the tested behavior concerns what fits on screen at a particular scroll position. Use fullPage: true to compare the full scrollable page. For a component whose surrounding page contains unrelated dynamic content, a locator screenshot can make the failing region easier to interpret.
// Compare a component rather than unrelated page content.
await expect(page.getByTestId('checkout-summary')).toHaveScreenshot('checkout-summary.png');
// Capture the full scrollable page.
await expect(page).toHaveScreenshot('article-page.png', { fullPage: true });
// Mask only content outside the behavior under test.
await expect(page).toHaveScreenshot('account-page.png', {
mask: [page.getByTestId('current-time'), page.getByTestId('avatar')],
maskColor: '#777777',
});
A mask covers the selected locator’s bounding box. If that element moves or resizes, its mask can conceal the very layout change you need to catch. Mask timestamps or randomized avatars only when their contents are irrelevant to the test; do not mask a shifting ad slot or banner if its effect on page geometry matters.
You can apply a stylesheet during a screenshot assertion to hide or alter known volatile elements. Playwright’s stylePath stylesheet applies through shadow DOM and inner frames. Keep such overrides narrowly scoped and documented so they do not hide regressions in the area being tested.
5. Tune comparison tolerance deliberately
Playwright offers three main difference controls: threshold sets the accepted perceived color difference for an individual pixel; maxDiffPixels limits the absolute number of differing pixels; and maxDiffPixelRatio limits the differing fraction of the image. The documentation does not prescribe a universal value. Start strict, inspect representative diffs, and set a tolerance that reflects the visual risk in that test.
await expect(page).toHaveScreenshot('dashboard.png', {
threshold: 0.2,
maxDiffPixels: 80,
// Alternatively, set maxDiffPixelRatio for a size-relative allowance.
});
Do not combine broad masks and generous tolerances to make a flaky test pass without understanding why it changes. A tolerance can reduce insignificant rendering noise, but it can also allow a real regression through. Prefer stabilizing data, resources, and environment first.
6. Diagnose a failing screenshot before updating it
- Inspect the actual, expected, and diff images. Find whether the difference is a global shift, a changed component, missing content, or a small rendering variation.
- Confirm the application state. Check whether the test captured a loading state, stale results, or content before a delayed update.
- Look for common layout-shift sources. Asynchronous resources, dynamically inserted elements, images or video without reserved dimensions, font substitution, and resizing third-party content can move existing content. See web.dev’s explanation of Cumulative Layout Shift.
- Compare run conditions. Check browser, OS image, viewport, device scale, fonts, locale, data, and headless settings against the baseline run.
- Fix the cause or define intentional volatility. Reserve media dimensions, wait on the real UI state, use stable test fixtures, or narrowly mask content outside the test’s scope.
- Review before updating. Only regenerate a baseline after confirming the change is intentional. Run
npx playwright test --update-snapshotsand review the resulting image diff before committing it.
CLS and screenshot diffs answer different questions. CLS measures unexpected user-visible movement during a page experience; screenshot comparison asks whether the captured pixels differ from a baseline. A clean screenshot diff does not prove there was no movement before capture, and a CLS result does not identify whether a particular baseline image changed.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot assertion times out | The page never produces two consecutive matching captures, or the expected state is not reached. | Check the app-state assertion and inspect animations, polling updates, rotating content, and delayed resources. Stabilize the source or isolate the relevant component before increasing the timeout. |
| Baseline differs only in CI | OS, browser build, fonts, hardware, headless mode, viewport, or device scale differs. | Generate and compare snapshots in the same environment. Pin the browser/container image and set viewport and scale explicitly. |
| Text wraps differently | A font did not load or a different font version is installed; the viewport may also differ. | Wait for the intended font and app state, install consistent fonts in CI, and confirm viewport dimensions. |
| Content jumps after the assertion starts | A later asynchronous task, image, embed, ad, or widget changed the layout after an initially stable moment. | Wait on the component’s actual completion signal and inspect resource timing. Reserve image dimensions or control the third-party content in test setup when appropriate. |
| A mask hides an unexpected gap or shift | The masked element’s bounding box moved, covering relevant pixels. | Remove or narrow the mask, or mask only the text/content while preserving geometry. Keep layout itself in scope if it matters. |
| Many harmless antialiasing pixels fail | Rendering differences or a too-strict per-pixel color threshold. | First align environment and fonts. Then adjust threshold or a small pixel allowance and inspect the diff to ensure meaningful changes still fail. |
| Updating snapshots makes the suite pass but uncertainty remains | The new baseline was accepted without deciding whether the visual change is intended. | Review actual and expected images with the change, confirm the owner-approved design behavior, then update and commit the baseline. |
8. Performance, reliability, and cost
Screenshot assertions take additional browser capture and image-comparison time, and full-page images contain more pixels than component captures. Keep tests focused on important visual states, reuse deterministic fixtures, and compare a component when page-wide content is irrelevant. A delay-based wait adds at least that delay to every run; a state assertion can finish as soon as the condition is satisfied.
Reliability comes from making state and environment repeatable, not from suppressing all differences. Preserve baseline images in version control, review diffs in code review, and treat updates as changes to expected behavior. For hosted visual review, BrowserStack Percy documentation describes its CI/CD visual-testing workflow. Applitools Eyes describes its vendor-provided visual testing approach. Evaluate either against your framework, browser coverage, review process, data requirements, and current pricing; this guide does not establish comparative performance or pricing.
Or skip the browser setup
For a one-off capture or an external page, ScreenshotNeo returns a screenshot from one GET request. See the API documentation for options, including viewport, full-page capture, waiting, caching, and output format.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say which page verdict applied and whether the request was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots each month with no card; paid plans start at $5 for 3,000 shots. Those benefits make it useful for clean captures and agent workflows, while a Playwright baseline test remains the appropriate way to assert that a particular application state matches an approved image.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Does Playwright’s screenshot assertion wait for all network requests?
No. It waits for two consecutive screenshots to match. Your test should separately assert the application state it needs, such as loaded results or a visible component.
Should I use CLS or screenshot comparison to catch layout movement?
Use CLS to assess unexpected movement over the page experience and screenshot comparison to detect pixel differences at a chosen capture point. They are complementary measurements.
When should I update a visual baseline?
After reviewing the diff and confirming the new appearance is intentional. An automatic baseline update can otherwise turn an unintended regression into the new expected image.
Can I compare screenshots across browsers?
Yes, but keep an appropriate baseline for each browser and platform project. Different rendering environments can produce different pixels even when the page behavior is correct.


