ScreenshotNeo

BlogAI agents

How to compare AI-agent website screenshots across desktop and tablet breakpoints

Compare AI-agent screenshots at desktop and tablet widths with stable Playwright captures, separate visual baselines, and a repeatable review process.

By the ScreenshotNeo team4 October 20267 min read

To compare AI-agent website screenshots across desktop and tablet breakpoints, capture the same page state at explicit viewport sizes, keep rendering conditions stable, and compare each size against its own reference image. With Playwright, define separate desktop and tablet projects, save separate visual baselines, and inspect any pixel difference before deciding whether it is a real layout or behavior problem.

Emulated viewports are useful for breakpoint coverage, but they simulate browser conditions rather than proving how every physical device renders. Playwright supports device profiles and viewport overrides for this purpose. Playwright emulation documentation

1. Decide what the comparison should answer

Start by writing down the page state and the question. A useful comparison controls the route, content, interaction state, viewport, capture extent, and pixel scale. Change one axis at a time when diagnosing a difference.

Choice Use it to answer
Viewport screenshot What fits in the current browser window, including responsive layout and above-the-fold behavior.
Full-page screenshot How the entire scrollable page renders, including below-the-fold content. Keep this mode identical for reference and current captures.
CSS-pixel scale Whether CSS layout dimensions and breakpoint behavior look correct.
Device-pixel scale Whether high-density rendered detail is correct. This can produce larger images.
Explicit viewport A controlled CSS breakpoint width with fewer device-specific variables.
Emulated device profile A broader simulated device setup that may include user agent and touch behavior as well as viewport dimensions.

For a first responsive check, use explicit desktop and tablet widths with viewport screenshots at CSS-pixel scale. Add full-page or device-pixel captures when the question requires them.

2. Configure separate desktop and tablet captures

The example below uses Playwright Test and TypeScript. It creates two projects with explicit viewport dimensions and gives each project a distinct screenshot baseline. Adjust the widths to the breakpoints your interface actually supports; the values below are example test targets, not universal definitions of desktop and tablet.

// playwright.config.ts
import { defineConfig } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  projects: [
    {
      name: 'desktop',
      use: { viewport: { width: 1440, height: 900 } },
    },
    {
      name: 'tablet',
      use: { viewport: { width: 820, height: 1180 } },
    },
  ],
  expect: {
    toHaveScreenshot: {
      // Tune only after reviewing legitimate rendering variation.
      animations: 'disabled',
    },
  },
});

Install the project dependency and browser once in your development environment with npm install -D @playwright/test and npx playwright install. Then add a test such as:

// tests/responsive.spec.ts
import { test, expect } from '@playwright/test';

test('pricing page matches the breakpoint reference', async ({ page }, testInfo) => {
  await page.goto('https://example.com/pricing', { waitUntil: 'networkidle' });

  // Prefer deterministic state: dismiss dialogs, seed data, or reach
  // the same interaction state before taking either screenshot.
  await page.locator('main').waitFor({ state: 'visible' });

  const name = `${testInfo.project.name}-pricing.png`;
  await expect(page).toHaveScreenshot(name, {
    fullPage: false,
    animations: 'disabled',
  });
});

Replace the example route and readiness condition with your application. If you need to inspect the complete page, set fullPage: true in both projects. Keep the names distinct: Playwright stores and compares a reference for each named screenshot, so desktop output cannot silently stand in for tablet output. The first run creates reference images; inspect and preserve them in version control or your chosen artifact store. Later runs compare against those references. Update a baseline only after reviewing and accepting an intended design change. Playwright visual comparisons

Use an emulated profile when device behavior matters

If touch behavior or a device-like user agent is part of the question, start with a Playwright device profile and then override its viewport to the target breakpoint. Explicitly set the final dimensions after spreading the profile so the chosen test width is clear.

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  projects: [
    {
      name: 'tablet-touch',
      use: {
        ...devices['iPad (gen 7)'],
        viewport: { width: 820, height: 1180 },
      },
    },
  ],
});

Choose a profile available in the installed Playwright version. A profile is still an emulation; use actual hardware testing when you need evidence about a specific physical device, browser build, or operating-system behavior. Device emulation and viewport overrides

3. Keep the capture state repeatable

A visual comparison is only useful when a difference is attributable to the page change you are investigating. Playwright notes that operating system, browser version, settings, hardware, and headless mode can affect screenshot output. Keep baseline generation and comparison in the same environment where practical. Playwright visual comparison guidance

  • Use the same browser engine and browser version for reference and current runs.
  • Run both viewport projects in the same operating system and rendering mode.
  • Use stable test data, locale, timezone, and authentication state.
  • Wait for the specific content under test rather than relying on a short arbitrary delay.
  • Disable animations for visual assertions when motion is not the test subject.
  • Make fonts and external assets available consistently; late font loading can shift line breaks.
  • Set the same page interaction state for both breakpoints, such as an open menu or dismissed consent prompt.
  • Use the same capture extent and scale for reference and current images.

networkidle can be a convenient navigation condition, but pages with polling or persistent connections may never become idle. In those cases wait for a meaningful selector or application-ready signal instead.

4. Review diffs at both breakpoints

  1. Run the desktop and tablet projects and let Playwright report mismatches.
  2. Open the reference, actual, and diff images for each failing project.
  3. Classify the visible change: moved, wrapped, missing, overflowing, clipped, or obscured content; changed typography; or rendering noise.
  4. Check the affected width in the browser and determine whether the difference is expected, a regression, or an unstable capture.
  5. Use DOM or accessibility assertions to verify structure and interaction when the visual image alone cannot answer the question.
  6. Update only the affected baseline after confirming the change is intended.

A pixel diff identifies changed pixels; by itself, it does not establish user impact or explain the cause. A one-pixel antialiasing change may be harmless, while a small shift in a button can affect usability. Review the region and test the behavior. Playwright’s screenshot documentation distinguishes visual captures from accessibility snapshots, which expose page structure and interaction references. Playwright screenshot documentation

5. Compare with an AI agent without surrendering the evidence

An AI agent can help summarize observed differences if you provide the desktop reference/current pair and tablet reference/current pair, along with the target widths and intended page state. Keep the two breakpoint pairs labeled separately. Ask it to describe locations and observable changes, then verify those claims against the actual images and the page. The cited Playwright documentation describes screenshot and structural inspection tools; it does not establish that an agent can reliably judge the semantic severity of every visual difference.

A useful review prompt asks the agent to report: which breakpoint is affected, the approximate region, what changed visually, whether content is missing or obscured, and what follow-up DOM or interaction check would confirm the observation. Ask it to distinguish direct visual observations from guesses about cause.

6. Common problems and fixes

Symptom Likely cause Fix
Many pixels differ on every run Different OS, browser, fonts, headless mode, or unstable dynamic content. Run comparisons in a consistent environment; stabilize content and assets; inspect whether the difference is rendering noise.
Tablet screenshot looks like desktop Viewport override was not applied, or a spread device profile replaced the intended viewport. Set viewport explicitly in the final project configuration and confirm the runtime page dimensions.
Only one breakpoint gets a baseline Both captures use the same snapshot name or only one project ran. Include the project name in the screenshot name and run both projects.
Screenshot is blank or incomplete Capture occurred before the route or target content was ready. Wait for a stable selector or app-ready condition and confirm navigation succeeded.
Full-page images differ in height Content length, lazy loading, or page state changed, or capture extents differ. Stabilize test data, ensure relevant content is loaded, and use the same full-page setting.
Diff highlights animated elements Animation or caret state changed between captures. Disable animations for the assertion or deliberately control the animation state.
Reference update hides a regression Baseline was accepted without reviewing actual and diff images. Review each changed region first, then update only the intended project baseline.

7. Performance, reliability, and cost

Each project renders the page separately, so adding tablet coverage adds browser navigation and capture work. Keep the core comparison focused on representative breakpoint widths, and reserve full-page or high-density captures for cases that need them. Device-pixel screenshots can be substantially larger than CSS-pixel captures. Reuse a stable CI environment and preserve reference artifacts so failures can be reviewed after a run.

Playwright is a do-it-yourself browser setup: your team manages browser installation, execution environment, reference images, and CI artifacts. The research sources give no benchmark or monetary cost for this workflow, so evaluate it using your own page complexity, CI capacity, and maintenance needs.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API can capture the same target URL at desktop and tablet viewport dimensions; use separate calls and filenames so each condition stays distinct. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com --data-urlencode width=1440 --data-urlencode height=900 -o desktop.webp
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com --data-urlencode width=820 --data-urlencode height=1180 -o tablet.webp

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Responses include page-verdict and billing headers. Sign up free and get 1,000 screenshots a month with no card.

9. FAQ

Should the tablet width match a specific device?

Use the breakpoint or width your team needs to support. Use a named device profile when user-agent or touch behavior matters, and treat emulation as simulated coverage.

Should visual baselines be committed?

Keep references somewhere versioned and reviewable so a later run compares against an intentional known state. The Playwright workflow generates references on the first run and compares later runs against them.

Does a screenshot prove accessibility?

No. It shows rendered appearance. Use accessibility and DOM checks for semantics, names, roles, and interaction behavior.

When should I use full-page capture?

Use it when below-the-fold layout is part of the question. For a focused breakpoint check of what a user initially sees, viewport capture is usually easier to interpret.