ScreenshotNeo

BlogHow-to

How to Test Browser Compatibility with Headless Browsers

Build a reproducible browser matrix with Playwright, validate failures in real browsers, and know when headless testing is not enough.

By the ScreenshotNeo team1 October 20267 min read

Headless testing is an execution mode, not a browser-coverage strategy. To test compatibility reliably, define a browser matrix from your real users and product risks, run identical journeys in each matrix cell, pin browser binaries, and preserve enough evidence to reproduce every failure.

What headless browser compatibility testing covers

A compatibility test asks whether the same user-visible behavior works across combinations of:

  • Browser engine: Chromium, Firefox, or WebKit.
  • Browser channel and version: bundled Chromium, branded Chrome, Edge, or a pinned release.
  • Operating system and architecture.
  • Viewport, device characteristics, and mobile emulation.
  • Permissions, storage, media codecs, extensions, downloads, and other environment features.

Playwright is a strong default because its official browser documentation covers Chromium, Firefox, WebKit, branded Chrome and Edge channels, device emulation, and headless execution. Selenium WebDriver remains a strong choice when an existing WebDriver ecosystem, Grid, or browser-specific capability matters. MDN describes WebDriver as a platform- and language-neutral protocol for remotely inspecting and controlling user agents and cross-browser testing.

Sources: Playwright browser documentation, MDN WebDriver, and Selenium documentation.

1. Design a browser matrix from user risk

Start with analytics, support contracts, accessibility commitments, and the browser-sensitive features your product uses. A practical baseline is Chromium, Firefox, and WebKit (the Safari-equivalent engine). Add cells only when evidence justifies them.

Dimension Baseline Add when
Engine Chromium, Firefox, WebKit Your users or feature risk is concentrated in another engine
Channel Playwright-managed browsers Chrome or Edge branding, codecs, extensions, or enterprise policies matter
Version Current pinned version Support policy requires older releases or a rolling latest-minus-one policy
OS CI runner OS Font rendering, input, downloads, permissions, or OS integration differs
Device Desktop viewport plus one mobile profile Responsive layouts, touch, orientation, or mobile APIs are important

Hosted grids commonly make browser name, browser version, OS, and device explicit. BrowserStack documents selectors such as latest, latest - 1, and latest - 2; use an equivalent, documented policy rather than an untracked floating browser.

2. Pin Playwright and its browser binaries

Each Playwright release requires specific browser binaries. Commit your package lockfile and install the matching browsers in CI so a result can be tied to known versions.

npm init playwright@latest
npm install
npx playwright install --with-deps

Keep the Playwright package, lockfile, CI image, and browser cache under version control or documented cache keys. A cache hit must still correspond to the lockfile; otherwise a test result can silently use a different browser binary.

3. Create identical cross-browser journeys

Test behavior rather than implementation details. Cover navigation, authentication, forms, keyboard and pointer input, responsive breakpoints, media, downloads, permissions, storage, and browser-sensitive APIs. Assert visible outcomes and important console or network errors; DOM snapshots alone can miss a broken interaction.

import { test, expect } from '@playwright/test';

test('user can search and open a result', async ({ page }) => {
  const consoleErrors: string[] = [];
  page.on('console', message => {
    if (message.type() === 'error') consoleErrors.push(message.text());
  });
  page.on('requestfailed', request => {
    consoleErrors.push(`${request.method()} ${request.url()} ${request.failure()?.errorText ?? ''}`);
  });

  await page.goto('https://example.com/search', { waitUntil: 'domcontentloaded' });
  await page.getByRole('searchbox').fill('headless testing');
  await page.getByRole('button', { name: 'Search' }).click();
  await expect(page.getByRole('main')).toContainText('headless testing');
  await expect(page.getByRole('link').first()).toBeVisible();
  expect(consoleErrors).toEqual([]);
});

Use stable roles, labels, and test identifiers. Avoid timing assumptions such as arbitrary sleeps when a visible state, selector, or network condition can express readiness.

4. Configure one Playwright project per browser

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  retries: process.env.CI ? 1 : 0,
  reporter: [['html'], ['line']],
  use: {
    baseURL: 'https://example.com',
    headless: true,
    trace: 'on-first-retry',
    screenshot: 'only-on-failure',
    video: 'retain-on-failure'
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
    { name: 'firefox', use: { ...devices['Desktop Firefox'] } },
    { name: 'webkit', use: { ...devices['Desktop Safari'] } },
    { name: 'mobile-chromium', use: { ...devices['Pixel 5'] } },
    { name: 'mobile-webkit', use: { ...devices['iPhone 13'] } }
  ]
});

Run the complete matrix locally or select projects while narrowing a failure:

npx playwright test
npx playwright test --project=firefox tests/search.spec.ts
npx playwright show-report

Keep retries limited. A retry can collect evidence, but unlimited retries hide real flakiness and make a broken cell look green.

5. Preserve evidence for every failure

For each result, record the test revision, Playwright version, browser name and version, OS, viewport, device profile, locale, timezone, permissions, and relevant feature flags. Store:

  • Failure screenshot and, where useful, a full-page screenshot.
  • Trace with actions, snapshots, console output, and network activity.
  • Video for animation, timing, media, or input problems.
  • Console messages and failed requests.
  • Browser and operating-system metadata.

A matrix failure in one engine or version suggests a compatibility issue. A failure in every cell more often indicates an application, data, or fixture problem. Re-run the smallest failing test with the same binary and environment before changing code.

6. Know when headless is insufficient

Playwright distinguishes its Chromium headless shell from its newer headless mode, which uses the real Chrome browser. The real-browser mode is more authentic and offers more features. Confirm high-risk failures in headed or branded mode when testing:

  • Visual rendering, fonts, and pixel-sensitive layouts.
  • Video, audio, media codecs, or WebGL.
  • Browser extensions and enterprise policies.
  • Downloads, print flows, permissions, and file pickers.
  • Safari or Chrome-specific behavior that WebKit or bundled Chromium may not represent exactly.
npx playwright test --project=chromium --headed
npx playwright test --project=chromium --config=playwright.chrome.config.ts

Automation can also be observable. MDN documents that Chrome sets navigator.webdriver when launched with --enable-automation or --headless, while Firefox sets it with Marionette controls. Treat bot-detection behavior as a separate compatibility case; do not assume a headless result represents a normal visitor.

7. Use a hosted grid when the matrix outgrows CI

Managed services can provide operating-system, browser-version, and device combinations that are expensive to maintain locally. Keep the same test code and assertions, declare capabilities explicitly, and record the provider capability set with every result. Compare local and hosted coverage by engine, version, OS, device, reproducibility, trace quality, queue time, and cost.

Common failures and fixes

Symptom Likely cause Fix
Browser executable missing Playwright package and browser cache are out of sync Run npx playwright install --with-deps for the locked package and repair CI cache keys.
Works in Chromium, fails in Firefox or WebKit Engine-specific CSS, API, timing, or input behavior Reduce to the smallest failing test, inspect console and network errors, then use a standards-based implementation or an explicit capability branch.
Only CI fails Different OS, fonts, timezone, locale, viewport, permissions, or service dependencies Record and reproduce CI metadata locally; set timezone, locale, viewport, and permissions explicitly.
Flaky timeout Race with navigation, rendering, or an external request Wait for a user-visible state or specific response; remove arbitrary sleeps; capture a trace on retry.
Screenshot differs by browser Font availability, anti-aliasing, viewport, device scale, or genuine layout difference Pin fonts and dimensions, compare semantic assertions first, and reserve pixel thresholds for controlled visual tests.
Download or permission test fails headlessly Execution mode does not match the user flow Repeat in headed or real-browser headless mode with explicit permissions and download handling.
Bot check appears only in CI Automation signals such as navigator.webdriver, IP reputation, or missing browser features Test the protected flow in an approved environment and treat the result as a deployment or anti-bot integration concern.
Failures disappear after retries Retry count is masking nondeterminism Keep retries low, track flaky tests separately, and inspect traces instead of accepting repeated retries.

Performance, reliability, and cost

  • Parallelism: Run independent matrix cells in parallel, but cap workers to the CI CPU and the capacity of test dependencies.
  • Selective execution: Run the full matrix on release branches and a risk-based subset on every commit. Always run a focused regression set for browser-sensitive changes.
  • Warm setup: Reuse browser installation caches and authenticated storage only when test isolation remains safe.
  • External services: Stub unstable third-party calls for deterministic tests, then keep a smaller integration suite against the real service.
  • Artifacts: Retain traces and videos for failures, not every passing test, unless audit requirements justify the storage cost.
  • Hosted grids: Price by session or minute, matrix breadth, parallel capacity, and artifact retention. A smaller, well-chosen matrix is often more useful than many unreviewed cells.

Or skip the browser setup

If you need rendered evidence rather than an interactive compatibility suite, ScreenshotNeo provides a website screenshot API and MCP server. Its request accepts a URL and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.

curl -G https://api.screenshotneo.com/v1/shot -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients take screenshots, inspect pages, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge.

FAQ

Is headless testing enough for browser compatibility?

No. It is efficient for CI, but headed or real-browser confirmation is appropriate for media, extensions, downloads, permissions, visual fidelity, and branded-browser behavior.

Should every version be tested?

Test versions your users, contracts, or risk profile require. Pin the version used by each result and add a documented rolling policy only when you can reproduce it.

Can Playwright test Safari?

Playwright tests WebKit, which is Safari-equivalent at the engine level. Validate high-risk Safari behavior in the Safari environments your support policy promises.

When should Selenium replace Playwright?

Choose Selenium when your organization depends on WebDriver tooling, an existing Grid, or browser-specific capabilities that fit Selenium better. Choose Playwright when its bundled multi-engine projects and tracing match your workflow.

What is the smallest useful matrix?

Start with Chromium, Firefox, and WebKit on a pinned CI environment, one desktop viewport, and a representative mobile profile. Expand from real usage and failures.