ScreenshotNeo

BlogHow-to

How to Write Stable Cross-Browser Tests

Build cross-browser tests around user-visible behavior, isolated state, controlled dependencies, and useful failure evidence. Includes a runnable Playwright setup and CI guidance.

By the ScreenshotNeo team4 October 202610 min read

Stable cross-browser tests come from clear user-facing assertions, independent test data and browser state, controlled dependencies, and a browser matrix that matches product risk. No browser tool makes a poorly isolated test reliable by itself. Selenium’s guidance puts it plainly: “No one approach works for all situations.” Selenium Test Practices.

This guide uses Playwright with TypeScript for runnable examples. The same design principles apply to Selenium and other browser automation frameworks. The examples test an application you control; they do not depend on a third-party site.

1. Decide what “cross-browser” means for your product

Start with the browsers, versions, operating systems, and device conditions your product promises to support. Add a browser or platform when it covers a real user risk: a Safari-specific issue, a mobile layout, branded Chrome or Edge requirements, or an API that differs by platform. A matrix should answer those questions, rather than grow simply because more configurations are available.

Playwright projects can run Chromium, Firefox, WebKit, branded Chrome and Edge channels, and emulated devices. The distinction between engines and branded browsers matters: Playwright’s bundled WebKit is not branded Safari, and its bundled Firefox is not the branded Firefox application. Playwright describes its WebKit build as deriving from recent WebKit development; when Safari fidelity matters, it recommends WebKit on macOS as the closest experience. Official browser binaries can also matter for capabilities such as media codecs. See Playwright’s browser documentation.

Requirement Useful coverage choice What to keep in mind
General engine coverage Chromium, Firefox, WebKit These are framework-managed browser builds; they do not represent every branded browser and operating system combination.
Released Chrome or Edge behavior Playwright’s branded Chrome or Edge channel Use the channel when the product requirement names the branded browser.
Safari-specific behavior WebKit on macOS, and Safari validation where required Do not describe Playwright WebKit as Safari.
Responsive or touch layout Relevant emulated device or viewport project Emulation checks layout and selected device traits; it does not reproduce every physical device condition.
Media or platform APIs The branded browser and OS users run Confirm that the selected browser build provides the capability under test.

Playwright notes its Chromium can run ahead of branded stable releases. That can provide early warning about browser changes, but use branded stable channels when the requirement is specifically the currently released Chrome or Edge. Keep the matrix small enough to run frequently, then run broader coverage on an appropriate CI schedule.

2. Install Playwright and configure browser projects

Create a TypeScript project and install Playwright Test. The install command downloads Playwright’s supported browser builds; use the official installation guide for operating-system dependencies and current setup details.

npm init -y
npm install --save-dev @playwright/test
npx playwright install

Create playwright.config.ts to define the application URL and a focused engine matrix:

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  forbidOnly: Boolean(process.env.CI),
  retries: process.env.CI ? 1 : 0,
  reporter: process.env.CI ? [['html', { open: 'never' }]] : 'list',
  use: {
    baseURL: 'http://127.0.0.1:3000',
    trace: 'retain-on-failure',
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
    { name: 'firefox', use: { ...devices['Desktop Firefox'] } },
    { name: 'webkit', use: { ...devices['Desktop Safari'] } },
  ],
  webServer: {
    command: 'npm run start:test',
    url: 'http://127.0.0.1:3000/health',
    reuseExistingServer: !process.env.CI,
    timeout: 60_000,
  },
});

Add a start script appropriate to your application in package.json, for example "start:test": "your-app-command". The health URL should report readiness only after the application can serve the pages and APIs the tests need. Replace the example host, command, and health path with your own app’s values. For browser projects and channel options, consult the official browser documentation.

3. Assert user-visible behavior

Use roles, accessible names, labels, or visible text when they are stable interface contracts. Use a dedicated test identifier when the product intentionally exposes one. Avoid selectors based on styling classes or incidental DOM nesting: a CSS refactor should not break a test that checks a user workflow. Playwright recommends testing user-visible behavior and provides locator auto-waiting and retry behavior. See Playwright Best Practices.

This example assumes an application with a labeled sign-in form and a page that displays an account heading after valid submission. Adapt the labels, route, and expected result to the application’s actual contract.

import { test, expect } from '@playwright/test';

test('a user can sign in and see their account', async ({ page }) => {
  await page.goto('/sign-in');
  await page.getByLabel('Email').fill('ada@example.test');
  await page.getByLabel('Password').fill('test-password');
  await page.getByRole('button', { name: 'Sign in' }).click();

  await expect(page.getByRole('heading', { name: 'Your account' }))
    .toBeVisible();
});

A successful click only proves that the click was dispatched. Assert the resulting state that matters to a user: a confirmation, a changed status, a new page heading, or an accessible error. Prefer a locator or assertion that waits for that state to become true instead of adding a fixed sleep. A guessed delay can be too short on a slow run and needlessly long on a fast one; this is a practical consequence of using state-based locator waiting.

4. Make each test independent

A test should pass when run alone, after another test, or in a different order. Give it known data, reset mutable state at a clear boundary, and avoid relying on an earlier test to create a user, record, or browser session. Use a separate browser context or the framework’s normal per-test isolation so cookies and storage do not leak between tests.

When signing in is expensive, a controlled setup can create reusable authentication state. Keep each test’s mutable records independent, and avoid running tests in parallel against the same account data if one test can alter what another reads. Playwright’s guidance recommends independent tests and isolated state; Selenium also recommends a fresh browser per test and avoiding shared state. See Playwright Best Practices and Selenium’s encouraged behaviors.

5. Control services and data outside the test

An end-to-end test of your application should not fail because an unrelated payment sandbox, analytics endpoint, email provider, or third-party page is slow or unavailable. Mock or stub external services when their behavior is not the feature being tested. Generate or seed application state explicitly so tests do not depend on whatever data happens to exist.

If the integration itself is the subject, cover it with a focused integration test and make its dependency and failure signal explicit. Keep the ordinary user-flow suite focused on behavior your team controls. Playwright recommends testing what you control; Selenium recommends mocking external services and generating application state in its testing guidance.

6. Run a purposeful matrix in CI

Run a compact, relevant cross-browser set often, and expand it for release checks or browser-specific risk. Use a repeatable operating system and browser version for visual comparisons: otherwise an OS, font, or browser update can change rendering independently of the application change. Keep dependency and browser updates deliberate, review failures after updates, and avoid letting local and CI environments drift without a reason.

A minimal CI workflow should install dependencies, install the browser builds required by the selected projects, start the application, run the tests, and retain reports and traces when a run fails. The exact YAML depends on the CI provider and application build; do not copy a provider-specific workflow without matching its runner image and dependency installation. Playwright recommends running tests frequently and keeping visual regression environments consistent in its best practices.

For a test run, select the projects explicitly when useful:

npx playwright test --project=chromium --project=firefox --project=webkit

For a branded browser channel, add a project with the corresponding Playwright channel configuration, and install or provision that browser according to the browser documentation. Use branded channels when the product requirement calls for those releases; use bundled engines for framework-managed engine coverage.

7. Investigate failures with evidence

Keep the report, trace, and relevant CI logs for failed runs. Playwright traces include an action timeline, DOM snapshots, and network requests, which can help distinguish a product defect from a selector that stopped matching, missing test data, an environment difference, or an uncontrolled dependency. The Selenium project also recommends improved reporting. This cause grouping is a practical synthesis of the documented failure sources, not a measured ranking.

Retries can capture more evidence about an intermittent failure, but a pass on retry does not explain why the first attempt failed. Preserve first-run context and find the cause rather than treating retries as proof that the test is healthy.

8. Troubleshooting common failures

Symptom Likely cause Fix
Locator times out in one browser The accessible name or role differs, the page did not reach the expected state, or the selector depends on browser-specific markup. Inspect the trace DOM snapshot and rendered accessibility information. Assert the intended user-facing contract and correct the app if its behavior differs unexpectedly.
Passes locally but fails in CI Different OS/browser build, missing setup, slow readiness, timezone, locale, or test data collision. Compare runner and local versions, seed data deterministically, wait on state rather than time, and make environment settings explicit where they matter.
Fails only when the suite runs in parallel Tests share mutable records, accounts, files, or external resources. Give tests separate data or serialize the genuinely shared resource; reset state at the test boundary.
Intermittent network or third-party error The test depends on a service outside the product team’s control. Mock that service for the product-flow test, or move the integration check into a focused test with explicit dependency handling.
Unexpected differences in media or browser APIs A bundled engine differs from the branded browser, or the OS/browser build lacks the target capability. Run the branded browser and operating system required by the feature; confirm the selected channel and browser build.
Visual snapshots change after a seemingly unrelated update Browser, operating system, fonts, or rendering environment changed. Compare on a consistent OS and browser version, then review whether the rendering change is expected before updating snapshots.
Retry passes but initial attempt fails Timing, shared state, or an uncontrolled dependency is still intermittent. Use the first attempt’s trace and logs to identify the cause; do not use the retry as the fix.

9. Keep runtime, reliability, and maintenance costs visible

Each extra browser, device, and operating system project adds execution and maintenance work. Choose projects from supported-browser commitments and risk, run the focused set frequently, and reserve broader combinations for the cases that justify them. Split or parallelize work only when tests are independent and the CI capacity makes that useful.

Reliability comes from repeatable setup, isolated data, controlled dependencies, and assertions against meaningful states. Browser and framework updates can expose legitimate compatibility changes, so update deliberately and keep the environment identifiable in reports. There is no source-backed universal browser count or flakiness reduction percentage; choose based on the product’s support promises and risk.

10. Screenshot evidence for review and debugging

When a failure needs a visual artifact, a screenshot can show what the page looked like at capture time. Use the browser test’s trace and screenshot output for evidence tied to a particular test run. For separate URL-based captures, ScreenshotNeo is a website screenshot API and MCP server for developers; its clean-capture options accept consent banners and remove supported consent platforms, newsletter popups, and chat widgets before capture. Its verdict and billing headers distinguish clean captures from bot checks, blank pages, failed loads, and cache hits.

Or skip the browser setup

If your task is to capture a page rather than validate interactive behavior across browser engines, one API request can return an image. This does not replace cross-browser tests, but it avoids setting up browser automation for URL-based screenshots. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo removes supported cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card.

FAQ

How do I stop cross-browser tests from flaking?

Make data and browser state independent, control external dependencies, and wait for the user-visible state the test cares about. Use traces to investigate intermittent failures instead of masking them with longer sleeps or retries.

Should I test Safari or WebKit?

Use the environment that matches the requirement. Playwright WebKit is not branded Safari; Playwright identifies WebKit on macOS as the closest Safari experience. Validate in Safari when the product’s support commitment requires the branded browser.

How many browsers should an end-to-end suite cover?

Cover the engines and branded browsers your product supports, then add device and OS combinations tied to specific risks. No fixed number fits every product.

Do cross-browser tests need screenshots for every assertion?

No. Assert behavior for ordinary workflows. Keep screenshots, traces, and visual comparisons where they add evidence or verify rendering requirements.