How to Automate Functional End-to-End Tests Across Platforms
Build reliable cross-browser end-to-end tests with Playwright projects, isolated workflows, repeatable CI, and traces that make failures easier to diagnose.
For a browser-based web application, automate functional end-to-end tests by choosing a few important user journeys, writing each test so it can run independently, and using Playwright projects to run those tests across the browser engines and device profiles your users need. Run the suite in CI with a stable configuration, then use traces to diagnose failures. Browser projects can cover Chromium, Firefox, WebKit, and emulated device profiles; they do not by themselves prove that a native iOS or Android app works.
This guide builds a runnable Playwright example, shows a deliberate cross-browser matrix, and explains setup, CI, debugging, scaling, and the boundary between browser testing and native app testing.
1. Define what “across platforms” means for your application
Before choosing tools, list the surfaces you need to validate. “Platforms” can mean different browser engines, desktop and mobile browser viewports, separate deployment environments, or native mobile and desktop applications. These are different test scopes.
| Scope | What this guide covers | What to keep in mind |
|---|---|---|
| Web app across browser engines | Chromium, Firefox, and WebKit projects running the same tests | Useful for detecting browser-specific behavior in web applications. |
| Responsive or mobile web | Emulated viewport and device profiles in browser projects | Emulation is not the same as validating every physical device or OS behavior. |
| Staging and production checks | Separate project configurations with different base URLs | Use production checks carefully; avoid destructive test data or actions. |
| Native iOS, Android, or desktop app | Not established by browser projects alone | Choose a platform-specific automation approach and include real-device validation where native behavior matters. |
Playwright’s project configuration lets one suite run against multiple browsers, emulated devices, and environments. Choose combinations according to your supported users and failure risks. Running every test on every conceivable combination can make feedback slow without adding useful coverage.
2. Start with user-visible acceptance outcomes
Pick a small set of important journeys, such as creating an account, signing in, searching, or completing a checkout. For each journey, describe the outcome a user can observe and the conditions that count as a failure.
Playwright’s Best Practices documentation advises: “Automated tests should verify that the application code works for the end users, and avoid relying on implementation details such as things which users will not typically use, see, or even know about such as the name of a function, whether something is an array, or the CSS class of some element.” In practice, assert visible text, accessible roles, navigation, and meaningful page state rather than internal function names or fragile styling classes.
Prefer accessible locators such as getByRole and getByLabel. A role and accessible name usually describe how a user encounters a control. If your app lacks accessible labels, improve the app markup where possible; use a stable test identifier for controls that cannot be located reliably by their user-facing semantics.
3. Create an isolated Playwright suite
Install Playwright Test in a JavaScript project and create a minimal configuration and test. The example assumes your web application is available at http://127.0.0.1:3000 and has a sign-in link or button and a page heading after navigation. Change the URL and assertions to match the application under test.
npm init -y
npm install --save-dev @playwright/test
npx playwright install
Create playwright.config.ts:
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
timeout: 30_000,
expect: { timeout: 5_000 },
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 1 : 0,
reporter: process.env.CI
? [['list'], ['html', { open: 'never' }]]
: 'list',
use: {
baseURL: process.env.BASE_URL ?? 'http://127.0.0.1:3000',
trace: 'on-first-retry',
},
projects: [
{
name: 'chromium',
use: { ...devices['Desktop Chrome'] },
},
{
name: 'firefox',
use: { ...devices['Desktop Firefox'] },
},
{
name: 'webkit',
use: { ...devices['Desktop Safari'] },
},
{
name: 'mobile-chrome',
use: { ...devices['Pixel 7'] },
},
],
});
Create tests/sign-in.spec.ts:
import { test, expect } from '@playwright/test';
test('visitor can open the sign-in page', async ({ page }) => {
await page.goto('/');
await page.getByRole('link', { name: /sign in/i }).click();
await expect(page).toHaveURL(/sign-in|login/);
await expect(
page.getByRole('heading', { name: /sign in|log in/i })
).toBeVisible();
});
Run the suite:
npx playwright test
npx playwright show-report
The test uses Playwright’s built-in page fixture, which provides a fresh page and isolated browser context for each test. Avoid making one test depend on another test having created a record or signed in. The Best Practices guide states: “Each test should be completely isolated from another test and should run independently with its own local storage, session storage, data, cookies etc.” Create the data and browser state a scenario needs as part of that scenario’s setup, and clean up test data when appropriate.
4. Configure the platform matrix deliberately
Playwright’s projects array is the main way to run the same test files under different configurations. A project can select a browser engine, use a device profile, or supply configuration such as a base URL.
Browser engines
The example config runs desktop projects for Chromium, Firefox, and WebKit, plus a mobile Chrome emulation profile. You can run one project while developing, then run the full matrix in CI:
npx playwright test --project=chromium
npx playwright test --project=webkit
npx playwright test
Use a browser project for each engine that matters to your supported users. Branded browser channels can also be configured when validating a specific installed browser channel is part of your requirement; consult the official project documentation for the current configuration options.
Device profiles and viewport coverage
Playwright’s device descriptors provide settings such as viewport dimensions and mobile behavior for browser emulation. Replace Pixel 7 with a descriptor from the installed Playwright version, or define your own viewport and device settings. A viewport-only project can be simpler when the requirement is responsive layout rather than a particular emulated device profile.
import { defineConfig } from '@playwright/test';
export default defineConfig({
projects: [
{
name: 'tablet-webkit',
use: {
browserName: 'webkit',
viewport: { width: 820, height: 1180 },
isMobile: true,
hasTouch: true,
},
},
],
});
Emulation checks browser behavior with configured device characteristics; it does not establish that the app behaves correctly on every physical phone, carrier, OS release, or native app runtime. Add real-device checks when those conditions are relevant.
Environment projects
When the same tests need to target separate deployments, use project-specific settings. Be deliberate about which projects run against production, and keep tests non-destructive there.
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
projects: [
{
name: 'staging-chromium',
use: {
...devices['Desktop Chrome'],
baseURL: 'https://staging.example.com',
},
},
{
name: 'production-smoke',
testMatch: /.*\.smoke\.spec\.ts/,
use: {
...devices['Desktop Chrome'],
baseURL: 'https://www.example.com',
},
},
],
});
Keep credentials and environment-specific values in CI secrets or environment variables rather than committing them to the repository. If tests require authenticated state, create it through a controlled setup step and ensure that tests remain isolated from each other.
5. Run the suite reliably in continuous integration
A CI run should install the project dependencies and the browsers and OS dependencies required by Playwright, then run the intended test projects on commits or pull requests. The exact YAML varies by CI provider, but the essential commands are:
npm ci
npx playwright install --with-deps
npx playwright test
Use a pinned lockfile so CI installs repeatable package versions. Keep the Playwright package and installed browser revisions aligned by installing browsers through the Playwright CLI for that project. Save the HTML report and relevant failure artifacts through your CI provider’s artifact mechanism so a failed run can be inspected.
Playwright recommends setting workers to “1” in CI environments to prioritize stability and reproducibility. A single worker is a sensible starting point, especially on shared or constrained runners. Once test data is isolated and runner capacity is understood, increase workers or split the project matrix across CI jobs. Playwright supports sharding for distributing a suite among jobs.
Use a container when a consistent OS and dependency set is helpful. For hosted browser execution at scale, Microsoft documents Playwright Workspaces as an option; service configuration and availability depend on Azure. See the Playwright Workspaces quickstart.
6. Diagnose failures with traces
A pass/fail result tells you which test failed, but often not why. The configuration above records a trace on the first retry. Open a trace from a failed test with:
npx playwright show-trace path/to/trace.zip
The Trace Viewer provides a timeline, DOM snapshots, and network requests. Use those records to see whether the test clicked too early, a request failed, a navigation did not happen, or an assertion expected the wrong state. A trace is generally more useful than adding a fixed sleep: waits can hide timing problems and make the test slower without addressing the underlying cause.
Keep a full-suite run in the validation process. Playwright’s CI documentation notes that --only-changed is heuristic and can miss tests, so it can help with preliminary selection but should not replace full-suite validation.
7. Keep failures actionable and tests stable
- Make setup explicit: create the user or records a scenario needs and clean them up when possible.
- Use condition-based waits: prefer locator assertions and navigation expectations over arbitrary timeouts.
- Keep selectors meaningful: favor roles, labels, and visible text; use stable test IDs when user-facing semantics are insufficient.
- Separate test data: avoid shared mutable accounts or records when projects or workers execute concurrently.
- Retain useful artifacts: keep reports and traces for failures long enough for the team to inspect them.
- Revisit the matrix: add a browser, device profile, or environment when supported users or observed risk justify it.
8. Performance, reliability, and cost trade-offs
Cross-platform coverage multiplies executions. If a test suite has T tests and you run it in P projects, the rough number of test-project executions is T × P, before retries. Choose projects based on user impact and browser support; avoid adding a project that does not answer a useful compatibility question.
- Feedback time: one worker favors predictable execution but takes longer than parallel jobs. Increase concurrency only after validating isolation and runner capacity.
- Flakiness: retries can help collect traces and reduce transient CI noise, but repeated retries can also conceal a real reliability problem. Track and fix recurring failures.
- Infrastructure cost: self-hosted runners trade machine maintenance for control; hosted browser services can reduce browser infrastructure work but introduce service setup and usage costs. Check current vendor pricing and availability directly before choosing.
- Artifact storage: traces and reports consume storage. Retain them for failed or retried tests according to the time your team needs to diagnose problems.
- Browser maintenance: update Playwright and browser revisions deliberately so local and CI environments stay aligned.
9. Know when browser tests are not enough
Playwright projects document browser engines and emulated device profiles. They do not establish full native iOS or Android automation or native desktop application coverage. If the requirement includes native permissions, app lifecycle behavior, push notifications, platform-specific controls, or hardware interactions, define a separate native test scope and choose tools and devices for that scope. Browser tests remain useful for the web surface, including a mobile website, but should not be described as proof that a native app works.
The research used for this guide did not establish a current, authoritative comparison of native automation frameworks, so it would be misleading to name a universal framework or claim one suite covers browser, native mobile, and desktop software equally. A 2021 research preprint on one image-driven mobile replay prototype reported specific experimental results; those figures describe that prototype and experiment, not expected production accuracy, so they are not used as a general benchmark here.
10. Troubleshooting common Playwright E2E problems
| Symptom | Likely cause | Fix |
|---|---|---|
| “Executable doesn’t exist” or browser launch fails | Browser binaries were not installed for the Playwright version in the project, or the CI image lacks required OS dependencies. | Run npx playwright install --with-deps in CI and ensure the installed Playwright package and browser revisions match the lockfile. |
| Tests pass locally but fail in CI | Different environment, missing dependencies, shared test data, or execution timing differences. | Reproduce with the CI configuration, retain the trace and report, make setup independent, and inspect network and DOM state in Trace Viewer. |
| “Strict mode violation” for a locator | The locator matched more than one element. | Use a more specific role/name or scope the locator to a meaningful container. Avoid selecting an arbitrary first match unless order is part of the requirement. |
| Timeout waiting for an element | The element never appeared, the page is in a different state, or the app/request is slow or failing. | Inspect the trace and network activity, verify the expected page state, and wait on the actual condition rather than increasing timeouts blindly. |
| State leaks between tests | Tests reuse accounts, cookies, storage, or mutable backend records. | Use isolated browser contexts and unique test data; make every test pass when run alone and in a different order. |
| Mobile project behaves unlike a real phone | Browser device emulation is being treated as full physical-device coverage. | Use emulation for responsive browser checks and add real-device validation for OS, hardware, or native behavior. |
| Suite is too slow after adding projects | Every test now runs in every browser and device configuration. | Reduce the matrix to meaningful combinations, split projects into CI jobs, or shard the suite after confirming runner capacity. |
| Retries make the pipeline appear green despite recurring failures | Transient failures are being masked by retry success. | Review retried tests and their traces, fix the underlying issue, and use retries to gather diagnostic evidence rather than as a substitute for reliability. |
11. Or skip the browser setup
For screenshot evidence alongside a functional suite, ScreenshotNeo provides a website screenshot API and MCP server. It does not replace interaction and assertion tests: use Playwright to exercise the journey, and use a screenshot capture when a clean visual artifact is useful.
One GET request captures a page; see the ScreenshotNeo API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners are accepted and removed before the shot; known consent platforms, newsletter popups, and chat widgets are removed too.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses include page-verdict and billing headers.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, no card required.
12. Frequently asked questions
Can one Playwright test run on multiple browsers?
Yes. Configure browser projects and run the same test files in each project. Keep the set aligned with browsers your users and support policy require.
Does mobile emulation test a native mobile app?
No. It emulates browser behavior and device characteristics for web testing. Native app behavior needs its own platform-specific coverage.
Should every test run on every browser?
Not necessarily. Start with the combinations that cover meaningful users and risk, then expand where compatibility evidence calls for it.
Should CI run one worker or many?
Start with one worker for stability and reproducibility. Add workers or shards after confirming runner capacity and test isolation.
Are screenshots functional end-to-end tests?
A screenshot captures visual output at a point in time; it does not by itself verify that a user can complete an interaction. Pair visual evidence with functional assertions when both matter.


