End-to-End Testing for Websites: A Practical Guide
Build reliable website E2E tests by choosing critical journeys, controlling test data, and running an isolated browser suite in CI.
End-to-end (E2E) tests drive a real browser through important user journeys and verify that the visible site works with its backend and relevant integrations. Start with a small set of high-risk flows—such as sign-in, submitting a key form, or purchasing—then make each test own its data and run independently. Use component and API tests for narrower questions; reserve browser tests for behavior that needs the full application path. Cypress lists authentication, purchasing, multi-screen persistence, and pre-deployment smoke checks as common E2E scenarios. Cypress: testing types
1. Choose the right journeys and test layers
Write down the user goal, starting state, important action, and observable success condition for each candidate journey. Prioritize flows whose failure blocks a meaningful task or creates serious operational risk. Avoid turning every validation rule or visual detail into a browser test: E2E suites need more setup and maintenance than narrower tests.
| Test layer | Use it for | Example |
|---|---|---|
| Unit | Pure logic and small functions | Price calculation for a set of inputs |
| Component | A UI part in isolation | Form validation and error presentation |
| API | Backend contracts and fast setup | Create a test order or confirm an authorization response |
| E2E | A critical rendered journey across application pieces | Sign in, add an item, submit an order, and see confirmation |
Use API requests to prepare state when the behavior under test is not the setup itself. Then verify the user-visible outcome through the browser. Cypress describes API tests as faster than E2E tests because they do not render a page or simulate user interaction, while still reaching backend behavior. Cypress: testing types
2. Make test data and environments predictable
Use a dedicated test environment and accounts the team controls. Each scenario should explicitly create or reset the state it requires: an empty cart, a known user, or a seeded record. Do not rely on a teammate’s staging data or on tests running in a particular order. Cypress documents using Node tasks or HTTP requests to reset and seed data. Cypress task command
- Give each test unique data when tests may run concurrently.
- Make setup idempotent where possible, so retries do not create confusing duplicates.
- Clean up test-created records or use disposable environments.
- Keep credentials in CI secrets; never commit live credentials.
- Stub third-party responses when the third party is not part of the behavior you own. Test real integration behavior separately in a controlled environment.
Playwright recommends controlling database data and using a staging environment that does not change unexpectedly. Playwright best practices
3. Write a first Playwright test
The following example uses Playwright Test with TypeScript. It assumes the site provides an accessible sign-in form and a test-only API endpoint at /test-support/reset that creates a known user. Adapt that endpoint and the final URL to your application; do not expose test-support routes in production.
import { test, expect } from '@playwright/test';
test('a user can sign in and open their account', async ({ page, request }) => {
// Reset or seed server state through a test-only endpoint.
const seed = await request.post('/test-support/reset', {
data: {
user: {
email: 'e2e@example.test',
password: 'correct-horse-battery',
},
},
});
expect(seed.ok()).toBeTruthy();
await page.goto('/login');
await page.getByLabel('Email').fill('e2e@example.test');
await page.getByLabel('Password').fill('correct-horse-battery');
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page).toHaveURL(/\/account/);
await expect(page.getByRole('heading', { name: 'Your account' }))
.toBeVisible();
});
Install and run it in an existing Node project:
npm init playwright@latest
npx playwright test
The setup command creates starter configuration and examples. In a real application, configure the test base URL and provide the seed endpoint or another controlled setup mechanism. For example:
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
timeout: 30_000,
expect: { timeout: 5_000 },
retries: process.env.CI ? 1 : 0,
workers: process.env.CI ? 1 : undefined,
use: {
baseURL: process.env.E2E_BASE_URL ?? 'http://127.0.0.1:3000',
trace: 'retain-on-failure',
screenshot: 'only-on-failure',
video: 'retain-on-failure',
},
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
],
reporter: [['list'], ['html', { open: 'never' }]],
});
Use the retry setting as a diagnostic aid, not as a way to hide unstable tests. The example keeps local runs unretried and captures trace, screenshot, and video artifacts on failure; choose artifact retention and browser projects to suit your CI and privacy requirements.
4. Make browser tests reliable
Use locators tied to user-facing behavior
Prefer roles, accessible names, labels, and explicit test IDs when a stable contract is needed. Avoid long CSS and XPath chains or selectors based on incidental styling. A role locator does not itself prove that a page is accessible; keep accessibility checks explicit. Playwright locators auto-wait and retry, and its best-practices guidance recommends user-facing attributes and explicit contracts. Playwright best practices · Playwright locators
// Prefer an accessible name or label.
await page.getByRole('button', { name: 'Place order' }).click();
await expect(page.getByRole('status')).toHaveText('Order confirmed');
// Use a documented test ID if no useful user-facing locator exists.
await page.getByTestId('order-summary').getByText('Total').isVisible();
Wait for conditions, not guessed durations
Use web-first assertions such as toBeVisible() or toHaveURL(), which wait for the condition. Fixed sleeps such as waitForTimeout(2000) make tests slower when the app is fast and still flaky when it is slower than expected. If the app has a specific asynchronous state, wait for that state or the relevant response.
Isolate every test
Each test should work alone, with its own browser context and data setup. Playwright’s documentation says: “Each test should be completely isolated from another test and should run independently with its own local storage, session storage, data, cookies etc.” Playwright best practices
Shared authentication setup can save repeated login steps, but keep tests independent in their data and avoid one test modifying state another expects. Use unique records or serial execution only when the workflow truly requires shared state.
5. Compare Playwright and Cypress by fit
| Decision | Playwright | Cypress | How to choose |
|---|---|---|---|
| Browser coverage | One API drives Chromium, Firefox, and WebKit. | Documents cross-browser testing, including running CI tests across Firefox and Chrome-family browsers. | Match the configured matrix to the browsers your product supports; verify the specific browser and CI setup. |
| Workflow and scope | Playwright Test includes auto-waiting, assertions, tracing, and parallelism. | Documents E2E, component, API, and accessibility testing workflows. | Try each against your team’s debugging, authoring, and test-layer needs. |
| Locators | Recommends user-facing attributes and explicit contracts; locators auto-wait and retry. | Its guidance recognizes test IDs as resilient, while locator choice alone does not establish accessibility. | Agree on stable selectors and accessibility assertions regardless of framework. |
| Data setup | Recommends controlled data and stable staging. | Documents Node tasks and HTTP requests for reset and seed operations. | Check how the framework fits your backend and environment lifecycle. |
These documented capabilities do not establish a universal winner, or prove that one framework is faster or more reliable for every application. Compare both using a representative flow and the browsers, data setup, CI, and debugging tools your team actually needs. Playwright browsers · Cypress testing types · Cypress cross-browser testing
6. Run the suite in CI
Run critical E2E tests on pull requests or commits, then run a broader browser matrix on a schedule or before release if its runtime is too costly for every change. Install browser binaries and operating-system dependencies in the job. Playwright’s CI guide documents these setup steps and recommends one worker in CI as a stable starting point; sharding can distribute work across jobs. Playwright CI
# Typical Playwright CI commands
npm ci
npx playwright install --with-deps
npx playwright test
Example GitHub Actions workflow:
name: E2E tests
on:
pull_request:
push:
branches: [main]
jobs:
e2e:
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
cache: npm
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
env:
E2E_BASE_URL: ${{ secrets.E2E_BASE_URL }}
TEST_API_TOKEN: ${{ secrets.TEST_API_TOKEN }}
- uses: actions/upload-artifact@v5
if: ${{ !cancelled() }}
with:
name: playwright-report
path: playwright-report/
retention-days: 14
Pin Node and framework versions through the project lockfile and CI configuration. Upload reports and traces with a retention period appropriate for the data they may contain. Playwright notes that browser binaries are tied to Playwright versions; its CI guide says caching those binaries is not generally recommended because restore time can be comparable to downloading them. Playwright CI
7. Treat accessibility automation as one layer
Automated scans can catch known classes of accessibility issues, but they do not prove an interface is accessible. Cypress explicitly says manual testing is still needed. Cypress accessibility overview
Pair scans with journey-specific checks: confirm form fields have labels, buttons have understandable names, expected semantic elements are present, keyboard users can complete key actions, and focus moves or returns appropriately after dialogs and errors. Manually evaluate critical flows with assistive technology and keyboard-only navigation. A passing automated scan or a role-based locator is useful evidence, not a certification.
8. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Element not found or click times out | Wrong accessible name, UI state not reached, duplicate match, or brittle selector. | Inspect the rendered page and trace; use a role/name or label locator, narrow the locator, and wait on the meaningful UI condition. |
| Test passes alone but fails in suite | Shared server data, order dependency, or reused external state. | Reset or create data per test, give records unique names, and remove cross-test assumptions. |
| Intermittent navigation or assertion failure | Fixed sleeps, race with an asynchronous response, overloaded CI worker, or uncontrolled third party. | Wait for a web-first assertion or specific response, control/stub dependencies outside the test scope, and reduce CI concurrency while diagnosing. |
| Browser fails to launch in CI | Missing browser binaries or OS dependencies, or a browser/framework version mismatch. | Install with npx playwright install --with-deps for the installed Playwright version; use DEBUG=pw:browser npx playwright test to inspect launch logs. Playwright CI troubleshooting |
| Seed request fails | Wrong base URL, absent test-support route, invalid test credentials, or unavailable test environment. | Check the CI environment variables and service readiness; verify the seed endpoint is available only in the controlled test environment. |
| Screenshot differs across machines | Different browser, operating system, fonts, viewport, or dynamic content. | Standardize the environment and viewport, control time and data where possible, and mask genuinely dynamic regions rather than broad areas. |
9. Performance, reliability, and cost
Browser tests consume more time and compute than API or component tests because they launch browsers and exercise rendered pages. Keep the suite focused on a small number of high-value journeys, and move narrow logic and setup checks to faster layers. Use parallel workers or sharding only after tests are isolated; competing for the same accounts or records can create failures that look like application bugs. Keep traces and videos for failures to make CI investigation practical, while limiting artifact access and retention because they may contain page data.
Reliability comes from deterministic state, explicit waits, stable locators, and controlled dependencies. Retries can reveal intermittent failures, but a passing retry does not explain or repair the underlying flake. Track repeated failures, inspect traces, and fix the source. Costs depend on your browser matrix, run frequency, execution duration, and CI provider; no single framework comparison establishes a universal cost or speed result.
Or skip the browser setup
If you need screenshots as part of a review, report, or visual check, ScreenshotNeo is a website screenshot API and MCP server. A screenshot is useful evidence for visual review, but it does not replace an E2E assertion that a workflow succeeds.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for configuration. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.
FAQ
How many E2E tests should a website have?
There is no useful universal count. Cover the small set of journeys whose failure would block important user goals, and use narrower tests for the rest.
Should E2E tests call real third-party services?
Only when verifying that integration is the purpose of the test and you can control its environment. Otherwise, stub the dependency so its availability or content does not make an unrelated journey unstable.
Do passing accessibility scans mean a site is accessible?
No. Automated checks find some issues; manual evaluation and task-specific keyboard, focus, and assistive-technology checks remain necessary.
Can screenshots replace E2E tests?
No. Screenshots show rendered appearance at a point in time. E2E tests verify interactions and outcomes across the application journey.


