Real-World Testing: A Practical Guide to End-to-End Testing
Learn what belongs in end-to-end tests, how to balance them with faster checks, and how to make critical browser journeys reliable in CI.
End-to-end (E2E) testing checks whether a small number of important user journeys work across the running application, from the interface through backend services and relevant integrations. Keep E2E tests focused on critical and high-risk flows; use unit, component, API, and integration tests for the faster, narrower checks that make up most of the suite. Reliable browser tests use isolated data, user-facing interactions, and assertions that wait for the expected condition.
This guide uses Playwright with TypeScript for a runnable example. The same planning principles apply to other browser testing tools. For visual documentation of a running page, ScreenshotNeo is a separate screenshot API; a screenshot does not replace assertions about whether a user journey works.
1. What end-to-end testing checks
An E2E test exercises an application in a browser and checks behavior across the integrated system. A realistic test might sign in, complete a purchase, and confirm that the resulting order appears in the account. That can cover the UI, backend, persistence, and a payment integration in one journey. Cypress describes E2E testing as testing from the browser through the backend and third-party APIs or services.
Use E2E tests to answer an integration question: Can a user complete this important task through the running application? They can provide confidence across system boundaries, but failures may have a wider range of causes than a narrow unit or integration test, and the full environment takes more effort to run and maintain.
Good candidates for E2E coverage
- Authentication and account recovery paths that matter to users.
- Purchasing or another high-value transaction.
- Data that must persist and appear correctly on another screen.
- Critical User Journeys (CUJs): the user’s goal and the tasks needed to reach it.
- A small pre-deployment smoke suite that checks the core application path.
Google recommends identifying and documenting Critical User Journeys, then testing those journeys end to end. Cypress also lists authentication, purchasing, cross-screen persistence, and pre-deployment smoke checks as common E2E scenarios. [Cypress testing types](https://docs.cypress.io/app/core-concepts/testing-types) · [Google: How Much Testing is Enough?](https://www.googblogs.com/how-much-testing-is-enough/)
2. Decide what belongs in the E2E layer
Start with user and business risk, not a target number of browser tests. Write down the outcomes users need, the cost of a failure, and the system boundaries involved. Select journeys where a break would block a critical task or where the behavior depends on several layers working together.
- List critical journeys. Describe each from the user’s point of view, such as “A customer can place an order and see it in order history.”
- Identify failure impact. Consider lost revenue, blocked access, incorrect persisted data, or a failed release.
- Map the system seams. Note which UI, service, database, and external integration behaviors must work together.
- Assign checks to the narrowest useful level. Put individual rules and edge cases in unit or component tests; use API or integration tests for service contracts; retain E2E for the complete journey and its critical seams.
- Make the test operable. Decide how data is created and cleaned up, what dependencies are controlled, and where the test will run in CI.
A browser test is usually a poor place to enumerate every business rule or input combination. That makes the suite slower, harder to diagnose, and more dependent on full-stack setup. Cover detailed logic with narrower tests, then use E2E checks to verify that the important pieces still connect correctly. This is a practical synthesis of the scope and dependency tradeoffs described by [Cypress](https://docs.cypress.io/app/core-concepts/testing-types) and [Google](https://www.googblogs.com/how-much-testing-is-enough/).
3. Balance E2E tests with faster test levels
The testing pyramid is a useful starting point: many unit tests, a smaller integration layer, and a limited set of E2E tests for critical flows and high-risk areas. It is not a universal quota. The UK Home Office notes that system complexity, safety requirements, prototype work, and resource constraints can change the right shape. Google also recommends building a solid unit base and integration coverage before adding CUJ checks.
| Test level | What it checks | Good fit | Typical tradeoff |
|---|---|---|---|
| Unit | One function or small unit of logic | Business rules, boundary conditions, transformations | Fast and focused, but does not establish that layers integrate |
| Component | A mounted UI component and its behavior | Component states, interactions, validation display | More UI coverage without starting the whole application |
| API or integration | HTTP contracts or interactions among a smaller group of real units | Service behavior, persistence contracts, integration seams | Often fewer dependencies than a full browser journey; does not prove the full UI works |
| E2E | A user-visible journey through the integrated application | Critical journeys, high-risk flows, release smoke checks | Broad confidence, but more infrastructure and less localized failures |
Do not treat a historic 70/20/10 unit/integration/E2E split as a measured rule for every team. The Google Testing Blog called it a “good first guess” and said the mix differs; the guidance in the reviewed sources emphasizes context and risk rather than a general ideal percentage. [UK Home Office test pyramid](https://engineering.homeoffice.gov.uk/standards/test-pyramid/) · [Google Testing Blog](https://testing.googleblog.com/2015/04/just-say-no-to-more-end-to-end-tests.html)
4. Build a reliable Playwright E2E test
The example below tests a typical order journey. It assumes the application is available at http://127.0.0.1:3000, exposes accessible labels and roles matching the example, and has a test-only API endpoint for preparing and deleting data. Replace those application-specific paths and labels with your own. The test creates a unique email for each run so parallel workers do not reuse the same account.
Install and configure
npm init playwright@latest
Choose TypeScript when prompted. Add or adapt playwright.config.ts:
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: Boolean(process.env.CI),
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 2 : undefined,
reporter: process.env.CI ? 'line' : 'list',
use: {
baseURL: process.env.BASE_URL ?? 'http://127.0.0.1:3000',
trace: 'retain-on-failure',
screenshot: 'only-on-failure',
video: 'retain-on-failure',
},
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
],
webServer: {
command: 'npm run start:test',
url: 'http://127.0.0.1:3000',
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
});
Set start:test to the command that starts your application against a test environment. If CI starts the app outside Playwright, remove webServer and set BASE_URL to that deployment. Use the browsers, projects, worker count, retries, and reporters appropriate to your CI capacity; more browser projects and workers increase resource use.
Write the journey
Create tests/checkout.spec.ts. The example’s test API is illustrative and must be implemented by the application; keep such endpoints restricted to a test environment.
import { test, expect } from '@playwright/test';
const baseURL = process.env.BASE_URL ?? 'http://127.0.0.1:3000';
test('customer can place an order and see it in order history', async ({ page, request }) => {
const email = `e2e-${Date.now()}-${Math.random().toString(16).slice(2)}@example.test`;
let customerId: string | undefined;
// Prepare isolated state using the application's test-only API.
const setup = await request.post(`${baseURL}/test-support/customers`, {
data: { email, password: 'Test-only-password-123!' },
});
expect(setup.ok()).toBeTruthy();
({ customerId } = await setup.json());
try {
await page.goto('/login');
await page.getByLabel('Email').fill(email);
await page.getByLabel('Password').fill('Test-only-password-123!');
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page.getByRole('heading', { name: 'Products' })).toBeVisible();
await page.getByRole('link', { name: 'Sample item' }).click();
await page.getByRole('button', { name: 'Add to cart' }).click();
await page.getByRole('link', { name: 'Cart' }).click();
await expect(page.getByText('Sample item')).toBeVisible();
await page.getByRole('button', { name: 'Checkout' }).click();
await page.getByRole('button', { name: 'Place order' }).click();
await expect(page.getByRole('heading', { name: 'Order confirmed' })).toBeVisible();
const orderNumber = await page.getByTestId('order-number').textContent();
expect(orderNumber).toBeTruthy();
await page.getByRole('link', { name: 'Order history' }).click();
await expect(page.getByText(orderNumber!)).toBeVisible();
} finally {
if (customerId) {
await request.delete(`${baseURL}/test-support/customers/${customerId}`);
}
}
});
The unique customer and cleanup illustrate isolation. Adapt cleanup to your system: deleting a customer may be inappropriate if records must be retained, and asynchronous cleanup jobs may be more suitable. If setup fails before an ID is returned, the application’s test support should avoid leaving partial state behind.
Why these assertions and selectors
getByRole and getByLabel describe the controls as a user encounters them. A stable test ID is used for the generated order number because it is a specific application contract. Avoid coupling a test to incidental CSS classes or internal function names. Playwright recommends asserting user-visible behavior and keeping tests independent; its web-first assertions wait and retry for the expected condition instead of checking once and racing the UI. [Playwright best practices](https://playwright.dev/docs/best-practices)
5. Test data, dependencies, and CI
- Give each test its own state. Use unique records, isolated browser contexts, and independent cookies and storage. Avoid relying on a previous test having logged in or populated a cart.
- Prepare state through a supported interface. An API or fixture can set up data faster and more directly than clicking through an unrelated admin form. Keep the journey under test in the browser.
- Control external dependencies. Use test credentials and sandbox integrations where available. Decide whether a third-party service is part of the behavior being validated or an unreliable dependency to isolate for this test.
- Make environment assumptions explicit. Document required services, secrets, migrations, seeded data, and startup commands. Do not point test cleanup at production data.
- Keep CI capacity in view. Browser processes, parallel workers, videos, traces, and multiple browser projects consume resources. Start with the minimum coverage and concurrency that meets release needs.
- Preserve failure evidence selectively. Traces and screenshots on failure can help diagnose intermittent behavior. Retain artifacts according to your team’s data-handling needs.
Cypress notes that full E2E tests require more backend infrastructure and scenario setup than narrower tests; its API testing guidance also describes using APIs to prepare state faster than driving forms. Google highlights that integration tests can often run with fewer dependencies in a smaller environment. [Cypress](https://docs.cypress.io/app/core-concepts/testing-types) · [Google](https://www.googblogs.com/how-much-testing-is-enough/)
6. Common reliability failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Test passes locally but fails in CI | Different startup timing, environment variables, browser resources, or shared data | Make service readiness explicit, set required environment values, isolate records, and inspect the failure trace before increasing timeouts. |
| Intermittent timeout after clicking | The test assumes an action completed immediately or waits on an unrelated timing guess | Assert on the resulting visible state with a web-first assertion; use a specific response or condition only when it represents the behavior being tested. |
| “Element not found” after a UI update | Selector reflects implementation details, or the element is not yet rendered | Prefer accessible roles and labels that match the user interface, then assert visibility. Add a test ID only for a stable contract without a suitable user-facing selector. |
| One test fails after another test ran | Shared account, cookies, local storage, or backend records | Give each test independent state and data; make setup and teardown safe for parallel runs. |
| Duplicate orders or records appear on retry | A retry repeats a non-idempotent action after a partial success | Use unique test data, clean up created state, and make test setup detectable. Do not blindly retry application operations in test code. |
| Checkout fails only when a payment provider is unavailable | A real external dependency is part of the test environment but is not controlled | Use the provider’s test environment for a deliberate integration journey; isolate it or cover the contract at a narrower level for other tests. |
| Test server does not start | Wrong start command, port, health URL, migration, or unavailable dependency | Run the command locally, check the configured readiness URL and logs, and verify required services and migrations are ready before browser tests begin. |
| Cleanup removes the wrong data | Test endpoint or credentials target a shared or non-test environment | Restrict test support endpoints and credentials to test environments, scope deletion to the unique record created by that test, and verify the configured base URL. |
7. Performance, reliability, and cost
E2E tests have operating costs: they need a browser, a running application, backend services, test data, and sometimes third-party integrations. Their execution and maintenance burden can grow with the number of full-stack scenarios. Keep detailed permutations in faster layers; run the critical journey suite where it provides release confidence, and keep setup and data cleanup predictable.
Reliability improves when tests are independent, assert observable conditions, and control the state they depend on. Retries can help a CI suite collect evidence about intermittent failures, but a test that passes only after retries still deserves investigation. Fixed sleeps usually make runs slower while leaving races unresolved; condition-based assertions wait for the condition that matters.
There is no source-backed universal test count, ideal E2E percentage, runtime target, or cost figure for all teams. Estimate your own suite’s cost from CI runtime, browser and service capacity, failure investigation, and maintenance effort. Google’s guidance frames how much testing is enough as dependent on the software’s type, purpose, and audience. [Google](https://www.googblogs.com/how-much-testing-is-enough/) · [UK Home Office](https://engineering.homeoffice.gov.uk/standards/test-pyramid/)
8. Capture a page for visual documentation
A screenshot can document what a page looked like during an investigation or provide an artifact alongside a test report. It cannot establish that a button works, a record persisted, or a user completed a journey; keep those claims in browser assertions and application checks. If you need a screenshot of a live page outside the Playwright test, ScreenshotNeo provides a website screenshot API and MCP server for developers. Its API options include full-page capture, element capture by CSS selector, device and viewport choices, wait conditions, custom CSS or JavaScript, and image formats including PNG, JPEG, and WebP. [ScreenshotNeo](https://screenshotneo.com) · [API documentation](https://screenshotneo.com/docs/)
Or skip the browser setup
Make one GET request to capture a page. Replace the target URL as needed and use your API key. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.
9. Frequently asked questions
How much testing is enough to qualify a software release?
There is no universal number. Base the decision on the software’s purpose and audience, the impact of failures, the coverage of important journeys, and the evidence from faster test layers. Google’s guidance recommends documenting and testing critical user journeys end to end.
Should every test run in a real browser?
No. Use a browser when the user-visible journey and integration across layers are what you need to verify. Use unit, component, API, or integration tests when they answer the question with less setup and a more focused failure.
Does a screenshot prove an E2E test passed?
No. A screenshot is visual evidence of a rendered page. A passing journey also needs assertions about actions and outcomes, such as confirmation that a created order appears in the user’s history.
Should E2E tests call real third-party services?
Only when the external integration itself is in scope and the test environment can support it reliably. Otherwise, isolate the dependency and test its contract at a narrower level, while retaining a deliberate journey for the integration risk that matters.


