ScreenshotNeo

BlogEngineering

End-to-End Testing vs. Integration Testing: Key Differences

Integration tests check collaboration across a limited boundary; end-to-end tests check broader workflows. Learn how to choose, combine, and debug both.

By the ScreenshotNeo team4 October 202610 min read

Integration tests check whether a small group of components works together across a boundary, such as an application and a database. End-to-end (E2E) tests check whether a broader, integrated system can complete an important workflow. Use integration tests to cover boundary behavior with focused feedback, and reserve E2E tests for critical journeys whose success depends on multiple parts of the system.

The names are used inconsistently across teams. For every test, document what it exercises, which dependencies are real, and which are replaced with test doubles. Scope matters more than the label.

1. Key differences

Dimension Integration test End-to-end test
Scope A limited set of components or one integration point A broad, integrated workflow or system outcome
Question Do these components communicate and handle data correctly? Can the system achieve this user-facing goal?
Typical boundaries Database, API, queue, filesystem, serialization, or service client Multiple features and services composed into a critical user journey
Dependencies Usually a smaller environment; may use a real dependency or a test double More of the application and its dependencies are exercised
Feedback and diagnosis Often faster and easier to localize to a boundary Often slower, with more possible causes when it fails
Best use Boundary behavior and component collaboration A small set of critical user journeys and whole-workflow confidence

These are tendencies, not guarantees. A test called an integration test can cover a large system, and a well-designed E2E test can be stable and useful. Google describes an integration test as exercising a small group of units, often two; Martin Fowler describes a narrower approach that tests one integration point at a time. Both sources note that teams use the terms differently. See Google’s discussion of E2E test scope and Fowler’s practical test pyramid.

2. What an integration test checks

An integration test verifies behavior across a limited boundary. It can check that a repository writes and reads the expected database records, that an API client handles a service response, or that a queue consumer decodes and processes a message correctly.

For example, a test that starts a test database, calls the application’s repository, and verifies the stored record has a clear scope: it exercises the application-to-database boundary. State whether the database is a local real instance, an ephemeral test instance, or a substitute. If the test instead calls several services and completes a purchase, describe that broader scope even if your team calls it an integration test.

Use a focused integration test when a defect could arise from:

  • Serialization or parsing mismatches.
  • Database queries, constraints, migrations, or transaction behavior.
  • HTTP request construction, response handling, or authentication at a service boundary.
  • Queue message formats, delivery handling, or filesystem interaction.

Where practical, use a local dependency or test instance. Fowler cautions against automated tests that bombard a production service. Keep external side effects controlled and test data isolated.

3. What an end-to-end test checks

An E2E test checks a broad outcome across an integrated system, usually a critical user journey: for example, a user signs in, selects an item, pays, and sees an order confirmation. It is useful when success depends on multiple features or services working together.

End-to-end describes scope, not a required interface. A browser test is one common form, but an API-driven test can exercise a broad server-side workflow without a graphical UI. A UI test may also replace an external payment provider with a test double. Record which parts are real so the test’s confidence is clear. Fowler’s guide discusses this continuum of test scope and dependencies: The Practical Test Pyramid.

Choose a few user journeys whose failure would meaningfully affect users or business operations. A broad test can catch orchestration gaps that narrow tests cannot establish alone, but it should not be the only check for every boundary.

4. How to decide which layer to use

  1. Name the risk. Is the risk in one component boundary, or does it require several parts of the system to interact?
  2. Choose the narrowest useful test. If a database mapping or service response is the risk, test that boundary directly. If the risk is a complete user journey, add an E2E check.
  3. Make dependencies explicit. Note which services, databases, queues, and browsers are real, local, simulated, or replaced.
  4. Keep the broad suite purposeful. Cover critical workflows rather than repeating every lower-level assertion through a browser.
  5. Use failures to improve placement. If an E2E failure is consistently caused by one boundary, add a focused test there so future feedback is more direct.

A useful rule is: test the boundary at the integration layer; test the complete user goal at the E2E layer. They complement each other. Integration tests can find interface defects with less setup, while E2E tests provide confidence that important workflows still compose correctly.

5. Build a balanced test portfolio

Google’s 2015 testing guidance offers 70% unit, 20% integration, and 10% E2E as a “good first guess,” while explicitly noting that the right mix differs by team. Treat it as a starting heuristic, not a quota or research-proven optimum. Google’s later discussion retains the general pyramid idea but points out that growing suites require additional tradeoff thinking. Sources: Just Say No to More End-to-End Tests and SMURF: Beyond the Test Pyramid.

Watch for a test hourglass: many unit and E2E tests but few useful medium-scope integration tests. That shape can leave boundary behavior under-tested and make the broad suite carry too much diagnostic burden. Google discusses the hourglass and ways to address it in Fixing a Test Hourglass.

Track whether tests detect the risks they target, how long their feedback takes, and whether failures can be diagnosed from their output. Adjust the distribution to your architecture, failure history, and release needs rather than enforcing a fixed percentage.

6. A practical JavaScript example

The following runnable Node.js example uses only built-in modules. It demonstrates the distinction without requiring a framework: the integration test checks a database boundary (represented here by an in-memory adapter), and the E2E test checks a user journey through an application service. In a production repository, substitute the adapter with the real test database for a database integration test, and wire the journey to your actual application environment for an E2E test.

// Save as testing-layers.mjs and run with: node testing-layers.mjs
import assert from 'node:assert/strict';

// Application component and collaborator for the boundary test.
class MemoryOrderStore {
  #rows = new Map();
  async insert(order) { this.#rows.set(order.id, structuredClone(order)); }
  async find(id) { return structuredClone(this.#rows.get(id) ?? null); }
}

class OrderRepository {
  constructor(store) { this.store = store; }
  async save(order) {
    await this.store.insert({ id: order.id, totalCents: order.totalCents });
  }
  async get(id) { return this.store.find(id); }
}

// Integration scope: repository and storage adapter collaborate.
async function integrationTest() {
  const repository = new OrderRepository(new MemoryOrderStore());
  await repository.save({ id: 'order-42', totalCents: 2599 });
  assert.deepEqual(await repository.get('order-42'), {
    id: 'order-42', totalCents: 2599
  });
}

// Broader workflow components.
class CheckoutService {
  constructor(repository, payment) {
    this.repository = repository;
    this.payment = payment;
  }
  async checkout({ id, totalCents }) {
    await this.payment.charge(totalCents);
    await this.repository.save({ id, totalCents });
    return { status: 'confirmed', orderId: id };
  }
}

// End-to-end scope in this example: exercise the composed checkout workflow.
// The payment collaborator is a test double, so this is not a production-provider test.
async function endToEndTest() {
  const repository = new OrderRepository(new MemoryOrderStore());
  const charges = [];
  const payment = { async charge(cents) { charges.push(cents); } };
  const checkout = new CheckoutService(repository, payment);

  const result = await checkout.checkout({ id: 'order-43', totalCents: 4200 });
  assert.deepEqual(result, { status: 'confirmed', orderId: 'order-43' });
  assert.deepEqual(charges, [4200]);
  assert.deepEqual(await repository.get('order-43'), {
    id: 'order-43', totalCents: 4200
  });
}

await integrationTest();
await endToEndTest();
console.log('Both example checks passed.');

The example makes the scope visible, but its in-memory store does not test a real database protocol. To make it a database integration test, run the repository against an isolated test database and assert persisted behavior. To make the checkout check a true system E2E test, exercise the deployed or test application boundary and the dependencies relevant to the journey. Keep payment-provider sandbox behavior separate if that provider is not part of the test’s intended scope.

7. Browser-driven E2E tests and screenshots

For a browser workflow, a screenshot can help diagnose the page state at a failure point. It is evidence for debugging, not a replacement for assertions: assert the outcome that matters, such as a confirmation state or persisted order, and capture a screenshot when it helps explain a failure. Avoid making a screenshot comparison the only proof that a workflow succeeded.

When collecting screenshots, use the same environment and viewport where practical, wait for a meaningful page condition, and avoid relying on arbitrary delays where a selector or application-ready signal is available. Keep secrets and personal data out of saved artifacts.

8. Or skip the browser setup

If you need a screenshot artifact for a browser workflow or test investigation, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Cookie and consent banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

9. Troubleshooting test failures

Symptom Likely cause What to do
Integration test passes with a fake but fails with the real dependency The test double does not model a relevant protocol, constraint, or response Add a focused test against an isolated real dependency for that boundary; keep the fake for behavior the test double represents accurately.
Broad test fails intermittently Timing, shared state, test order, or an unstable dependency may affect the workflow Capture the failing state, isolate test data, wait for an observable condition, and inspect dependency health before adding retries.
E2E failure gives little diagnostic information The test crosses many components but reports only a final timeout or assertion Include step-level context, relevant logs, and a screenshot or trace when useful; add a narrower boundary test for the suspected interface.
Test suite slows down as coverage grows Too many tests repeat broad setup and workflow checks Keep E2E coverage for critical journeys; move boundary assertions to focused integration tests where they expose the same risk.
Test passes locally but fails in CI Environment configuration, dependency versions, state, or readiness differs Make required services and configuration explicit, reset test state, and wait for readiness rather than assuming startup timing.
Screenshot is blank or shows a transient state Capture occurred before the page reached the expected condition, or the target page did not load Wait for an application-specific ready selector or state and inspect the test’s navigation and network errors.

10. Performance, reliability, and cost

Integration tests commonly need fewer dependencies than full E2E tests, so they are often faster and easier to diagnose. Google’s 2021 guidance describes integration tests with smaller environments as faster and more reliable than full E2E tests with their full dependency set: How Much Testing is Enough? This is a general tendency, not a runtime guarantee for every suite.

Measure the feedback time and maintenance burden of your own suite. Parallel execution can shorten wall-clock time, but tests that share mutable data or scarce resources may become less reliable. Isolate state, make dependencies reproducible, and keep external services out of automated loops unless the test specifically covers that integration.

The cost of a test includes compute and environment setup as well as the engineering time spent diagnosing and maintaining it. A small number of broad tests can be valuable when they cover important orchestration risks; a large pile of overlapping broad checks can make feedback expensive. Place each assertion at the narrowest layer that still exercises the risk.

11. Frequently asked questions

Can an integration test use a real database?

Yes. Testing an application component against a real, isolated database is a common way to verify the application-to-database boundary. Be explicit about the database and test data lifecycle.

Do end-to-end tests have to use the UI?

No. A UI-driven browser test is common, but E2E refers to broad system scope. An API-level workflow can also exercise a broad integrated path.

Can one test count as both integration and end-to-end?

A test may cover several boundaries and a broad workflow. Since labels vary, describe the actual components, dependencies, and outcome it exercises instead of forcing a universal category.

Should every feature have an E2E test?

Usually, keep E2E checks focused on critical user journeys. Cover routine boundary behavior with narrower tests and add broad coverage where integrated behavior matters.

12. A checklist for naming and reviewing tests

  • Does the test state the behavior or user goal it verifies?
  • Can a reader tell which components and dependencies are exercised?
  • Is each real service isolated from production and from unrelated test data?
  • Could a narrower test expose the same boundary risk with clearer feedback?
  • Does each critical journey have enough broad coverage to catch orchestration failures?
  • Do test names and team documentation define local use of “integration” and “end-to-end”?

Use integration tests for focused confidence at component boundaries and E2E tests for selected workflows across the system. The best portfolio makes both the risk and the test’s scope easy to understand.