ScreenshotNeo

BlogEngineering

How Agile Teams Use the Test Automation Pyramid

Use the test automation pyramid to balance fast, focused checks with service-level and end-to-end confidence—without treating a suggested ratio as a rule.

By the ScreenshotNeo team4 October 20269 min read

The test automation pyramid is a planning heuristic for balancing automated checks at different scopes. Agile teams usually build a broad base of focused unit or component tests, a substantial middle layer of integration or service tests, and a smaller set of end-to-end tests for critical user-visible paths. The right mix depends on the system and its risks; the pyramid is not a quota or a requirement to automate every check.

The reason to think in layers is practical: scope affects feedback speed, how clearly a failure points to a defect, and the effort required to keep tests reliable. A narrow check can make everyday development faster, while a smaller number of broad checks can provide confidence that important parts work together.

What the test automation pyramid means

The pyramid describes a portfolio, not a particular framework or test runner. Its classic shape has many low-level checks, fewer integration checks, and still fewer broad-stack checks that exercise an integrated application. Martin Fowler describes the model as resting on assumptions about the cost of broad-stack tests and notes that exceptions exist. Teams also use terms such as “unit” and “integration” differently, so agree on what each layer means in your own codebase.

Fowler attributes the model’s popularization as the “Test Automation Pyramid” to Mike Cohn’s 2009 book Succeeding with Agile. Fowler recounts that Cohn first drew the idea in conversation with Lisa Crispin around 2003–04 and described it at a Scrum gathering in 2004. Treat that as Fowler’s account of the history.

The three layers and what to put in them

Layer Typical scope Useful for Common tradeoff
Unit or component A small piece of behavior, often with dependencies isolated or controlled Frequent feedback on business rules, transformations, validation, and component behavior May not reveal integration failures at boundaries the test replaces or excludes
Integration or service A meaningful boundary, such as application-to-database or service-to-service communication Contracts, persistence, serialization, wiring, and behavior that depends on real collaborators Usually needs more setup and can be slower to diagnose than a focused unit check
End-to-end or broad-stack An integrated system exercised through a user-visible or system-level path Critical workflows and confidence that major components work together Can be slower, more brittle, and more expensive to maintain or diagnose

These labels do not map universally to tools, teams, or ownership. A “unit test” in one architecture might involve more collaborators than a “unit test” in another. Name layers by the boundaries and dependencies they actually exercise.

How Agile teams apply the model

  1. Start with a risk or behavior, not a target count. Identify what can fail, how users or other services are affected, and what evidence would catch the failure.
  2. Choose the narrowest reliable scope. Ask whether a focused test can check the behavior deterministically. Use a broader check when the risk lies in an interaction that the narrower test omits.
  3. Put boundary risks in the middle layer. For service-heavy or distributed systems, test important API, data, and service boundaries directly. These checks can catch interface and wiring problems without driving the whole UI.
  4. Keep broad checks for important integrated behavior. Choose a small set of high-value user journeys or system paths. Avoid making every rule prove itself through a browser journey if it can be checked reliably closer to the code.
  5. Use failures to adjust the portfolio. When a defect escapes, ask which layer could have caught it with useful feedback. When a broad test is consistently stable, fast, and cheap to change, it may cover enough risk that a corresponding lower-level check is unnecessary.
  6. Review the tradeoffs as the architecture changes. New services, persistence boundaries, or user workflows can change where confidence is most valuable.

For each proposed test, compare four things: feedback speed, stability and maintenance effort, diagnostic clarity, and the risk or boundary covered. This is more useful than counting tests by layer without considering what they protect.

Should teams use Google’s 70/20/10 split?

Google’s Testing Blog suggested 70% unit tests, 20% integration tests, and 10% end-to-end tests as a first guess in 2015, while explicitly saying the mix varies by team. This is a heuristic from that post, not an industry-wide measurement, proven optimum, or required target. A team can use it as a prompt to inspect an unusually top-heavy portfolio, but should decide based on its architecture, risks, and test costs.

For example, a service-heavy product may need a meaningful number of API and integration checks. A system with a small, stable set of critical user journeys may benefit from selected end-to-end checks. Neither case is automatically wrong because the resulting diagram does not look like a textbook pyramid.

When the portfolio should have a different shape

Use the pyramid as a question—“Can this behavior be checked reliably at a narrower scope?”—rather than an outcome to optimize for its own sake. Alternative discussions, including the honeycomb and trophy shapes, challenge whether every system should maximize unit tests. They represent alternative emphases, not settled replacements for the pyramid.

A broad test that is fast, stable, and inexpensive to change may reduce the need for a corresponding narrow test. Conversely, a test that crosses many moving parts and fails without clear diagnosis may be better split: retain a small system-level check for the integrated path, and cover detailed rules and boundaries lower down.

Practical example: choosing scope for a checkout change

Suppose a team changes discount calculation in an online checkout. A useful portfolio might include:

  • A focused component test for discount rules and edge cases such as an expired promotion or a minimum-spend threshold.
  • An integration or service test that verifies the checkout service reads the promotion data and returns the expected total through its API.
  • A small end-to-end check that a shopper can apply a valid promotion and see the updated total in the critical checkout journey.

The exact tests depend on the application and what its lower-level checks already cover. The point is to cover distinct risks at useful scopes, rather than repeat every assertion at every layer.

Browser checks and visual evidence

Some end-to-end checks need to verify what a user sees in a browser: a critical page renders, a key element appears, or a workflow reaches the expected state. Keep those checks focused on user-visible risks. A screenshot can help with visual review or provide an artifact for debugging, but by itself it does not prove that application logic, service contracts, or accessibility behavior is correct.

For a do-it-yourself screenshot in browser automation, the following Playwright example is runnable with Node.js after installing Playwright and its Chromium browser:

npm install --save-dev playwright
npx playwright install chromium
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  try {
    await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 30000 });
    await page.screenshot({ path: 'checkout.png', fullPage: true });
  } finally {
    await browser.close();
  }
})();

Replace the example URL with an environment you are authorized to access. For a production test, prefer a controlled test account and stable test data. Use a specific readiness condition when network activity never settles, and capture only the page or element needed for the check.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its screenshot API can be used to produce browser evidence without installing or managing a browser in this script. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month, with no card required.

Reliability, speed, and cost considerations

  • Feedback time: Run fast, focused checks frequently. Schedule or parallelize broader checks according to their runtime and the feedback your team needs.
  • Flakiness: Broad checks depend on more components and state. Control test data and environment, wait for meaningful application readiness, and avoid arbitrary sleeps when a condition can be observed directly.
  • Diagnosis: A failure that identifies a specific rule or boundary is easier to act on than a long workflow that fails somewhere near the end. Keep logs and artifacts that help identify where the failure occurred.
  • Maintenance: Every automated check has a change cost. Remove duplicate assertions and update tests when behavior changes; do not preserve a layer count at the expense of useful coverage.
  • Execution cost: Consider compute, browser infrastructure, setup, and developer time together. The sources support the general tradeoff that broad-stack checks can be slower and more costly to maintain; they do not establish a universal monetary cost or speed ratio.
  • Screenshot billing: If using ScreenshotNeo for screenshot artifacts, only clean shots are billed. Its response includes X-Page-Verdict and X-Billed headers; consult its documentation for the exact request and response options.

Troubleshooting test pyramid problems

Symptom Likely cause What to do
End-to-end suite is slow Too many checks repeat behavior that can be verified at a narrower boundary Keep critical integrated journeys and move detailed rules or contract assertions to component and service checks where reliable.
Browser tests fail intermittently Uncontrolled data, timing assumptions, network dependencies, or unstable environments Stabilize fixtures and environment, wait for a real readiness condition, and capture diagnostics around the failing step.
A test passes but integration still breaks The check mocked or bypassed the boundary that failed Add a test at the relevant service, database, or API boundary using the real integration needed to expose the defect.
One failure has many possible causes The test covers a broad workflow with several moving parts Retain broad coverage for the user journey, then add focused checks at the likely failure boundaries to improve diagnosis.
Team debates whether a check is “unit” or “integration” Test labels differ across teams and architectures Describe the actual scope, dependencies, and risk covered; agree on local definitions only when they help planning.
Screenshot request returns an error or unexpected page Invalid credentials or URL, an inaccessible target, a timeout, or a page that did not render usable content Check the request parameters and target access first. With ScreenshotNeo, inspect X-Page-Verdict and X-Billed in the response to distinguish a clean shot from a failed or unbillable capture.

Frequently asked questions

Does the pyramid tell a team how many tests to write?

No. It helps reason about scope and tradeoffs. Choose checks based on risk, architecture, feedback needs, and maintenance cost.

Are end-to-end tests bad?

No. They provide useful system-level confidence for important integrated behavior. The concern is relying on broad checks for every behavior when narrower checks can give faster, clearer feedback.

Does a screenshot test replace an end-to-end assertion?

No. A screenshot records rendered output. It does not establish that every functional requirement or backend boundary behaved correctly.

Is the trophy or honeycomb the new standard?

No universal replacement is established by the cited discussions. Treat each shape as a way to question portfolio emphasis, then select tests that cover your system’s risks.

What should a team do first?

Map a few important risks to the checks that cover them, note where those checks run, then look for slow, duplicate, or poorly diagnosed coverage. Adjust one boundary at a time.

Sources and further reading