How to Use the Testing Pyramid to Build a Faster Test Suite
Use the testing pyramid to balance fast feedback with confidence: choose the right scope for each check, then order your CI pipeline by speed and risk.
Use the testing pyramid as a guide for where to put checks, not as a required test-count formula. Put isolated behavior and edge cases in fast unit tests, interactions across important boundaries in integration tests, and a small set of critical user journeys in end-to-end tests. In CI, run quick, narrow, reliable checks first and broader checks later. The right mix depends on which layer can verify each risk cheaply and reliably.
The practical goal is a suite that gives useful feedback quickly while still checking the behavior and boundaries that matter. A larger number of tests is not automatically better: each check should add confidence that is worth its runtime, maintenance, and resource cost.
What the testing pyramid means
The testing pyramid is a metaphor for a portfolio of automated checks at different scopes. Its traditional shape has many small unit tests at the base, fewer integration or service tests in the middle, and a relatively small number of broad end-to-end tests at the top. Martin Fowler describes the core idea as having many more low-level unit tests than high-level, broad-stack tests. The shape reflects typical tradeoffs: broader tests often take longer to run and maintain, and can be harder to diagnose when they fail.
It is a heuristic, not a law. A higher-level test can belong earlier in a pipeline if it is fast, reliable, and cheap to change. Teams also use terms such as “unit” and “integration” differently. Agree on what dependencies and boundaries each label includes so the pyramid describes your suite consistently.
Choose the right scope for each behavior
Unit tests: isolated rules and edge cases
Use a unit test for a small piece of behavior whose result can be checked without bringing up the application’s full environment. Good examples include validation rules, calculations, state transitions, and boundary cases. Small, focused tests are generally faster and easier to diagnose. Keep them hermetic where practical so results do not depend on unrelated network, filesystem, clock, or shared-state conditions.
- Check normal inputs, boundary values, invalid inputs, and meaningful branches.
- Keep a failure local enough that the relevant behavior is obvious.
- Use test doubles for external dependencies when the test is about your code’s decision-making rather than that dependency’s real behavior.
- Keep fakes and stubs trustworthy; test them or use a real dependency when the interaction itself is the risk.
Integration tests: interactions and boundaries
Use integration tests to verify that a small group of units, services, or dependencies work together. They can catch issues that isolated tests miss, such as incorrect serialization, query behavior, configuration, or a mismatch between components. Keep the environment focused on the boundary being checked: a database-backed repository test need not launch every service in the product.
Avoid an hourglass suite that has many unit tests and broad end-to-end tests but too few checks of ordinary component interactions. Integration tests can provide more realistic confidence than isolated tests without the setup and dependency burden of a full end-to-end environment. Google’s testing guidance says that smaller environments can make integration tests faster and more reliable than full end-to-end tests.
End-to-end tests: critical user journeys
Use end-to-end (E2E) tests for a limited set of workflows where confidence depends on multiple real parts of the system working together. Choose critical user journeys (CUJs): actions that represent important user goals, such as signing in, completing a purchase, or creating a key resource. Keep the set focused on outcomes that lower-level tests cannot establish on their own.
A broad UI test is costly when it repeats every branch already covered below, but it can be valuable when it checks a real user journey or a boundary that narrower tests cannot reproduce. UI involvement alone does not determine a test’s level; scope and dependencies do.
Build or reshape the suite in seven steps
- List behaviors and risks. Group them into isolated rules, interactions across boundaries, and complete user workflows. Include the failures that would be costly to users or operators.
- Cover isolated behavior at the unit level. Add focused tests for important branches and edge cases. Prefer deterministic checks that do not need unrelated services.
- Add integration checks around important boundaries. Exercise the actual interactions that carry risk, such as application-to-database behavior or service contracts. Keep the environment only as large as needed.
- Select critical user journeys. Write E2E checks for a few essential outcomes, not every combination of lower-level behavior.
- Order CI by feedback value. Run the fast, narrow checks first, then progressively broader or slower checks. Place each test according to its measured runtime, reliability, scope, and value, not only its label.
- Turn discovered defects into focused regressions. When an E2E test finds a bug, add a lower-level regression test if that layer can reproduce it. Keep the E2E check if it still verifies a distinct journey or boundary.
- Review overlap and cost. Remove or redesign checks that repeat the same conditions without adding distinct confidence. Revisit the balance as architecture and risks change.
Order the tests in CI for faster feedback
A useful pipeline tends to report mistakes as soon as possible. Martin Fowler’s deployment-pipeline guidance puts this succinctly: “A good build pipeline tells you that you messed up as quick as possible.” A practical starting sequence is:
- Commit stage: formatting, static checks, and fast deterministic unit tests.
- Early integration stage: narrow tests for high-risk contracts or boundaries that are still fast and dependable.
- Broader verification: larger integration environments and slower suites.
- Critical journey stage: selected E2E checks, often with parallel execution where the test infrastructure supports it.
This is a starting arrangement, not a mandatory mapping. A fast integration test may be more useful in the early stage than a slow or flaky unit test. Track test duration and failure patterns so stage placement reflects actual feedback latency. A test that takes a long time to run or frequently fails for environmental reasons deserves attention at any layer.
How many unit, integration, and E2E tests should you have?
There is no universal target ratio. Google’s 2015 Testing Blog offered 70% unit, 20% integration, and 10% end-to-end as a first guess, while noting that the mix varies by team. Treat it as a prompt to inspect an E2E-heavy suite, not as a quota or an empirical guarantee of speed.
Test counts alone hide meaningful differences: one E2E check may cover a critical workflow, while many small unit checks cover separate edge cases. Instead, review the suite against these questions:
- Does each test cover a real behavior or risk?
- Does it add confidence that another test does not already provide?
- How long does it take to get a useful result?
- Is it reliable, and can someone diagnose its failure quickly?
- What environment, maintenance, and compute cost does it require?
- Does its scope reflect the production condition you need to verify?
Common pyramid shapes that cause trouble
| Pattern | What it looks like | Why it causes trouble | What to change |
|---|---|---|---|
| Ice-cream cone or inverted pyramid | Many broad UI or E2E tests and relatively few focused checks | Feedback can be slow, failures harder to localize, and maintenance heavier | Move independently reproducible rules and boundaries into smaller tests; retain E2E coverage for distinct critical journeys |
| Hourglass | A large unit layer and broad E2E tests with little integration coverage | Ordinary component interactions are left to an expensive full-stack environment | Add focused integration tests for the boundaries where the system most often fails or where confidence matters |
| Fixed-ratio pyramid | Teams add or remove tests to match a percentage target | Counts do not show reliability, runtime, risk, or unique confidence | Use ratios only as a discussion starter; assess the actual behaviors and costs |
| Duplicated pyramid | The same condition is asserted at every layer without a distinct purpose | Repeated checks add runtime and upkeep while obscuring which test owns the behavior | Keep lower-level regression coverage and retain higher-level coverage only when it verifies something distinct |
Measure suite health beyond test counts
Google’s 2024 SMURF framework describes five useful dimensions: Speed, Maintainability, Utilization, Reliability, and Fidelity. Apply them to the suite as a whole and to important individual checks:
- Speed: How long do developers wait for feedback, including queue time and setup?
- Maintainability: How much work does it take to update a test after a legitimate product change?
- Utilization: Does the test infrastructure spend resources efficiently, or do costly environments sit idle?
- Reliability: Does a test give consistent results for the same product state?
- Fidelity: Does it represent the real dependencies and operating conditions relevant to the risk?
These properties can trade off. A broad test may provide useful production-like fidelity, while a smaller test may be much cheaper and faster. A suite can also be functionally strong while missing nonfunctional needs: performance and load, fault tolerance, security, accessibility, localization, privacy, and usability require suitable coverage of their own.
Troubleshooting a slow or unreliable suite
CI takes too long even though most tests are unit tests
Likely cause: “Unit” tests may share expensive setup, perform I/O, or wait on external resources. Test labels do not guarantee speed or isolation.
Fix: Profile setup and execution time, remove unnecessary shared environment work, isolate independent tests, and move real dependency checks to focused integration tests where they are easier to understand.
E2E failures are hard to reproduce
Likely cause: The test depends on timing, shared state, external services, or an environment that differs between runs.
Fix: Make inputs and state explicit, reduce unrelated dependencies, and preserve E2E tests for the workflows that need broad coverage. When possible, reproduce the underlying defect with a smaller regression test.
The suite is fast but still misses integration defects
Likely cause: Isolated tests verify local logic but not the contracts or configuration between components.
Fix: Add integration coverage at the affected boundary, using the real dependency if its behavior is part of the risk. Keep the environment smaller than the full application when possible.
Teams argue over whether a test is unit or integration
Likely cause: The organization has no shared definition of scope or allowed dependencies.
Fix: Document what each level includes in your system—for example, whether a repository test may use a real database. Decide pipeline placement from runtime and confidence, even when labels remain imperfect.
Test counts rise but feedback does not improve
Likely cause: New tests duplicate existing conditions, test low-risk details, or are too slow or unreliable to help developers.
Fix: Review which distinct risk each test covers. Remove redundant assertions, improve flaky tests, or replace broad duplication with a focused check plus a small number of meaningful user journeys.
Limits of the testing pyramid
The pyramid cannot choose test scope without knowledge of the system’s risks. Teams with strong service boundaries, a complex client application, or unusually fast browser infrastructure may need a different balance. Some high-level checks are economical; some purportedly small tests are slow and fragile. The model also says little about nonfunctional testing. Use it to ask where a check belongs and what confidence it adds, then evaluate the evidence from your own pipeline.
Or skip the browser setup
For a visual regression or rendered-page check, the DIY pyramid still needs browser capture and test plumbing. If the job is simply to capture a website, ScreenshotNeo provides a website screenshot API and MCP server. The one-call API can return an image or PDF; the parameters other screenshot APIs use also work.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners and consent notices are accepted like a visitor and removed, along with supported newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month with no card.
FAQ
Does a test pyramid mean every project needs three named layers?
No. The labels are a useful way to discuss scope, but the valuable decision is which check can verify a risk reliably and at reasonable cost. Define the layers in terms of your system’s boundaries.
Should every E2E bug get another E2E test?
Not necessarily. Add a lower-level regression test when it can reproduce the defect. Keep the E2E test when it continues to verify a distinct user journey or system boundary.
Does the pyramid cover load or accessibility testing?
No. Those are separate testing needs. The pyramid helps distribute functional checks across scopes; it does not replace dedicated nonfunctional coverage.
Sources
- Martin Fowler, “The Practical Test Pyramid” — pyramid rationale, pipeline sequencing, and the value of quick feedback.
- Google Testing Blog, “Just Say No to More End-to-End Tests” — the 70/20/10 first-guess distribution and the qualification that the mix varies.
- Google Testing Blog, “How Much Testing is Enough?” — test scope, integration coverage, and critical user journeys.
- Google Testing Blog, “The SMURF Framework for Evaluating Test Suites” — Speed, Maintainability, Utilization, Reliability, and Fidelity.


