How to Choose a Software Testing Strategy: The Testing Pyramid
Build a testing strategy around risk, feedback speed, and maintenance cost. Learn what belongs at each layer and when end-to-end tests earn their place.
A useful software testing strategy puts focused checks close to the behavior they protect, tests important interactions between components, and reserves end-to-end tests for a small set of critical user journeys and whole-system risks. Treat the testing pyramid as a way to reason about the balance—not a quota. Choose a test scope that exposes the risk with acceptable execution time, failure diagnosis, and maintenance cost.
There is no universal number of unit, integration, and end-to-end tests. Google’s 2015 Testing Blog article offers 70/20/10 as a “good first guess” and says the exact mix differs by team. That is a published heuristic, not a measured universal optimum or a claim about current Google-wide practice.
1. What the testing pyramid means
The pyramid describes a portfolio of automated tests at different scopes. The lower the layer, the fewer dependencies a test usually needs and the narrower the behavior it exercises. The names vary between teams, so define each test by what it actually runs and depends on.
| Layer | Typical scope | Good fit | Common trade-off |
|---|---|---|---|
| Unit | A function, module, class, or similarly focused behavior, often with external dependencies replaced by fakes or mocks. | Business rules, validation, formatting, edge cases, and error handling. | Fast and usually easy to localize, but a test with simulated dependencies cannot prove the real components work together. |
| Integration | A small group of components working together across a boundary, such as application-to-database or service-to-service. | Persistence mappings, serialization, queues, API contracts, configuration, and important component interfaces. | More realistic than isolated tests, while usually needing fewer dependencies than a complete deployed system. |
| End-to-end (E2E) | A larger system exercised through a user-visible interface or production-like entry point. | Critical user journeys, deployment wiring, and behavior whose correctness depends on the complete system. | High realism, but broader failures can take longer to run, diagnose, and maintain. |
Martin Fowler’s Test Pyramid explanation captures the central idea: emphasize low-level tests over broad, GUI-driven tests. It does not mean that every test called a “unit test” must target a single method, or that every end-to-end test is a browser test. Scope and dependencies matter more than labels.
2. Start with risks, not percentages
Before choosing a ratio or adding a test, write down what could fail, how users or systems would be affected, and where the failure can be detected most cheaply. A practical strategy begins with a small risk inventory and grows from real gaps.
- List important behaviors and boundaries. Include user goals, business rules, persistence, external services, permissions, and deployment configuration.
- Rank the consequences. Consider user impact, likelihood, detectability, and recovery cost. A rare payment or data-loss failure can deserve more coverage than a frequently used but harmless display detail.
- Choose the narrowest test that proves the behavior. Use focused tests for local rules, integration tests for real boundaries, and E2E checks where whole-system behavior is the risk.
- Identify critical user journeys. Select the short workflows whose failure would block a user’s main goal, such as account creation or completing a purchase.
- Run tests at useful feedback points. Developers need fast feedback during changes; broader checks can run at appropriate CI stages. Keep the schedule aligned with the cost and risk of a missed failure.
- Review failures and production feedback. A recurring incident may expose a missing boundary test. A slow, flaky test that duplicates narrower checks may need redesign or removal.
Google’s guidance on how much testing is enough similarly recommends a solid unit-test base, meaningful integration testing, and E2E verification of critical user journeys. The right level of rigor depends on the product’s purpose and risks.
3. Decide what belongs in each layer
Unit tests: behavior with a small boundary
Use unit tests for deterministic logic that benefits from quick, precise feedback. Examples include price calculations, permission decisions, input validation, state transitions, and handling of unusual values. Test observable behavior rather than implementation details that change without changing the contract.
- Cover normal cases and boundary values, including empty, missing, malformed, minimum, and maximum inputs where relevant.
- Make dependencies explicit. A mock can verify interactions; a fake can provide a lightweight working substitute. Avoid replacing the behavior that the test is meant to verify.
- Keep tests isolated from shared mutable state so they can run repeatedly and in any order.
- When a bug is found, add a focused regression test at the narrowest useful scope.
Integration tests: prove the seams work
Use integration tests where individually correct components can still disagree. Examples include database schema and query behavior, HTTP request and response contracts, message encoding, cache invalidation, and authentication middleware wired to application routes.
- Use real dependencies when their behavior is the risk; use a fake when it gives reliable, faster coverage of the same contract.
- Keep the environment small and controlled. Provision only the services and data needed for the boundary under test.
- Test failures as well as success: timeouts, rejected messages, invalid responses, transaction rollback, and retry behavior.
- Separate setup failures from product assertions so a broken environment does not masquerade as a regression.
End-to-end tests: protect complete outcomes
Use E2E tests when the claim depends on a working system path, not merely a collection of isolated parts. A short list of high-value journeys can check that users can reach their goal through the deployed application’s real entry points.
- Choose journeys by user impact and system-wide risk, not by the number of screens or features.
- Assert meaningful outcomes: a completed order, a persisted setting, or a generated report—not every decorative detail on the page.
- Keep test data controlled, isolate parallel runs, and make cleanup reliable.
- Prefer stable selectors and explicit synchronization over arbitrary sleeps. Wait for the condition that represents readiness.
- Use a browser screenshot as a visual artifact when visual appearance is part of the acceptance criteria. A screenshot alone does not prove that a workflow, backend operation, or business rule succeeded.
4. How many tests should each layer have?
Count is a weak target: one meaningful test can cover more risk than many redundant cases. Google’s 2015 70/20/10 split—70% unit, 20% integration, 10% E2E—is a starting point for discussion, not a universal standard. It is also about a distribution of tests, not a required distribution of runtime, test files, or engineering effort.
Instead of enforcing the ratio, ask whether every important behavior and boundary has suitable coverage, whether feedback arrives fast enough, and whether failures are actionable. A system whose main risk is component wiring may justify a larger integration layer. A product with a handful of crucial cross-system workflows may need more E2E coverage than a simple library. The architecture and risk profile determine the shape.
| Question | What a useful answer tells you |
|---|---|
| Which components and user-visible behavior does this test exercise? | Whether its scope matches the risk it is meant to cover. |
| How long does it take, and how often can it run? | Whether it provides feedback at the point developers need it. |
| Does it rely on unstable services, environments, or shared data? | Where flakiness and operational dependencies may come from. |
| Can a failure be localized quickly? | How much diagnosis and maintenance the test imposes. |
| Does it cover a meaningful failure mode not covered elsewhere? | Whether the test adds confidence or mostly duplicates existing checks. |
5. Diagnose the shape of your suite
Pyramid: a useful starting shape
A broad base of focused tests, a substantial layer for important component interactions, and a smaller set of E2E checks often gives teams a balance of feedback speed and whole-system confidence. The visual shape is a prompt for discussion, not a health score.
Ice-cream cone: too much weight at the top
A suite dominated by broad UI-driven tests can make feedback slow and failures difficult to diagnose. When an E2E test is the only protection for a simple business rule, consider adding a focused test closer to that behavior. Keep the E2E check only if it still verifies an important whole-system outcome.
Hourglass: a missing integration middle
An hourglass has many small unit tests and many broad E2E tests, but few tests of component interactions. It can leave teams choosing between a simulated test that misses real boundary defects and an expensive whole-system test. Google’s article on fixing a test hourglass describes building testable boundaries and using suitable fake backends to make integration checks faster and more reliable.
Other shapes can fit
Fowler discusses alternatives such as the honeycomb and trophy, which put more emphasis on integration or higher-level checks in some contexts. Google’s 2024 SMURF: Beyond the Test Pyramid likewise encourages considering trade-offs among test types. The useful question is not “Does our diagram look right?” but “Do our tests provide reliable, maintainable confidence at the boundaries that matter?”
6. Keep E2E coverage small and valuable
Build an explicit list of critical user journeys and give each a reason to exist. A journey usually merits E2E coverage when multiple parts of the system must work together and a narrower test cannot establish the user’s actual outcome.
- Write the user goal in one sentence.
- Identify the important system boundaries in the path.
- Check whether unit or integration tests already cover the likely failure modes.
- Add one E2E check for the outcome that depends on the whole path.
- Record its owner, required data, expected runtime, and what a failure should prompt.
- Review it when the workflow changes or it becomes flaky, redundant, or hard to diagnose.
Do not turn every possible input combination into an E2E scenario. Cover detailed permutations closer to the logic where they are cheaper to run and easier to debug. Keep broader tests for representative paths and high-impact failures.
7. Beyond the pyramid: other testing needs
The unit/integration/E2E layers mainly describe functional scope. They do not replace testing for other qualities. Depending on the product and risk, a strategy may also include performance and load testing, fault-tolerance checks, security testing, accessibility, localization, privacy, and usability research. Google’s testing strategy guidance lists these as additional concerns.
Plan these around specific risks. For example, a unit test can verify an algorithmic limit, while a load test checks behavior under realistic concurrency. An accessibility automation check can catch some markup issues, while it cannot by itself establish that an entire workflow is usable with assistive technology.
8. Use browser screenshots as visual evidence
For browser-based products, screenshots can help review visual regressions or attach a page artifact to a release check. Treat screenshot comparison as one signal in the testing strategy: dynamic content, fonts, animations, viewport dimensions, and browser differences can all affect pixels. Stabilize those inputs, compare only relevant regions when appropriate, and retain functional assertions for behavior that pixels cannot establish.
To capture a page yourself, use a browser automation framework already used by your team. Keep the page deterministic, set a fixed viewport, wait for a meaningful readiness condition, and save the image as a build artifact. Screenshot capture belongs alongside your tests; it does not replace tests at any layer.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. Here is a runnable cURL example for a page screenshot; see the API documentation for the available options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners, newsletter popups, and chat widgets can be removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, no card required.
9. Performance, reliability, and cost
Performance and feedback
Test runtime is part of the strategy because it affects how often developers can get useful feedback. Measure time by layer and identify the slowest suites, setup steps, and shared bottlenecks. Run independent tests in parallel only when their data and environments are isolated. Avoid optimizing runtime by removing the only test that covers an important risk.
Reliability
A flaky test makes it harder to distinguish a code regression from infrastructure noise. Track intermittent failures, preserve diagnostic logs and artifacts, and identify whether the cause is timing, external dependencies, test data, resource contention, or environment drift. Fix the cause or redesign the test; repeatedly rerunning without recording the failure pattern hides the reliability problem.
Cost and maintenance
Test cost includes compute and service usage, but also developer waiting time, maintenance, debugging, and the cost of defects that escape. Narrow tests tend to be cheaper to run and diagnose, while broader tests can buy confidence about real interactions. Compare the incremental confidence a test provides with its ongoing cost. Delete or consolidate tests only when their risk coverage is preserved elsewhere.
10. Troubleshooting common strategy problems
| Symptom | Likely cause | Practical fix |
|---|---|---|
| CI takes too long to give routine feedback. | Too many broad tests run for every change, slow setup, or serial execution. | Profile by suite, move detailed cases to narrower tests, reduce unnecessary setup, and schedule broad checks at a cadence that still matches risk. |
| E2E failures have vague causes. | The test covers too many steps or depends on unstable services and data. | Split the journey at meaningful boundaries, isolate test data, improve assertions and diagnostics, and cover local behavior in focused tests. |
| Unit tests pass but integration defects reach users. | Important real interfaces are mocked away or the middle layer is missing. | Add tests around the actual persistence, API, queue, or service boundary that failed. |
| Tests pass locally but fail in CI. | Environment mismatch, shared state, ordering assumptions, timing, or resource contention. | Make configuration explicit, isolate state, remove order dependence, wait on conditions, and capture CI diagnostics. |
| Coverage is high but regressions still escape. | Execution coverage is being treated as proof of meaningful assertions or risk coverage. | Review whether tests assert outcomes, cover critical behavior and boundaries, and detect realistic failure cases. |
| Developers ignore flaky failures. | Noise has reduced trust in the suite. | Assign ownership, track flakes separately, prioritize root-cause fixes, and avoid treating repeated reruns as a permanent solution. |
| Many tests duplicate the same scenario. | Coverage was added without reviewing existing tests or deciding which scope best exposes the risk. | Keep the clearest, most reliable test at each useful scope; remove duplication when equivalent risk coverage remains. |
11. A practical checklist for revising your strategy
- Have we listed the product’s most important risks and critical user journeys?
- Do focused tests cover business rules and important edge cases?
- Do integration tests exercise the boundaries where components can disagree?
- Does each E2E test protect a meaningful complete outcome?
- Can developers identify the likely cause of a failure without reproducing the entire system?
- Are test data, external services, and environments controlled well enough for repeatable results?
- Do runtime and flakiness trends have owners and follow-up actions?
- Do incidents and product changes feed back into the test portfolio?
FAQ
Is the testing pyramid still useful?
Yes, as a conversation tool for balancing scope, feedback, and maintenance. It is not a universal architecture or test-count rule; alternative shapes can better fit some systems.
Are acceptance tests always end-to-end tests?
No. Acceptance describes whether agreed behavior is satisfied; the test can run at different scopes. Fowler’s Practical Test Pyramid treats acceptance criteria as distinct from the pyramid’s test-level distribution.
Can a team have more integration tests than unit tests?
It can be reasonable when the system’s main risks lie in interactions and the integration tests remain fast and maintainable. Choose based on risk and observed feedback, not diagram symmetry.
Should every pull request run every test?
Not necessarily. Decide which checks must gate each change based on feedback time and failure impact, then ensure broader checks still run reliably before release or at another suitable point.
Does a screenshot prove the page works?
No. It records rendered appearance at a point in time. Use functional assertions to verify behavior, and use screenshots when visual output is itself relevant.


