How to Choose the Right Layers for a Test Automation Strategy
Choose test layers by the risks they cover: fast focused checks for local behavior, boundary tests for interactions, and selective end-to-end checks for critical journeys.
Choose test scope according to the failure risk and the confidence you need. Use focused checks for local behavior, integration or service tests for interactions across boundaries, and a small set of end-to-end tests for critical journeys that depend on the assembled system. Treat the test pyramid as a starting heuristic, not a required ratio.
The right portfolio is the one that catches meaningful failures quickly and reliably, is maintainable, and reflects the parts of production behavior that matter. Test names alone do not tell you what a test covers: write down which processes, dependencies, and interfaces each layer exercises.
1. Start with failure boundaries
For each important user outcome or technical risk, ask: What is the narrowest test scope that can detect this failure reliably? Place the check at that boundary. This keeps local rules quick to diagnose while reserving broad tests for failures that only appear when components work together.
| Layer | Use it for | What it tells you | Costs and cautions |
|---|---|---|---|
| Unit or component | Local logic and behavior within an isolated function or small component boundary | Whether a focused rule or component behaves as expected | Usually quick and resource-light. Excessive simulation or narrow isolation can drift from integrated behavior. Teams use “unit” and “component” differently. |
| Integration, service, or API | Contracts and interactions between components, services, databases, or dependencies | Whether the relevant boundary works with connected parts | Requires more setup and resources than focused checks. Be explicit about which dependencies are real, simulated, or omitted. |
| End-to-end or UI | A few critical user journeys or risks that depend on the assembled system and user-facing path | Whether the whole configured path works together | Typically slower and more exposed to timing, browser, environment, and dependency instability. Keep scenarios valuable and maintainable. |
| Exploratory or manual | Usability and unexpected concerns that are difficult to express as repeatable assertions | What a human can discover by probing the product in context | Does not itself create a repeatable regression check. Turn important discoveries into an automated check when a stable assertion makes sense. |
UI and end-to-end are not synonyms: a UI test can be narrowly scoped, and an end-to-end check can exercise a system without using a graphical interface. “Customer-facing” describes an audience or environment, not necessarily a test scope.
2. Use the pyramid as a heuristic, not a quota
A pyramid is a useful picture of a suite with many focused checks, fewer boundary checks, and still fewer broad checks. Its underlying rationale is that broader tests are often slower, more costly to maintain, and less reliable. That pattern is not universal: if broad tests in your system are fast, stable, and cheap to change, you may need fewer narrow tests than the picture implies.
Google’s 2015 article offers a 70% unit, 20% integration, and 10% end-to-end split as a “good first guess,” while also saying the mix differs by team. It is guidance, not an empirical universal optimum or a target every suite must hit. The category definitions vary, so comparing percentages across teams can be misleading. Google’s suggested starting split is best treated as a discussion prompt.
When choosing between plausible layers for the same behavior, compare:
- Failure boundary: What components and dependencies must participate for the defect to appear?
- Feedback speed: How soon does the result reach the person who can act on it?
- Reliability: Does the check fail for product defects, or frequently for timing and environment noise?
- Fidelity: How closely does the exercised path reflect the real configuration and production behavior?
- Cost: What are the runtime, infrastructure, debugging, and maintenance costs?
- Added confidence: Does the broader check catch a meaningful risk that narrower checks do not?
Google’s SMURF framework names five useful suite qualities: speed, maintainability, utilization, reliability, and fidelity. These qualities can pull against one another. For example, using real dependencies may improve fidelity while increasing runtime and setup work. Consider the tradeoff rather than optimizing only for the number of tests. See SMURF: Beyond the Test Pyramid.
3. Build a portfolio from risks
- List important outcomes and boundaries. Include the user actions that protect core product value, the components involved, external dependencies, and known high-risk failure modes.
- Assign each risk to its narrowest reliable scope. Put local rules in focused checks, interactions at the relevant service or component boundary, and whole-system risks in end-to-end scenarios.
- Define your layer vocabulary. Document what “unit,” “integration,” “service,” and “end-to-end” mean in your team. State which interfaces, processes, and dependencies a test exercises. The terms are used inconsistently across the industry.
- Choose a small set of high-value journeys. Add end-to-end coverage when a critical path depends on the assembled system and lower layers cannot provide the needed confidence. Do not copy every lower-layer edge case into a browser test.
- Order checks for useful feedback. Run quick, narrowly scoped checks early. Put slower broad checks later when doing so gives developers actionable feedback sooner. Test labels alone should not decide pipeline order.
- Review failures and escaped defects. When a broad test finds a defect, add a narrower regression check where practical. Keep the broad check if it protects a distinct whole-system risk.
- Keep exploratory testing in the plan. Use human-directed exploration for usability and unexpected behavior, then decide whether each finding needs a repeatable automated assertion.
Martin Fowler’s Practical Test Pyramid discusses choosing layers, sequencing feedback, avoiding duplication, and keeping exploratory testing in view. The UK Home Office’s test pyramid guidance advises that its teams avoid large numbers of end-to-end tests; that is a department standard, not a universal rule for every organization.
4. Keep end-to-end automation selective
End-to-end tests earn their place when the product risk crosses boundaries and confidence depends on the actual assembled path. A few examples might cover a core sign-up, checkout, or publishing flow, if those are critical outcomes in your application. Pick scenarios based on your own risks rather than adopting a generic list.
For each proposed end-to-end test, write down:
- Which user or business outcome it protects.
- Which boundaries and dependencies it traverses.
- What defect it can detect that lower-level checks cannot.
- How it will be made deterministic and diagnosed when it fails.
- Whether a focused regression check can cover the defect after discovery.
Keep the test only while it adds meaningful confidence. If it repeats assertions already covered below and does not validate a distinct assembled-system risk, simplify or remove the duplicate. Broader scope is useful when it adds evidence, not simply because it looks closer to a user action.
5. Sequence tests for fast feedback
A practical pipeline often starts with checks that run quickly and have a small failure surface, then runs service or integration checks, and finally runs broader journey tests. The exact stages depend on runtime, parallelization, and how quickly failures need to reach developers. A test that is reliable and cheap may run earlier than its label suggests.
Separate the questions your pipeline answers. A fast change-level stage can catch regressions before merge; broader scheduled or pre-release stages can cover slower scenarios when they are not needed on every local edit. If a risk requires an end-to-end check before release, schedule it at a point where its result can still influence the release decision.
Do not hide flaky tests in a permanently ignored stage. Record intermittent failures, identify whether the cause is product behavior or test/environment instability, and fix or retire checks that repeatedly fail for reasons unrelated to useful assertions.
6. Measure suite health
There are no universal target values for the following measures. Track them over time and use them to find costly or missing coverage:
| Measure | What to inspect | Question it helps answer |
|---|---|---|
| Execution time | Per test, stage, and full suite; include queue and setup time where available | Where does feedback become slow? |
| Unreliable tests | Intermittent failures, reruns, and failures that disappear without a product change | Which checks consume attention without dependable evidence? |
| Defects found or missed by level | Where defects are detected and where they escaped to later stages or users | Are important risks missing a suitable check? |
| Defect density | Defects associated with components or boundaries over time | Where might additional focused or boundary coverage help? |
| Automation coverage | Important outcomes or risks with a repeatable check, alongside uncovered areas | Are checks aligned with risk, or are counts masking gaps? |
| Maintenance and utilization | Time spent fixing tests, and whether results are used in decisions | Does the suite provide enough value to justify its ongoing cost? |
Pair metrics with review of real failures. A high test count or coverage percentage cannot show by itself whether assertions are clear, reliable, or attached to the right failure boundary.
7. Troubleshoot common strategy problems
| Symptom | Likely cause | What to do |
|---|---|---|
| The suite is slow and developers avoid running it | Too much work is concentrated in broad tests, or the feedback sequence delays quick results | Measure runtime by stage, move focused checks earlier, parallelize where useful, and retain broad checks only for risks they uniquely cover. |
| UI tests fail intermittently | Timing sensitivity, unstable test data, browser/environment variation, or dependencies outside the assertion’s control | Inspect failure evidence, stabilize data and setup, wait on meaningful conditions instead of arbitrary timing where possible, and remove checks that cannot be made dependable. |
| Unit tests pass but integrated behavior breaks | Mocks or narrow isolation do not represent a failing contract or dependency interaction | Add a focused integration or service test at the boundary that failed. Keep local checks for local rules. |
| Teams disagree about whether a test is integration or end-to-end | Layer names conceal different assumptions about process, interface, and dependency scope | Describe what runs and connects in concrete terms, then use the shared definition in pipeline and reporting. |
| The same assertion appears at several layers | Coverage was added by copying tests without identifying distinct evidence | Keep duplication only where each level protects a different failure boundary or gives necessary diagnostic confidence. |
| A defect is found only by a broad test | The failure boundary was not represented below, or the issue truly requires the assembled system | Add a narrower regression test if it can reproduce the defect reliably, and retain the broad scenario if it still validates a critical path. |
| Coverage metrics rise but escaped defects do not fall | Counts measure test volume rather than risk coverage or assertion quality | Review missed defects, critical outcomes, test reliability, and whether the tested scope matches actual failure boundaries. |
8. ScreenshotNeo for screenshot checks
Visual checks can be one part of a test strategy when a rendered page or component is itself the behavior you need to inspect. They do not replace unit, service, or end-to-end assertions about application logic. To keep a visual check useful, decide which route or component is important, what state it must be in, and what differences should count as a meaningful failure.
For a do-it-yourself browser screenshot check, use a browser automation framework your team already maintains. Here is a runnable Playwright example in Node.js that captures a page after the network becomes idle. Install Playwright with npm install -D playwright and install its browser with npx playwright install chromium.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 30000 });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
})();
For this approach, pin the browser and framework versions in your project, isolate test data, use a stable test environment, and avoid depending on third-party content that changes independently. Set timeouts deliberately: network-idle can be unsuitable for pages with persistent connections or polling. In those cases, wait for a specific selector or application-ready state. Browser screenshots can vary with viewport, device scale, fonts, animation, locale, and time-dependent content; fix those inputs when comparing captures. Browser setup and maintenance are part of the test’s cost.
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. A single GET request returns a PNG, JPEG, WebP, or PDF. It can help when a test or workflow needs a captured page without maintaining a browser process. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Keep the API key out of source control and client-side code. For production use, handle non-success responses and timeouts, and inspect the response headers: X-Page-Verdict and X-Billed report the page outcome and whether the capture was billed. Cookie banners are accepted and 60+ known consent platforms, newsletter popups, and chat widgets can be removed before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
There are 1,000 screenshots per month on the free plan with no card. Paid plans start at $5 for 3,000 screenshots; all features are on every plan, and yearly billing gives two months free. If this fits your visual-check workflow, sign up for 1,000 free screenshots a month with no card. Learn more about ScreenshotNeo.
10. FAQ
Should every important feature have an end-to-end test?
No. Give each important risk a reliable check at the narrowest scope that can detect it. Use end-to-end coverage when whole-system behavior adds needed confidence.
Is the 70/20/10 split a proven optimum?
No. Google presented it as a starting guess and said the right mix differs by team. Treat it as a prompt for discussion, not a quota.
Can a team use a test shape other than a pyramid?
Yes. The useful question is what each test exercises and whether the suite is fast, maintainable, reliable, used, and faithful enough for its purpose.
Does automated testing remove the need for exploratory testing?
No. Exploratory work can reveal usability and unexpected issues that are difficult to encode as repeatable checks. Automate findings when stable regression assertions make sense.


