ScreenshotNeo

BlogEngineering

Complementary Testing: How to Combine Testing Approaches

Combine automated, risk-based, and exploratory testing to catch different kinds of defects. Use this practical guide to build and maintain a test portfolio.

By the ScreenshotNeo team4 October 20269 min read

Combine testing approaches by matching each one to the question it answers. Use narrow, repeatable automated checks for fast regression feedback; prioritize work according to risk; investigate uncertain behavior with exploratory testing; and reserve end-to-end checks for critical journeys and high-risk areas. The right balance depends on your product, failure risks, and quality objectives—there is no universal test count or coverage percentage.

These approaches complement one another because they have different strengths. A script can repeatedly verify specified behavior, but may miss an interaction nobody anticipated. Exploration can uncover surprising behavior, usability problems, and edge cases, but does not automatically create a repeatable regression check. Risk assessment helps decide which gaps deserve attention first.

What each testing approach contributes

Approach Useful for Tradeoff to manage
Automated scripted checks Repeatable verification of known behavior and regression feedback after changes They cover the scenarios encoded in them; upkeep is needed when behavior or interfaces change
Exploratory testing Investigating uncertain areas, unexpected behavior, usability, and edge cases Findings can be harder to reproduce unless the session and evidence are recorded
Risk-based testing Choosing what to test first and where to spend limited effort Priorities need revisiting when features, dependencies, or impact change
End-to-end testing Checking that a critical user journey works across connected parts of a system These checks can be complex and costly to maintain, so broad coverage is not automatically better

ISTQB’s Advanced Level Agile Tester syllabus describes the gap between scripted and real use: “Scripted end-to-end tests may miss unexpected behaviors that arise from real usage.” It also explains that exploratory testing can complement scripts by examining quality from multiple perspectives and helping uncover unexpected behavior, usability defects, and edge cases.

Build a test portfolio around risk

  1. Identify important risks. List critical user journeys, recently changed components, important integrations, complex rules, and failures with significant user or operational impact. Record why each item matters.
  2. Choose the narrowest useful check. Where feasible, use focused checks for isolated behavior. Add integration checks at boundaries where components, services, or data stores interact.
  3. Automate stable regression scenarios. Automate scenarios that are repeatable, valuable after changes, and maintainable. Keep the result useful to the people who need to act on failures.
  4. Explore what the scripts do not describe. Run focused exploratory sessions around uncertainty, newly changed behavior, realistic usage, and combinations that are difficult to enumerate in advance.
  5. Use end-to-end checks selectively. Cover critical journeys and high-risk behavior. Avoid making broad end-to-end coverage a substitute for narrower checks.
  6. Review and adjust. Revisit priorities when the product changes, new failure patterns emerge, or a check becomes costly or redundant.

This sequence is a practical synthesis of the cited guidance, not a prescribed formula. The UK Home Office test-pyramid guidance describes multiple levels from unit through end-to-end testing and recommends limiting end-to-end checks to critical flows and high-risk areas because of their complexity and maintenance costs.

Select test levels deliberately

Test levels describe the scope of a check, not whether it is automated or manual. A useful portfolio commonly includes narrow checks, checks across important integration boundaries, and a small set of broader user journeys.

  • Narrow checks: verify a small unit of behavior with few dependencies. They can provide focused feedback about logic.
  • Integration checks: verify that important components or services work together at their boundaries.
  • End-to-end checks: verify a user-visible flow across connected parts of the system. Choose these for journeys whose failure matters enough to justify their wider scope and upkeep.

Do not assume every feature needs the same mix. A high-risk payment flow may justify deeper integration and end-to-end attention than a low-impact informational screen. A new, uncertain interaction may need exploration before the team knows which scripted cases are worth keeping.

Make exploratory testing repeatable enough to learn from

Exploratory testing is purposeful investigation, not unstructured clicking. Give a session a target and a time boundary, then record enough context for another person to understand what was covered and reproduce a finding.

  1. State the area, change, or risk you are investigating.
  2. Note relevant assumptions, data, environment, and starting conditions.
  3. Try realistic paths and vary inputs, timing, state, and sequences where relevant.
  4. Capture unexpected results with steps, evidence, and the expected behavior if it is known.
  5. After the session, decide whether a finding needs a defect report, a new automated regression check, a design change, or more investigation.

Exploration and automation can feed each other. A session may reveal a repeatable defect worth turning into a regression check. Conversely, an automated failure can point to an area that needs human investigation because the cause or user impact is unclear.

Prioritize with a practical risk record

For each candidate area, write down the failure mode, its likely impact, how much uncertainty remains, and what evidence would reduce that uncertainty. Use the record to explain priorities rather than to claim precise risk mathematics unsupported by evidence.

Risk question Example evidence to consider Possible testing response
Who or what is affected if it fails? Critical user journey, data integrity, security-sensitive behavior, or operational consequence Prioritize focused checks and appropriate boundary or journey coverage
What changed? Code change, dependency change, configuration, migration, or interface change Run relevant regression checks and explore changed interactions
Where are assumptions uncertain? New behavior, complex combinations, unclear requirements, or unfamiliar browser behavior Use exploratory sessions to learn what scenarios matter
How costly is a missed defect? Impact and recovery difficulty, based on product context Invest more evidence-gathering effort where the consequences justify it

Use browser screenshots as supporting evidence

Visual evidence can help a reviewer understand a rendering defect or compare a page before and after a change. A screenshot is an observation of a page at a particular state and time; it does not prove that the underlying behavior, accessibility, or workflow is correct. Pair visual checks with the relevant functional and exploratory work.

Capture a page with a browser you control

For a local page or a test environment that requires your own browser session, Playwright can navigate to a URL and save a screenshot. Install Playwright and its Chromium browser in your project first:

npm install -D playwright
npx playwright install chromium

Save this as capture.mjs, then run node capture.mjs https://example.com. The output is a full-page PNG. Use a test URL you are authorized to access.

import { chromium } from 'playwright';

const url = process.argv[2];
if (!url) throw new Error('Usage: node capture.mjs <url>');

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

domcontentloaded avoids waiting for every network request to finish, which can be useful on pages with long-lived requests. If the content you need appears later, wait for a meaningful selector or application-ready signal before capturing. Set viewport, device scale, authentication, and page state to match the question the evidence should answer. Avoid treating a single screenshot as a complete visual regression system.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; the API documentation describes the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners are accepted and removed before capture; 60+ known consent platforms, newsletter popups, and chat widgets can be removed, and each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers identify the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.

Sign up free for 1,000 screenshots a month, with no card.

Keep the portfolio reliable and affordable

Performance and feedback

Run fast, focused checks frequently enough to inform changes while they are still easy to investigate. Broader checks can take more coordination and are more exposed to interactions across components. Organize execution so failures give a useful signal: identify the affected area, preserve relevant logs or artifacts, and make it clear which check failed.

Reliability and maintenance

  • Keep checks independent where practical, with controlled data and explicit setup and cleanup.
  • Use stable assertions tied to meaningful behavior. Avoid relying on incidental timing or fragile presentation details unless those details are the subject of the test.
  • When a check fails, distinguish a product defect from an environment issue, bad test data, or an unstable test. Record enough evidence to make that diagnosis.
  • Review checks that repeatedly fail without useful findings. Repair, replace, or remove them so the suite remains actionable.
  • Retain exploratory session notes and screenshots when they make a finding easier to reproduce, while following your team’s data-handling rules.

Cost and effort

Testing effort includes writing and maintaining checks, running them, investigating failures, preparing environments, and performing exploratory work. Compare approaches by their purpose, scope, feedback speed, repeatability, upkeep, and ability to uncover behavior outside predefined cases. The research cited here provides no numerical benchmark for test counts, coverage targets, defect reduction, or return on investment, so choose priorities from your product’s observed risks and quality goals.

Troubleshooting common portfolio problems

Symptom Likely cause What to do
Many checks pass, but users still find surprising failures The suite mostly verifies known scripted scenarios Explore changed and uncertain areas; use findings to update requirements, design, or regression checks
End-to-end suite is slow or hard to maintain Too many checks span the full system or depend on unstable setup Keep end-to-end coverage focused on critical flows; move suitable assertions to narrower levels
Exploratory findings are difficult to reproduce Session target, starting state, inputs, or evidence were not recorded Record conditions and steps as you explore, then add a repeatable check when the behavior warrants it
Automated check fails intermittently Timing, shared state, environment variability, or test data may be uncontrolled Inspect logs and state, make setup deterministic, wait on meaningful conditions, and isolate shared resources
Visual captures differ between runs Viewport, device scale, content, fonts, animation, or page readiness differs Control capture conditions, wait for the relevant state, and determine whether the difference is meaningful before changing an assertion
Testing effort is spread evenly across all features Priorities are not linked to impact, uncertainty, and change Review critical journeys, recent changes, complex boundaries, and consequences of failure; allocate effort accordingly

A review checklist

  • Have we identified critical journeys and meaningful failure risks?
  • Do narrow and integration checks cover suitable behavior and boundaries?
  • Are stable, valuable regression scenarios automated and actionable?
  • Do we investigate uncertain behavior and realistic usage with exploration?
  • Are end-to-end checks limited to flows whose risk justifies their maintenance?
  • Can someone reproduce a failure from the evidence we keep?
  • Do we revisit priorities as the product and failure patterns change?

FAQ

Is exploratory testing the same as testing without a plan?

No. Give an exploratory session a target and record its conditions and findings; the investigation remains flexible while still producing useful evidence.

Should every exploratory finding become an automated test?

No. Automate findings that represent valuable, repeatable regression scenarios. Some findings instead call for a product change, clearer requirements, or further investigation.

Is a screenshot enough to verify a page?

No. It documents visual output at one moment. Use it alongside checks for the behavior and user journey you need to verify.

How many end-to-end tests should a team have?

The cited guidance supports selecting them for critical flows and high-risk areas, not a universal count. Decide from your product’s risks and the checks’ maintenance cost.

Sources