ScreenshotNeo

BlogEngineering

Testing Best Practices: Dos and Don’ts for QA Teams

Build a risk-based testing strategy that combines the right test levels, techniques, and quality checks—and makes release evidence clear.

By the ScreenshotNeo team4 October 202610 min read

Good testing starts with the failures that would matter most to users. Identify important workflows, likely failure modes, affected users, and the cost of failure; then choose test levels and techniques that produce useful evidence about those risks. There is no universal test count or test-level ratio that proves a release is ready.

Exhaustive testing is impractical, so teams sample and prioritize. A passing suite says that the tests run passed under their conditions. It does not prove the absence of defects. Release decisions should state what was tested, what remains uncertain, and why the remaining risk is acceptable.

1. Start with risk and user impact

Before deciding what to automate or how much coverage to target, make the product risks explicit. For each critical area, record:

  • User or workflow: Who depends on it, and what are they trying to accomplish?
  • Failure mode: What could go wrong, including invalid data, dependency failure, or an unsafe state?
  • Impact and likelihood: How many users or systems could be affected, and how serious would the outcome be?
  • Evidence needed: What result would help the team release, investigate, or remediate?

Use this assessment to prioritize test design and execution. A payment authorization, account recovery flow, or data migration may warrant more varied evidence than a low-impact display preference. The right priority depends on your users, product, and consequences of failure.

ISO/IEC/IEEE 29119-1:2022 describes risk-based testing as a basis for test strategy and prioritization. Its series overview also covers test levels and types, techniques, documentation, environments, data, reporting, and defect management. See the ISO/IEC/IEEE 29119-1:2022 overview and the ISO software testing standards catalogue. Static reviews are covered by ISO/IEC/IEEE 20246, according to the series overview.

2. Use test levels for different questions

Unit, integration, and end-to-end tests provide different kinds of evidence. Build a mix that fits the architecture and risk; do not force every check into a full application workflow.

Level Question it helps answer Good fit Trade-off
Unit or component Does an individual function or component behave as intended? Branches, validation rules, calculations, boundary conditions Fast feedback, but collaborators and external systems may be represented by test doubles.
Integration Do connected components exchange and handle information correctly? Service contracts, persistence, queues, adapters, API boundaries Exercises real connections while limiting some of the dependencies of a full user journey.
End-to-end (E2E) Can a user complete a critical workflow through the assembled product? A small set of high-value journeys and release-critical paths Provides realistic workflow evidence, but broad dependencies can make these checks slower and less reliable.

Google’s testing guidance recommends smaller tests and integration checks alongside end-to-end tests for important user journeys. Integration tests typically have fewer dependencies than full end-to-end tests, which can make them faster and more reliable. Read Google’s guidance on end-to-end testing.

Google’s earlier testing-pyramid article offered 70/20/10 unit/integration/end-to-end as a first guess, while saying the mix varies by team. Treat it as a dated heuristic, not a required target or independently validated benchmark. Choose a distribution based on the feedback speed, realism, and maintenance cost your risks require.

3. Choose a technique that fits the behavior

Test design techniques help turn requirements and risks into cases. Select them based on the behavior or uncertainty you need to examine.

  • Boundary-value analysis: Test at, just below, and just above limits such as minimum age, maximum upload size, or retry count.
  • Equivalence partitioning: Group inputs expected to behave alike, then sample representative values from each group.
  • Decision tables: Enumerate combinations of conditions and expected actions, useful for permissions, pricing rules, or eligibility.
  • Use-case testing: Check behavior through a user goal and its alternate or failure paths.
  • Exploratory testing: Learn about the product while testing, guided by a charter or identified risk; record the discoveries and evidence.
  • Checklist-based testing and error guessing: Apply relevant team knowledge to recurring omissions and plausible failure cases.

These techniques are documented in ISO/IEC/IEEE 29119-4. ISTQB’s historical 2017–18 survey of more than 2,000 respondents in 92 countries reported use-case testing, exploratory testing, boundary-value analysis, checklist-based testing, and error guessing among common techniques. That survey describes its period and does not establish current prevalence. See the ISTQB survey summary.

4. Cover quality attributes beyond functionality

Functional behavior is only one part of product quality. Select non-functional checks where the product’s risks justify them, and begin early enough to address findings before release pressure makes them costly.

  • Performance and load: Measure response behavior under expected and peak demand.
  • Fault tolerance: Check recovery and graceful behavior when dependencies fail or become slow.
  • Security and privacy: Examine access boundaries, sensitive data handling, and relevant threat scenarios.
  • Accessibility and usability: Check that users can perceive, navigate, and complete important tasks.
  • Localization and globalization: Check language, date, number, currency, and layout behavior for supported locales.

Google’s guidance lists performance, load and scalability, fault tolerance, security, accessibility, privacy, usability, localization, and globalization among areas teams may need to test. The applicable set depends on the product’s purpose and audience. See Google’s testing education guidance.

5. Keep environments, data, and checks dependable

Test results are useful only when their conditions are understood. Manage test environments and data deliberately, and treat flaky or unexplained failures as reliability work.

  1. Document environment dependencies, configuration, and meaningful differences from production.
  2. Use representative test data, with appropriate controls for sensitive information.
  3. Make setup and cleanup repeatable so test outcomes do not depend on execution order.
  4. When a test fails, capture enough context to reproduce it: inputs, environment, relevant logs, and dependency state.
  5. Retest fixes and run regression checks selected for the affected risks.

ISO/IEC/IEEE 29119 includes environment and data management, communication and reporting, and defect or incident management as supporting testing activities. A flaky test can obscure a real regression and consume investigation time; identify whether the instability comes from timing, shared state, external dependencies, or an unclear assertion before relying on its results.

6. Treat coverage as one signal, not a release score

Code coverage indicates which code structures tests exercised under a particular run. It does not establish that requirements were correct, important user journeys work, or relevant quality attributes are safe. Pair coverage with evidence tied to explicit risks and test objectives.

For a release review, report:

  • the critical workflows and risks assessed;
  • test levels, techniques, and environments used;
  • important failures, fixes, and unresolved defects;
  • material gaps, including untested platforms, data states, or dependencies;
  • the remaining risk and the reasoning behind the release decision.

This makes the limits of the evidence visible. ISO/IEC/IEEE 29119-1:2022 is an informative introduction to the series; distinguish that overview from normative requirements in other parts of the standards.

7. Adjust the strategy to the product and lifecycle

Testing strategy should change when the product, architecture, users, or consequences change. A service with financial or safety consequences may need stronger controls and more evidence than a low-impact internal tool. A system with many external dependencies may need contract and failure-mode checks alongside a small number of complete journeys. A rapidly changing feature may benefit from exploratory sessions early, then repeatable regression checks as behavior stabilizes.

Ask at planning and release points: What changed? Which users or risks are newly affected? Which earlier assumptions no longer hold? What evidence would change the release decision? These questions keep the test effort connected to decisions instead of accumulating checks without a clear purpose.

8. Testing AI-based systems

AI-based systems can make expected results difficult to define. Some outputs are probabilistic or context-sensitive, so a single exact expected string may not be an adequate oracle. Define acceptance criteria around the task and risks, explain how outputs will be evaluated, and record where human judgment or statistical evaluation is involved.

ISO/IEC TR 29119-11:2020 discusses testing AI-based systems, including the test oracle problem and black-box and neural-network white-box approaches. The official listing marks this technical report as under review, so check its status before presenting it as the latest guidance: ISO/IEC TR 29119-11:2020.

9. A practical release checklist

  1. Identify critical user journeys and the consequences of their failure.
  2. Prioritize risks and choose test objectives that address them.
  3. Cover behavior at component, integration, and end-to-end levels as appropriate.
  4. Select design techniques suited to input boundaries, rule combinations, and user workflows.
  5. Include relevant performance, resilience, security, accessibility, privacy, usability, and localization checks.
  6. Confirm test data, environments, and dependencies are understood and reproducible.
  7. Review failures, unresolved defects, coverage limitations, and residual risk.
  8. Make the release decision using the evidence and its limits, not a single pass rate or coverage percentage.

10. Capture visual evidence for UI checks

When visual rendering is part of a critical workflow, screenshots can help document the page state under test. Choose viewport, device scale, wait condition, and target element to match the question: a full-page image can reveal layout or content issues across a long page, while an element capture can focus on a component. Screenshots are evidence of a rendered state; they do not replace interaction, accessibility, or functional checks.

DIY browser capture with Playwright

This Node.js example captures a page after its main content selector appears. Install Playwright with npm install playwright and install its browser with npx playwright install chromium. Save the following as capture.mjs and run node capture.mjs https://example.com.

import { chromium } from 'playwright';

const target = process.argv[2];
if (!target) throw new Error('Usage: node capture.mjs https://example.com');

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 1000 }, deviceScaleFactor: 1 });
  await page.goto(target, { waitUntil: 'domcontentloaded', timeout: 30000 });
  await page.locator('body').waitFor({ state: 'visible', timeout: 10000 });
  await page.screenshot({ path: 'shot.png', fullPage: true });
} finally {
  await browser.close();
}

For an element-focused capture, replace the screenshot call with await page.locator('main').screenshot({ path: 'main.png' });. Choose a selector that is stable for the page. Increase the timeout only when the page’s expected load behavior warrants it, and avoid waiting for every network request to finish if the site keeps long-lived connections open.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. One GET request returns an image or PDF. See the ScreenshotNeo website and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and whether the shot was billed.
  • An MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
  • The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card.

11. Troubleshooting common testing problems

Symptom Likely cause What to do
A large suite passes but users still find defects Tests may miss important risks, states, platforms, or quality attributes. Review escaped defects against risk assumptions and add evidence that addresses the missed condition.
End-to-end checks fail intermittently Timing, shared state, unstable selectors, or external dependencies. Capture failure context, isolate the unstable dependency where appropriate, and reserve full workflows for valuable journeys.
High code coverage gives false confidence Executed code paths may not assert meaningful outcomes or cover user risks. Review assertions and link test objectives to behaviors and quality risks.
Test data differs from realistic cases Fixtures may omit boundaries, invalid states, or representative variation. Add equivalence classes and boundary cases based on actual rules; document data assumptions.
Release reviews cannot explain readiness Reports show pass rates without gaps, failures, or residual risk. Report what was tested, what was not, unresolved defects, and the decision rationale.
AI outputs vary and exact-match checks fail Expected results may not be deterministic or exact text may not express acceptance. Define task-based criteria and an explicit evaluation method, including human review where needed.
Playwright capture times out The page or selector did not reach the chosen condition within the timeout, or navigation is waiting on a persistent request. Wait for a meaningful stable selector, use an appropriate navigation condition, and inspect the page error and network dependencies.
Screenshot is blank or incomplete The capture may run before content appears, or lazy content may require scrolling. Wait for the relevant selector or app-ready state; for DIY full-page capture, scroll or otherwise trigger lazy content before capturing if needed.

12. FAQ

How much testing is enough to qualify a software release?

It depends on the software type, purpose, audience, and consequences of failure. Google’s testing guidance makes the same context-dependent point. Define risk-based exit reasoning and state the limits of the evidence rather than looking for a universal test count.

Should every feature have an end-to-end test?

No fixed rule fits every system. Use end-to-end checks for important complete journeys and use component or integration checks for narrower behavior where they provide faster, clearer evidence.

Does a passing suite prove the product has no defects?

No. It establishes that the tests performed passed under the conditions used. Sampling, untested states, and changing environments leave residual risk.

Is a fixed testing pyramid ratio required?

No. A published 70/20/10 split was presented as a first guess, not a universal standard. Choose the mix that fits your product risks and feedback needs.

Can screenshot evidence replace UI automation?

No. A screenshot records rendered appearance at a moment. It does not establish that controls work, keyboard navigation is usable, or the workflow succeeds.

References