Why Manual Testing Still Matters
Manual testing still matters because automated checks can only judge what teams have encoded. Learn where human exploration adds value and how to pair it with automation.
Why does manual testing still matter when teams have automated tests? Because automated tests check expectations a team has already expressed in code, while a person can investigate unfamiliar behavior, change the next test based on what they observe, and judge whether the result makes sense for a user. Manual testing is most useful alongside automation: automate stable, repeatable checks, and use focused human exploration where context, discovery, or judgment matters.
Manual testing is not simply a slower way to replay a script. In exploratory testing, test design, execution, and evaluation happen together as the tester learns about the software. That makes it useful for questions the team has not yet translated into assertions. It does not make manual testing a substitute for a dependable automated regression suite.
1. What manual testing contributes
A test script follows a known path and evaluates stated conditions. A manual tester can notice an unexpected result, form a new question, and probe it immediately. For example, while exploring a checkout flow, a tester might notice that changing the delivery address resets the selected shipping method. They can then vary the address, payment choice, session state, or sequence of actions to understand the behavior and whether it violates user expectations.
That cycle is the useful distinction: the human chooses what to investigate next based on evidence gathered so far. The goal is not random clicking. A focused exploratory session has a purpose, a time limit, and notes about what was tried and learned. ISTQB describes exploratory testing as tests being designed, executed, and evaluated together while the tester learns about the test object (ISTQB Foundation Level experience-based techniques).
Google’s account of testing Google Talk is a project example of planning both automated and manual work. It describes a test plan that identified where automation applied and the role manual testing still played before release. It is an illustration of one team’s practice, not a rule that every product needs the same split (Google Testing Blog: Exploratory Testing on Chat).
2. What should you test manually?
Prioritize behavior that is new, changing, risky, ambiguous, or sensitive to user context. Manual investigation is especially useful when the team does not yet know all the relevant states or when deciding whether an outcome is acceptable requires interpretation.
- New features and unfamiliar workflows: Explore how users can reach a feature, what happens when they change direction, and whether surrounding behavior remains coherent.
- High-risk changes: Investigate affected workflows where failure could cause data loss, expose information, block a core task, or create a costly operational issue. Use the product’s own risk assessment to decide priority.
- Boundary and unusual states: Try empty, long, malformed, stale, repeated, interrupted, or conflicting inputs where those states are relevant. Follow surprising results with targeted probes.
- Cross-step journeys: Examine whether a critical journey feels understandable and complete from the user’s point of view, including recovery from validation errors and interruptions.
- Visual and browser behavior: Review layout, content, responsive states, and browser-specific rendering when those are part of the change. A screenshot can preserve what a page looked like at a particular point, but it does not by itself establish that the interaction or result was correct.
- Features with judgment-based expectations: Check whether the behavior fulfills the user’s intent, not merely whether the page rendered or a request returned success.
Manual testing is less efficient for repeatedly checking a stable, deterministic condition across many builds. Once an exploratory session reveals a reproducible defect or an important invariant, record it as a test case and automate it where the check is reliable and valuable.
3. Can automation replace exploratory testing?
Automation can execute encoded checks consistently and frequently. It cannot automatically answer questions the team has not represented as checks. This is a limit of the current test design, not a claim that software cannot automate more sophisticated analysis: a test can only provide evidence about the conditions it actually evaluates.
A passing suite therefore means that the checks which ran passed under the conditions of that run. It does not prove that no defects exist or that users will be satisfied. Google’s testing guidance cautions that code coverage alone does not establish that covered code is bug-free, and recommends a layered strategy with unit tests, integration tests, end-to-end tests for critical user journeys, and attention to code and functional coverage (Google Testing Blog: How Much Testing is Enough?).
| Testing need | Useful emphasis | Reason |
|---|---|---|
| Stable checks after code changes | Automated unit and integration tests | They can rerun encoded expectations consistently. |
| Critical end-to-end journeys | Automation plus human review | Automated assertions check known steps; a person can judge whether the journey remains clear and useful. |
| Unfamiliar or changing behavior | Manual exploratory testing | The tester can adapt the next probe as new evidence appears. |
| User intent and unclear acceptance criteria | Manual judgment informed by requirements and user context | A script cannot establish user satisfaction unless the team has defined an appropriate observable measure. |
| AI-based behavior | Planned combination of technical checks and human evaluation | Outputs may be probabilistic, non-deterministic, and dependent on data. |
4. How manual and automated testing work together
- Build a repeatable base. Cover small, well-defined rules with unit tests and interactions between components with integration tests. Keep these checks focused on stable expectations.
- Automate the critical journeys. Choose a limited set of end-to-end paths that matter to users. Assert outcomes that can be checked reliably, and avoid making every visual detail a brittle assertion.
- Plan exploratory sessions around risk. Give each session a charter, such as “Explore how an interrupted payment can be resumed.” Set a time box and identify the build, environment, account state, and relevant data.
- Explore, observe, and adapt. Start from the charter, vary inputs and sequence where useful, and follow unexpected results. Keep notes so another person can understand the path and reproduce a suspected defect.
- Turn discoveries into durable checks. Report reproducible bugs with steps, expected and actual behavior, and useful evidence. After fixing a defect, add an automated regression check when the behavior can be asserted consistently.
- Review gaps and remaining risk. Coverage numbers describe which code or functions were exercised; they do not, by themselves, describe the quality of assertions or whether the test plan matches user risk. Document what remains uncertain.
A short session record can include the charter, build and environment, data used, paths explored, observations, defects filed, evidence captured, and questions left open. This is practical team guidance: the point is to make learning and remaining risk visible, not to produce paperwork for its own sake.
5. Manual testing for AI-based features
AI-based features can produce probabilistic or non-deterministic behavior and depend on input data. That makes a single fixed expected output a poor fit for some checks. Teams can still automate deterministic requirements, input validation, safety constraints, and measurable properties, while human evaluators examine examples where quality depends on context or intended use.
Plan evaluations around representative inputs and meaningful failure cases. Record the model or system version, relevant configuration, prompt or input, and observed output so a result can be interpreted later. For generative features, consider how wording, context, and unusual inputs affect the outcome; judge results against explicit criteria rather than intuition alone. ISTQB’s CT-AI Version 2.0 covers AI-system testing across the lifecycle, including data, models, and machine-learning development, and identifies characteristics such as probabilistic behavior and non-determinism (ISTQB CT-AI; CT-AI syllabus version 2.0 announcement).
6. Capturing visual evidence from manual sessions
When a defect depends on what appeared in a browser, a screenshot helps communicate the observed state. Capture the relevant viewport, include enough context to identify the affected element, and record the browser, viewport, account state, and steps that produced it. A screenshot documents appearance at one moment; pair it with reproduction steps and behavioral evidence when the issue involves interaction, timing, or data.
For repeatable browser evidence, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API can capture a page or element and return an image or PDF. This can support evidence gathering around manual review, but it does not replace a tester’s judgment about whether the page behaves correctly.
7. Or skip the browser setup
For a one-call capture, see the ScreenshotNeo API documentation and use the request below. Replace the key and target URL with your own.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card.
8. Common mistakes and troubleshooting
“We have high coverage, so manual review is unnecessary.”
Cause: Treating coverage as proof of correctness. Coverage shows which code or functions ran under a measure; it does not show that assertions are sufficient or that the outcomes make sense to users. Fix: Review assertions and user risks, and schedule exploratory sessions for new or high-risk behavior.
“Exploratory testing is just unstructured clicking.”
Cause: Starting without a question or recording what happened. Fix: Use a charter, time box, known environment, and concise session notes. Follow observations with deliberate probes.
“The suite passed, but users still report a problem.”
Cause: The failing condition, data, sequence, environment, or expectation may not be represented in the checks that ran. Fix: Reproduce the report, identify the missing condition, and add a stable regression check where possible. Keep an exploratory session focused on related states until the gap is understood.
“A screenshot proves the feature works.”
Cause: A still image shows appearance, not necessarily the interaction, result, or sequence that produced it. Fix: Attach the steps and relevant state, and use a behavioral check or recording where timing or interaction matters.
“An AI result changed, so the system is broken.”
Cause: Expecting identical output from a system whose behavior may be probabilistic or non-deterministic. Fix: Compare against explicit quality criteria and safety or product constraints, record relevant input and version details, and investigate whether the variation crosses an unacceptable boundary.
9. Reliability, performance, and cost
Automation is usually the practical choice for checks that must run repeatedly and quickly, but end-to-end checks include more dependencies and can be less reliable than smaller integration checks. Keep the automated suite layered so a failure can be localized, and reserve human time for questions where exploration is likely to reduce uncertainty.
Manual sessions cost focused tester time and do not provide the same repeatable execution as a script. Their value depends on selecting useful charters, recording discoveries, and converting important reproducible behavior into checks. There is no universal percentage of testing that should be manual: choose the mix based on product risk, rate of change, testability, and the consequence of missed behavior. The available sources do not establish a general defect-yield figure or ideal manual-testing share.
Browser screenshots also have a narrow role: they can preserve visual evidence, but they do not establish correctness or substitute for functional assertions. ScreenshotNeo bills only clean shots; its responses identify page verdict and billing status in headers. Its listed plans are Free (1,000 shots/month), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free, and every feature is on every plan. Choose a capture workflow based on whether saved visual evidence is useful to your process, not as a replacement for testing the application.
10. Frequently asked questions
Does manual testing mean testing without a plan?
No. Exploratory testing is most useful when guided by a clear objective and followed by notes about the paths and findings.
Should every manual test become an automated test?
No. Automate checks that are valuable, repeatable, and expressible as reliable assertions. Some judgments are better handled through periodic human review.
Can developers do exploratory testing?
Yes. Anyone with the needed product and technical context can explore behavior. A different tester may still bring useful user, domain, or accessibility perspectives.
How do we know when to stop exploring?
Use the session time box and charter, then stop when time is up or the session has answered its main questions. Record unresolved risks and decide whether another focused session is justified.


