ScreenshotNeo

BlogHow-to

How to Use Automated Exploratory Testing

Use timeboxed exploratory sessions to uncover unknown behavior, document evidence, and turn important discoveries into focused automated regression checks.

By the ScreenshotNeo team4 October 20268 min read

Automated exploratory testing combines human-led investigation with automated checks that preserve important discoveries. A tester explores a feature without following a script that dictates the expected outcome, records what happens, and then automates stable, repeatable scenarios that should not regress. Automation can help capture actions and verify behavior; it does not replace the tester’s judgment about what to investigate or what a result means.

The practical sequence is: set a mission, prepare a timeboxed session, explore and record evidence, triage discoveries, automate selected regression checks, and report follow-up work. GOV.UK describes the goal as exploring a system as a user would, without a script to test a predetermined outcome.

1. Define the session mission

Choose a feature or workflow with enough functionality to explore. State the user or business goal and the uncertainty you want to investigate. A charter gives the session boundaries without prescribing every action.

Charter field What to record
Goal The user outcome or risk you want to understand.
Scope The feature, workflow, roles, and relevant boundaries.
Tester and session Who is exploring, when, and for how long.
Environment Build or release, browser or device, and any relevant configuration.
Test data Accounts, records, permissions, and data reset approach.
Questions Unknowns that may guide probes, without predetermining the result.

Example charter: “Explore password reset for a returning user, focusing on expired links, repeated requests, and recovery after a network interruption. Use the staging build, a disposable account, and a 40-minute session.” This defines an area and purpose, but leaves the exact path open.

2. Prepare a focused, timeboxed session

  1. Confirm access to the application and the correct environment.
  2. Prepare test data that is safe to modify and can be reset or recreated.
  3. Choose a timebox long enough to investigate one meaningful area. Reserve time afterward to assess and report findings.
  4. Open a place to record notes and capture screenshots, logs, or other useful evidence.
  5. Check that the session will not disrupt shared test data or real users.

Session-management software is optional. GOV.UK notes that pen and paper are enough to begin. Use a dedicated workflow when coordinating many testers or collecting evidence centrally, not as a prerequisite for exploration.

3. Explore, observe, and adapt

Interact with the product as a user would. Make a probe, observe the result, and let what you learn guide the next probe. Follow unexpected behavior and use domain knowledge to investigate it. A charter is a mission, not a fixed test script.

For the password-reset example, probes might include using an expired link, requesting several links in a row, opening a link in another browser context, or interrupting the network while submitting. These are investigation ideas; the tester should adapt them to the product and to observations during the session.

  • Record what you did when it helps someone understand or reproduce an observation.
  • Capture relevant conditions such as account state, permissions, browser, and build.
  • Write down questions and follow-up probes as they arise.
  • Distinguish expected product behavior from assumptions. If the expected behavior is unclear, record the question for product or design follow-up.

4. Capture useful evidence

Evidence makes findings easier to investigate, repeat, and communicate. Notes can include the charter, features covered, relevant actions and conditions, observations, questions, bugs, and supporting screenshots or logs. Record enough context to make a finding actionable, while avoiding sensitive user data.

  • Observation: what happened, stated plainly.
  • Reproduction context: environment, account state, data, and relevant actions.
  • Expected behavior: only when it is known; otherwise mark it as a question.
  • Impact and uncertainty: who may be affected and what remains unclear.
  • Evidence: screenshot, log, trace, or other artifact, with enough context to locate it.

A screenshot is useful when visual state matters, such as a missing confirmation or a broken layout. It does not replace notes about the steps or conditions that produced the state. For browser UI evidence, a screenshot API can capture a page or element without requiring a local browser setup. ScreenshotNeo is a website screenshot API and MCP server; its options include full-page and selector-based capture, custom CSS and JavaScript, and browser/device settings. See the ScreenshotNeo API documentation for request parameters.

5. Triage discoveries after the session

Separate confirmed defects from questions, risks, and ideas for further exploration. Then decide which discoveries deserve a repeatable regression check.

Discovery Next step
Confirmed defect with a clear reproduction Report it with evidence; after the behavior and fix are understood, add a regression check if it is important and repeatable.
Unclear or inconsistent behavior Record the question and ask the relevant product or design owner to clarify expected behavior.
Risk with no failure observed Document the risk and plan a focused follow-up probe or test.
Useful new test idea Keep it as a follow-up charter or scenario; do not mistake an unverified idea for a defect.

Not every action taken during exploration should become an automated test. Prefer checks for valuable behavior with a clear expected result and a stable way to set up the needed state. Keep exploratory work available to discover behavior that existing checks do not anticipate.

6. Turn a discovery into a Playwright regression check

For browser workflows, Playwright can help record actions and assertions. Its generator produces code to inspect and adapt, not a finished test strategy. Playwright’s guidance emphasizes user-visible behavior, isolated tests, resilient user-facing locators, and assertions that wait and retry. It supports JavaScript/TypeScript, Python, Java, and .NET; use the language and runner that fit the existing project.

JavaScript example

Suppose exploration found that submitting a valid reset request should display a confirmation. This runnable example uses Playwright Test, an environment variable for the application URL, and a locator based on accessible user-facing text. Replace the example route and accessible name with those used by the application.

import { test, expect } from '@playwright/test';

test('password reset request shows confirmation', async ({ page }) => {
  const baseURL = process.env.APP_URL;
  if (!baseURL) throw new Error('Set APP_URL to the application base URL');

  await page.goto(new URL('/forgot-password', baseURL).toString());
  await page.getByLabel('Email address').fill('qa-reset@example.test');
  await page.getByRole('button', { name: 'Send reset link' }).click();

  await expect(page.getByRole('status')).toContainText('Check your email');
});

Install the project’s Playwright Test dependency and browser using the official installation guide, then run the test with APP_URL=https://your-staging-host.example npx playwright test. The example assumes the application labels its email field and button accessibly and exposes the confirmation with a status role; adapt those locators to the real UI.

Use code generation as a draft

Playwright’s test generator can record interaction and help pick locators. Review its output: remove incidental clicks, replace brittle selectors where possible, add the assertion that captures the discovered risk, and make setup isolated. A recording that only repeats actions without checking the expected user-visible result is not a useful regression check.

Python, Java, and .NET

Playwright supports these languages, but the test-runner integration and idiomatic setup differ. Use the project’s existing runner and follow the corresponding official language documentation rather than translating a JavaScript Test example mechanically. Keep the same test-design goals: isolate state, use resilient user-facing locators, and assert the behavior tied to the discovery.

7. Report results and revisit

Share what was explored, the charter and timebox, confirmed findings, unresolved questions, evidence, and recommended follow-up. If useful to the team, record time spent on setup, investigation, and reporting so the work is understandable. Track selected regression checks in the project’s regular workflow, and use Playwright traces or other project debugging evidence when a check fails.

Common mistakes and troubleshooting

Symptom Likely cause What to do
The session becomes a checklist run The charter dictates every action and expected result. Keep the goal and scope, but let observations determine the next probe.
A finding cannot be reproduced Notes omit account state, environment, data, or relevant conditions. Add the missing context and evidence; repeat the probe where possible.
Test code breaks after small UI changes Generated code uses brittle selectors or incidental structure. Prefer accessible roles, labels, and other user-facing locators; review generated code.
Automated test is flaky Shared state, timing assumptions, or an assertion that does not wait for the UI. Isolate test data and state, use web-first retrying assertions, and investigate traces rather than adding arbitrary sleeps.
Automated test passes but the defect returns The assertion does not express the user-visible behavior that failed. Revisit the original evidence and make the assertion verify the actual failure condition or required outcome.
Exploration produces many low-value tests Every probe is being automated regardless of repeatability or risk. Triage first; automate the stable, important scenarios and retain open questions for exploration.
Screenshot lacks useful context Only the image was saved, without the page state or steps. Pair it with session notes, environment and relevant logs. For dynamic pages, wait for the needed content before capture.

Performance, reliability, and cost

Exploration is timeboxed, so choose a mission narrow enough to investigate thoughtfully and leave time to report. Browser automation is most reliable when tests have isolated state, stable data, user-facing locators, and assertions that wait for the interface. Avoid treating fixed delays as proof that a page is ready.

The session itself can begin with notes and existing evidence tools; a paid session-management product is not required by the method. Automated checks add ongoing maintenance, so prioritize scenarios whose regression risk justifies that cost. Screenshot capture has its own latency and service costs; capture only evidence useful to the finding, and follow the chosen service’s documentation and billing rules.

Or skip the browser setup

ScreenshotNeo can return a screenshot or PDF from one GET request, and its API accepts familiar screenshot API parameter names. The following request saves a page screenshot; see the ScreenshotNeo documentation for output formats and capture options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing status in headers.
  • An MCP server lets AI agents, including Claude and Cursor, use screenshot tools.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is available on every plan.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does exploratory testing require a dedicated tool?

No. Notes, screenshots, and logs can be enough to start. Choose a session-management tool when its coordination or evidence workflow solves a real team need.

Should every exploratory session produce an automated test?

No. Preserve important, repeatable behavior as regression checks; keep uncertain areas open for further investigation.

Can AI or browser automation perform the exploration for a tester?

Automation can execute actions and check specified outcomes, but exploratory work depends on a tester interpreting observations and choosing what to investigate next.

When is a screenshot useful in a test report?

When visual state helps explain a finding. Include conditions and steps as well, since an image alone may not make the issue reproducible.

Sources