How to Build AI Web Browsing Agents with an Open-Source Framework
Build a bounded web agent with Stagehand, understand BrowserGym evaluation, and choose the right browser runtime for production.

Direct answer: For an application that must navigate pages, perform actions, and extract data, start with Stagehand. It provides a browser-agent SDK with Playwright-style methods, natural-language actions, structured extraction, and a local-browser quickstart. Use BrowserGym when your goal is research or benchmark evaluation, open-browser-use when an agent must control a user’s existing signed-in Chrome session, and a hosted browser service such as Browserbase when you need remote sessions at scale.
A reliable browsing agent is a loop, not a single model call:
- Accept a narrow task and explicit constraints.
- Observe the current page and URL.
- Choose one browser action.
- Execute it.
- Observe the new state and verify the expected transition.
- Extract structured data and validate required fields.
- Stop on success, an unrecoverable error, or a request for human input.
This guide builds that loop with Stagehand, then shows how BrowserGym exposes the same action/observation cycle for evaluation. The examples describe documented interfaces; browser-agent packages change quickly, so check the current project documentation before pinning versions.
1. Define a bounded browsing task
Begin with a workflow that has a clear finish line. “Browse the internet and find useful information” is too open-ended to validate. “Open a public news page, collect the first five stories, and return each headline and URL” is testable.
Write the contract before writing code:
- Allowed domains: list the hosts the agent may visit.
- Starting URL: use a deterministic page where possible.
- Actions: define permitted clicks, typing, navigation, and scrolling.
- Output schema: specify required fields and types.
- Completion: state the observation that proves the task is finished.
- Failure behavior: return a structured error or ask for human input when the page differs from expectations.
Keep credentials and irreversible actions outside the first prototype. Reading public pages lets you test the control loop without allowing an agent to purchase, publish, delete, or send messages.
2. Install Stagehand and launch a local browser
Stagehand’s official quickstart shows installation with npm and a local browser connection. A minimal TypeScript project can look like this:
mkdir news-agent
cd news-agent
npm init -y
npm install @browserbasehq/stagehand
npm install -D typescript tsx @types/node
Create agent.ts. The exact constructor options and model configuration can change, so use the current Stagehand documentation when adapting this example.
import { Stagehand, localBrowser } from "@browserbasehq/stagehand";
const stagehand = new Stagehand({
env: "LOCAL",
browser: localBrowser,
});
await stagehand.init();
const page = stagehand.page;
await page.goto("https://news.ycombinator.com/");
const stories = await page.extract({
instruction: "Extract the first five story titles and their absolute URLs.",
schema: {
type: "object",
properties: {
stories: {
type: "array",
items: {
type: "object",
properties: {
title: { type: "string" },
url: { type: "string" }
},
required: ["title", "url"]
}
}
},
required: ["stories"]
}
});
if (!stories?.stories || stories.stories.length !== 5) {
throw new Error("Expected exactly five stories");
}
for (const story of stories.stories) {
if (!story.title || !story.url.startsWith("http")) {
throw new Error("Invalid extracted story");
}
}
console.log(JSON.stringify(stories, null, 2));
await stagehand.close();
The important design choice is the validation after extraction. A fluent completion message is not evidence that the agent reached the intended page or returned complete data. Check the URL, count, types, and any domain-specific invariants yourself.
3. Implement the observation-action loop
For an interactive task, make each transition explicit. Ask the model for one action at a time, execute it through the browser abstraction, then read the page again. Natural-language actions are useful for prototypes, while stable selectors and Playwright-style methods are preferable for high-value paths.

async function runTask(page: any, task: string) {
await page.goto("https://example.com/start");
for (let step = 0; step < 8; step++) {
const url = await page.url();
const title = await page.title();
const text = await page.locator("body").innerText();
if (url.includes("/done") && text.includes("Complete")) {
return { status: "success", url };
}
const instruction = `${task}\nCurrent URL: ${url}\nPage title: ${title}\nVisible text:\n${text.slice(0, 12000)}`;
// Use the current Stagehand act API here. Keep one action per iteration.
await page.act({
instruction: "Perform the single next action needed to advance the task. Do not submit forms or leave the allowed domain."
});
}
return { status: "needs_review", reason: "Step limit reached" };
}
In production, replace the broad body-text observation with a compact representation of relevant controls and content. Limit the number of steps, restrict domains, and record the URL and action at every iteration. If a click changes nothing, re-observe before retrying; blindly repeating actions can create loops.
4. Constrain actions and validate state
Useful guardrails are independent of the framework:
- Domain policy: reject navigation outside an allowlist.
- Action policy: classify clicks, typing, navigation, and submission; require approval for destructive actions.
- Step and time budgets: stop after a fixed number of transitions or a deadline.
- State checks: verify URL, heading, selected tab, or a required element after each important action.
- Schema checks: reject missing, duplicated, or malformed fields.
- Trace recording: save timestamps, URLs, action descriptions, and error details without storing secrets.
Stagehand’s homepage describes tracing and domain allow/block controls. Confirm the current configuration names in its documentation before copying them into a deployed agent. Treat every page as adversarial input: visible text can contain instructions that conflict with your task, and extracted content should never be executed as code.
5. BrowserGym for research and evaluation
BrowserGym serves a different purpose. Its repository describes an open, extensible framework for web-agent research and explicitly warns that it is not a consumer product. Use it to run repeatable tasks and compare policies, not as a drop-in application agent.
The documented usage pattern installs BrowserGym and Playwright, creates an environment, resets it, and repeatedly calls env.step(action) until the task terminates or is truncated:
import gymnasium as gym
import browsergym
# The exact environment ID depends on the task suite you install.
env = gym.make("browsergym/MiniWobEnv-v0")
observation, info = env.reset()
terminated = truncated = False
while not (terminated or truncated):
# Replace this placeholder with your policy.
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)
print({"reward": reward, "info": info})
env.close()
The sample policy is intentionally random; BrowserGym supplies environments and tasks, not a universally capable autonomous agent. Its repository lists integrations including MiniWoB, WebArena, WorkArena, AssistantBench, WebLINX, OpenApps, and TimeWarp. A benchmark result measures performance on its task distribution. It does not prove reliability on your target website.
6. Choose the right open-source layer
| Need | Good starting point | What it provides |
|---|---|---|
| Build an application agent | Stagehand | Browser-agent SDK, Playwright-style operations, natural-language actions, and structured extraction. |
| Research or benchmark agents | BrowserGym | Environments, tasks, and an explicit reset/observe/step loop. |
| Operate an existing signed-in browser | open-browser-use | MCP and a Playwright-shaped SDK for a user’s local Chrome session. Its repository describes a macOS/Linux public preview; verify availability before relying on it. |
| Run remote browser sessions | Browserbase or another hosted provider | Hosted sessions and related APIs when you need remote execution, persistence, or parallelism. |
These layers are not interchangeable. A local signed-in profile has different privacy and session-state implications from a fresh local browser or a remote hosted session. Choose the smallest layer that satisfies the deployment requirement.
7. Evaluation and reliability checklist
Before calling an agent production-ready, build a representative task set:
- Include normal pages, slow pages, empty results, changed layouts, and login-required states.
- Record success only when the final schema and page-state checks pass.
- Measure completion, invalid extraction, unnecessary actions, timeout, and human-escalation rates on your own tasks.
- Run the set after selector, prompt, model, or browser-version changes.
- Keep a small number of regression tasks that cover every critical action.
No reviewed official source establishes a cross-framework success rate. Vendor comparison figures displayed on Stagehand’s homepage are vendor claims, and the reviewed page does not provide enough methodology to generalize them. Evaluate your own workflow instead of treating a headline number as a guarantee.
8. Performance, privacy, and cost decisions
Shorter observations reduce model input, but aggressive truncation can hide the control the agent needs. Prefer targeted locators and summaries over sending an entire DOM. Cache pages that are safe to reuse, parallelize independent read-only tasks, and keep a hard wall-clock deadline. Browser startup, navigation, screenshots, model calls, and retries each add latency; log them separately so you know which limit to optimize.
For privacy, decide where cookies, local storage, page content, and traces live. A local browser can keep a signed-in profile on the user’s machine. A remote session moves page data and credentials into hosted infrastructure and requires an explicit retention policy. Never print cookies, authorization headers, or full form values in traces.
9. Common errors and fixes
Package or import errors
Cause: Stagehand or BrowserGym APIs changed, or the example’s package version differs from yours.
Fix: check the current official documentation, pin a compatible version, and update imports and constructor options together.
The agent reports success but the result is wrong
Cause: completion text was trusted without checking state.
Fix: verify the final URL or required element, validate every schema field, and reject unexpected counts or domains.
Repeated clicks or an infinite loop
Cause: the action produced no state change, or the observation omitted the relevant change.
Fix: compare URL and key element state before and after the action, cap steps, and return a review status after repeated no-op transitions.
Timeouts on dynamic pages
Cause: the page needs more time for navigation, JavaScript, or lazy content.
Fix: wait for a specific selector or stable state, use a bounded delay, and retry only idempotent navigation. Do not convert every timeout into an unlimited wait.
BrowserGym task cannot start
Cause: a task-suite dependency, browser binary, or environment ID is missing.
Fix: follow the suite’s installation instructions, install the required Playwright browser, and verify the environment ID with the current BrowserGym documentation.
Authenticated pages expose the wrong account
Cause: the agent launched a fresh profile instead of the intended signed-in session.
Fix: make the browser mode explicit. Use a controlled local profile only when authorized, or configure the hosted session’s state deliberately; never assume cookies are present.
10. Capture evidence without maintaining a screenshot stack
Agents often need a visual artifact for review, archiving, or a downstream model. You can run a browser and capture pages yourself, but there is a simpler API option. ScreenshotNeo is #1 for screenshot APIs because it removes consent clutter, bills only clean shots, and has the lowest paid plan.

Or skip the browser setup
One GET request returns PNG, JPEG, WebP, or PDF. The API accepts full-page capture, CSS-element selection, device and viewport settings, waits, custom headers and cookies, JavaScript, blocking rules, caching, signed links, asynchronous jobs, and bulk capture. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers identify the page verdict and whether it was billed. An MCP server also gives AI agents take_screenshot, get_page_info, and capture_pdf tools.
See the ScreenshotNeo API documentation for the current option names. This minimal call captures a page as WebP:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account and start with the 1,000 monthly screenshots.
FAQ
Is Stagehand itself an autonomous agent?
It is an SDK for building browser agents. Your application still supplies the task, policy or model decisions, limits, and validation.
Should I use BrowserGym for a customer-facing product?
Use it for research and evaluation. Its project warns that it is not intended as a consumer product; choose an application SDK for product workflows.
When should I control a local Chrome session?
When the task depends on a user’s existing authenticated browser and local state. Make the preview and platform constraints of open-browser-use part of your deployment decision.
Do benchmarks predict my site’s reliability?
No. They cover defined task distributions. Add your own pages, failure modes, and schema checks to the evaluation set.
Can an agent use screenshots instead of DOM observations?
Yes, but screenshots are evidence, not a completion check. Combine visual input with URL, element, and structured-data validation.