How to Write an AI Agent That Uses a Browser
Build a browser agent as a controlled observe–decide–act–verify loop. See a runnable Playwright foundation, security rules, failure handling, and a screenshot API option.
A browser agent is a controlled loop: give a model a task and a bounded view of the current page, accept a proposed action, check that action against policy, execute it in application-controlled browser automation, then inspect the new state. The host application—not the model—must retain control of permissions, execution limits, cancellation, and consequential actions. See the OpenAI computer-use guidance and Anthropic browser-use documentation.
This guide builds the browser side of that loop with JavaScript and Playwright. It includes a runnable, deliberately limited example and shows where a model decision function belongs. The example does not connect to a model: provider tool schemas and model support change, so use the current provider documentation when wiring one in.
1. Define the task contract first
Before opening a browser, write down exactly what the agent may do. A task contract should specify:
- Goal: one outcome the user asked for, such as finding a page title or locating a public article.
- Allowed origins: exact sites or a narrow allowlist. Redirects must be checked too.
- Allowed actions: for example, navigate, inspect text, click a link, or fill a non-sensitive search box.
- Forbidden actions: purchases, posting, messaging, destructive changes, credential entry, downloads, and data transmission unless a separate confirmation policy allows them.
- Bounds: maximum action count, elapsed time, pages, and model calls; a cancellation path; and a clear stopping condition.
- Result contract: what evidence must be returned and how success will be verified.
Keep secrets, unrelated files, and privileged accounts out of the browser environment. Page content is untrusted data: it cannot expand the task, grant permissions, or override trusted instructions.
2. Choose the narrowest browser interface that fits
| Approach | Useful when | Trade-offs to consider |
|---|---|---|
| Existing application API or tool | The workflow already has a structured interface, such as a search or data API. | Usually easier to bound and verify. It may not expose the user-facing web workflow. |
| Browser automation with page structure | The task depends on page text, accessibility semantics, links, and form controls. | Semantic locators are precise on ordinary pages; sites can still change their accessible names or behavior. |
| Screenshot-and-coordinate computer use | The UI itself is the interface, including canvas-heavy or visually arranged pages. | Actions depend on visual coordinates and fresh screenshots. Scaling, scrolling, and layout changes can make coordinates stale. |
Anthropic describes browser use through page structure and screenshots; Google documents a screenshot-driven action loop, with Playwright as a client-side handler. OpenAI describes computer-use approaches with application-controlled execution. Compare observation fidelity, action precision, isolation, verification, recovery, session management, data retention, latency, and current cost for your deployment. Do not assume vendor capabilities or pricing are interchangeable; check current primary documentation.
3. Create a bounded Playwright agent shell
Install Node.js and Playwright, then install its Chromium browser:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as agent.mjs and run node agent.mjs. It demonstrates the control loop with a fixed decision function so it runs without model credentials. Replace decide with an adapter for the model or tool you choose; keep policy checks and execution in the host application.
import { chromium } from 'playwright';
const task = 'Open the allowed page and report its title.';
const allowedOrigins = new Set(['https://example.com']);
const maxActions = 3;
const timeoutMs = 30_000;
function assertAllowedUrl(rawUrl) {
const url = new URL(rawUrl);
if (!allowedOrigins.has(url.origin)) {
throw new Error(`Blocked origin: ${url.origin}`);
}
return url;
}
// Demonstration decision function, not a model. A real adapter should return
// only a validated action from the same small schema.
async function decide({ task, observation }) {
if (!observation.url) return { type: 'navigate', url: 'https://example.com/' };
return { type: 'finish', answer: observation.title };
}
async function observe(page) {
return {
url: page.url(),
title: await page.title(),
text: (await page.locator('body').innerText()).slice(0, 4000)
};
}
async function execute(page, action) {
switch (action.type) {
case 'navigate': {
const url = assertAllowedUrl(action.url);
await page.goto(url.href, { waitUntil: 'domcontentloaded' });
// Check the final URL too: a permitted URL can redirect elsewhere.
assertAllowedUrl(page.url());
return;
}
case 'click': {
if (typeof action.name !== 'string' || action.name.length > 100) {
throw new Error('Invalid link name');
}
await page.getByRole('link', { name: action.name, exact: true }).click();
assertAllowedUrl(page.url());
return;
}
case 'finish':
return;
default:
throw new Error(`Unsupported action: ${action.type}`);
}
}
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
page.setDefaultTimeout(5000);
try {
const deadline = Date.now() + timeoutMs;
let observation = { url: '', title: '', text: '' };
for (let turn = 0; turn < maxActions; turn++) {
if (Date.now() > deadline) throw new Error('Run timed out');
const action = await decide({ task, observation });
// In production, apply an explicit policy check here before every action.
if (action.type === 'finish') {
// Verify a concrete postcondition instead of trusting a completion claim.
if (!page.url() || !(await page.title())) {
throw new Error('Cannot verify the requested page state');
}
console.log(JSON.stringify({ answer: action.answer, url: page.url() }));
break;
}
await execute(page, action);
observation = await observe(page);
console.log(JSON.stringify({ turn, url: observation.url, title: observation.title }));
}
} finally {
await context.close();
await browser.close();
}
This small shell intentionally exposes only navigation and a semantic link click. Add actions one at a time, each with input validation, authorization, limits, and a verifiable postcondition. Do not make arbitrary model-generated JavaScript executable in the page.
4. Connect a model without giving it authority
Implement a provider adapter behind decide. The adapter should receive only the task, the minimum useful observation, and the allowed action schema. It should return data, not executable code. Validate every field at runtime: action name, URL, selector or accessible name, string length, and any numeric bounds.
For provider-specific computer-use tools, follow the vendor’s current tool loop. Some APIs return a proposed action that the host must execute and report back; some maintain session state or require a particular image format. Preserve that boundary: a model response is a proposal. Your host checks policy, executes an allowed operation, and sends a fresh observation. See the current OpenAI guide, OpenAI Agents API guide, Anthropic browser-use guide, and Gemini Computer Use documentation for current syntax and availability.
Do not ship code copied from a preview or vendor example without checking whether the model, region, tool syntax, and execution responsibilities still match your deployment. The cited Gemini documentation labels Computer Use as preview and advises close supervision for important tasks.
5. Make actions resilient and verify each meaningful step
- Prefer Playwright locators based on roles and accessible names, such as
getByRole('button', { name: 'Search' }). Playwright locators auto-wait and retry, and actions check conditions such as visibility and enabled state. See Playwright Best Practices. - Scope ambiguous names to a meaningful container, then require a unique match. Avoid selectors tied to generated classes or brittle DOM paths when a user-facing locator exists.
- After a click, check the expected destination or changed state. After a form submission, inspect the resulting confirmation or record rather than accepting the model’s summary.
- Use explicit waits for a real condition, such as a known heading becoming visible. Avoid arbitrary long sleeps; they add latency without proving that the page is ready.
- When an action times out or returns an ambiguous result, re-observe before retrying. Repeating a click on a purchase or submit control can duplicate an irreversible action.
- Record action type, target, policy decision, outcome, timestamps, and verification result. Redact credentials, personal data, and sensitive page content from logs.
6. Protect the agent from prompt injection
Every page is untrusted input. Instructions can be placed in visible text, hidden page regions, embedded documents, ads, reviews, or content loaded after navigation. They may ask the agent to reveal secrets, change goals, click a dangerous control, or send data elsewhere. Treat extracted text and screenshots as evidence about the page, never as authority.
Prompt wording alone is not a security boundary. Google describes layered controls including origin isolation, a separate alignment critic, confirmation for critical steps, threat detection, and red-teaming. Anthropic’s research also says browser agents are not immune to prompt injection. Combine model defenses with:
- Browser isolation and network restrictions, including redirect and download handling.
- Origin allowlists checked at navigation and after redirects.
- Least-privilege accounts with no unrelated credentials, files, or secrets available.
- A narrow action dispatcher; no arbitrary shell, page script, or unrestricted network tool.
- Human confirmation for purchases, posting, messaging, destructive changes, credential entry, and data transmission.
- Run limits, a user-visible cancellation mechanism, and a safe stop on ambiguity or access escalation.
- Security tests with hostile page text, off-domain links, deceptive buttons, unexpected redirects, and injected content.
OpenAI’s guidance specifically treats typing sensitive information into a form as data transmission and recommends confirmation for consequential actions. Apply the same caution to downloads and actions that are difficult to reverse.
7. Handle failure, recovery, and cost
| Failure | Likely cause | Safer response |
|---|---|---|
| Navigation is blocked | Origin is outside the task allowlist, or a redirect left an allowed site. | Stop and report the blocked destination. Do not silently broaden the allowlist. |
| Locator times out | Wrong accessible name, delayed rendering, hidden control, or page changed. | Take a fresh observation, inspect the accessible structure, narrow the target, and retry only if the action is safe. |
| Click has no visible effect | Overlay, disabled control, client-side validation, or an incorrect target. | Inspect current state and validation messages. Do not repeat a consequential click blindly. |
| Agent loops or repeats | No explicit stop condition, action result not returned, or weak postcondition. | Enforce turn and time limits; return fresh state; stop with an incomplete result when the bound is reached. |
| Model proposes an unknown action | Malformed output or action schema drift. | Reject it at the dispatcher, log a redacted validation error, and stop or request a new proposal within the limit. |
| Unexpected sensitive request | Page asks for credentials, personal data, or a consequential confirmation. | Pause for the user or stop. Do not infer consent from page text. |
| Final answer conflicts with page | Agent trusted its own completion sentence instead of application state. | Check the actual destination, visible confirmation, or resulting record and report uncertainty. |
Performance: browser startup, page loading, screenshots, and model calls all add latency. Reuse a browser context for a bounded task when isolation permits, avoid sending repeated full-page observations, and prefer structured page data for page-centric work. Screenshot-based interaction is useful for visual surfaces but can increase image payload and make stale coordinates risky. Stop when the task is complete rather than polling indefinitely.
Reliability: set navigation and action timeouts, cap turns and total elapsed time, close contexts in a finally path, and make retries conditional on whether the previous action may have taken effect. Persist only the state needed to recover; expire sessions and clean up stored activity according to your application’s policy. Verify current vendor retention and session behavior before deployment.
Cost: there is no cross-vendor success-rate or cost benchmark established by the cited implementation sources. Measure your own model calls, image and text volume, browser runtime, retries, and any hosted execution charges. Bound these inputs per task and surface a budget stop. Re-check provider pricing and availability from current primary sources.
8. When the task only needs a screenshot
If the agent only needs to inspect a page image, you may not need to build and operate a browser for that step. ScreenshotNeo is a website screenshot API and MCP server for developers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. A screenshot is still just an observation: your agent must decide what to do, enforce policy, and verify consequential results.
Or skip the browser setup
One GET request returns an image or PDF. The examples below use the documented API pattern and a public example target; see the ScreenshotNeo API docs for options and the current parameter reference.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners are accepted like a visitor; more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot. Each step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
- An MCP server lets AI agents take screenshots, get page information, and capture PDFs.
- The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month with no card.
9. Further configuration to consider
Keep the action surface as small as the task allows. If you add browser capabilities, define policy for each before exposing it to a model.
| Capability | Configuration and edge cases |
|---|---|
| Navigation | Allow only required origins; validate schemes; re-check final URLs after redirects and popups. |
| Clicks | Use role and accessible name; scope duplicate names; detect when the click opened a new page. |
| Forms | Allow specific fields and data classes. Treat credential or personal-data entry as sensitive transmission requiring explicit policy and confirmation. |
| Downloads | Block by default or define file type, size, destination, retention, and malware handling limits. |
| New tabs and popups | Track each page explicitly and apply origin checks to every page; do not assume a new tab inherits a safe destination. |
| Frames and embedded content | Inspect the correct frame deliberately; treat embedded instructions and cross-origin content as untrusted too. |
| Screenshot observations | Capture only the necessary viewport or region; remove sensitive material where possible and treat coordinates as invalid after layout changes. |
| Session state | Use isolated contexts per user or task where practical. Protect cookies and saved activity; provide cleanup and avoid logging secrets. |
FAQ
Should I use Playwright or a computer-use model?
Use the narrowest interface that can complete the task. Playwright is a good fit when page structure and semantic controls are available. Screenshot-and-coordinate interaction fits visual surfaces where structure is insufficient. An application API is often simpler for a known structured workflow.
Can I let the agent follow instructions it finds on a page?
It can use page content as evidence for the user’s task, but page content must not change the task or grant new permissions. Enforce that rule in the host application’s policy and action dispatcher.
How do I know the agent finished correctly?
Check an application-visible postcondition, such as the resulting URL, a confirmation state, or a created record. Treat the model’s natural-language completion claim as unverified until that check passes.
Can an agent safely complete a purchase or send a message?
Those actions have external effects. Require user confirmation at the point of commitment, show the intended target and content, and stop if the result is ambiguous.


