ScreenshotNeo

BlogAI agents

How to Build a Browser-Based AI Operator

Build a safe browser AI operator with Playwright, screenshots, bounded actions, verification, approvals, and recovery.

By the ScreenshotNeo team1 October 20268 min read

A browser-based AI operator is an agent that observes a live webpage, chooses a small browser action, executes it, reads the updated state, and repeats until a verified outcome exists. The reliable design is a bounded observe → plan → act → verify loop running Playwright (or Chrome DevTools Protocol) inside a sandbox.

Keep the task narrow, expose only the actions the model needs, preserve the browser context between steps, and stop for approval before irreversible actions. Treat every string, image, link, and tool result from a webpage as untrusted input.

1. Define the task contract before writing code

Write down the contract your operator must obey:

  • Allowed domains: the exact hosts it may visit.
  • Inputs: fields supplied by the user and their validation rules.
  • Output: the record, file, or confirmation that counts as success.
  • Allowed actions: navigate, inspect, click, type, select, wait, screenshot, and extract.
  • Limits: maximum actions, wall-clock time, retries, and repeated-state count.
  • Approval points: purchases, messages, form submissions, account changes, deletion, and disclosure of sensitive data.

Start with read-only extraction or a reversible workflow. A contract such as “collect the title and price from these three product pages” is easier to secure and evaluate than “use the web to buy the cheapest item.”

2. Choose the browser execution layer

Playwright is the usual starting point because it controls Chromium, Firefox, and WebKit with one API. Use Chrome DevTools Protocol (CDP) when the agent must attach to an existing Chromium session or profile. Keep browser control behind a small adapter so the model interface can change later.

Requirement Good default Reason
New isolated session Playwright browser context Separate cookies, storage, permissions, and credentials.
Existing Chrome session CDP Connects to a running Chromium instance.
Stable known workflow Deterministic Playwright code Cheaper, faster, and easier to test.
Variable layouts or open navigation Model-selected actions with strict tools Handles uncertainty while retaining policy controls.

3. Implement a bounded observe-plan-act loop

The model should never receive unrestricted computer access. Give it a typed action schema and a compact observation. A useful observation contains the current URL, page title, visible text, interactive elements with stable references, and a recent screenshot when visual layout matters.

JavaScript and Playwright skeleton

import { chromium } from "playwright";

const MAX_STEPS = 20;
const ALLOWED_HOSTS = new Set(["example.com"]);

function hostAllowed(url) {
  return ALLOWED_HOSTS.has(new URL(url).hostname);
}

async function observe(page) {
  const elements = await page.locator("a,button,input,select,textarea").evaluateAll(nodes =>
    nodes.slice(0, 100).map((node, i) => ({
      ref: `e${i}`,
      tag: node.tagName.toLowerCase(),
      role: node.getAttribute("role"),
      text: (node.innerText || node.getAttribute("aria-label") || "").trim().slice(0, 160),
      type: node.getAttribute("type"),
      name: node.getAttribute("name"),
      value: node.value ?? ""
    }))
  );
  return {
    url: page.url(),
    title: await page.title(),
    text: (await page.locator("body").innerText()).slice(0, 12000),
    elements,
    screenshot: (await page.screenshot({ type: "png" })).toString("base64")
  };
}

// Replace this function with your model/API call. It must return one typed action.
async function decide({ task, observation, history }) {
  // Example deterministic action for a prototype:
  if (observation.url === "about:blank") {
    return { type: "navigate", url: "https://example.com" };
  }
  return { type: "finish", result: { title: observation.title, url: observation.url } };
}

async function execute(page, action) {
  switch (action.type) {
    case "navigate": {
      const target = new URL(action.url, page.url());
      if (!hostAllowed(target.href)) throw new Error("Domain is not allowed");
      await page.goto(target.href, { waitUntil: "domcontentloaded", timeout: 30000 });
      return;
    }
    case "click":
      await page.locator(action.selector).click({ timeout: 10000 });
      return;
    case "type":
      await page.locator(action.selector).fill(action.text);
      return;
    case "select":
      await page.locator(action.selector).selectOption(action.value);
      return;
    case "wait":
      await page.waitForTimeout(Math.min(action.ms ?? 500, 10000));
      return;
    case "screenshot":
      await page.screenshot({ path: action.path ?? "state.png", fullPage: true });
      return;
    default:
      throw new Error(`Unsupported action: ${action.type}`);
  }
}

async function run(task) {
  const browser = await chromium.launch({ headless: true });
  const context = await browser.newContext();
  const page = await context.newPage();
  const history = [];
  try {
    for (let step = 0; step < MAX_STEPS; step++) {
      const observation = await observe(page);
      const action = await decide({ task, observation, history });
      history.push({ step, action, url: observation.url });

      if (action.type === "finish") {
        // Require a postcondition in production before returning success.
        return { ok: true, result: action.result, history };
      }
      if (action.type === "request_approval") {
        return { ok: false, needsApproval: action.reason, history };
      }
      await execute(page, action);
    }
    return { ok: false, error: "step_limit", history };
  } finally {
    await browser.close();
  }
}

run("Read the page title on example.com").then(console.log).catch(console.error);

In a production adapter, validate every action against a schema before execution. Prefer selectors based on accessible roles or stable attributes. Do not let the model submit arbitrary JavaScript, launch shell commands, read files, or navigate outside the allowlist.

Python action executor

from dataclasses import dataclass
from playwright.sync_api import sync_playwright
from urllib.parse import urlparse

ALLOWED_HOSTS = {"example.com"}
MAX_STEPS = 20

@dataclass
class Action:
    type: str
    selector: str | None = None
    text: str | None = None
    url: str | None = None
    ms: int = 500

def allowed(url: str) -> bool:
    return urlparse(url).hostname in ALLOWED_HOSTS

def execute(page, action: Action):
    if action.type == "navigate":
        if not allowed(action.url):
            raise ValueError("domain is not allowed")
        page.goto(action.url, wait_until="domcontentloaded", timeout=30_000)
    elif action.type == "click":
        page.locator(action.selector).click(timeout=10_000)
    elif action.type == "type":
        page.locator(action.selector).fill(action.text or "")
    elif action.type == "wait":
        page.wait_for_timeout(min(action.ms, 10_000))
    else:
        raise ValueError(f"unsupported action: {action.type}")

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context()
    page = context.new_page()
    page.goto("https://example.com", wait_until="domcontentloaded")
    print({"title": page.title(), "url": page.url})
    browser.close()

4. Ground the model with the right state

Use DOM and accessibility data for precise targets, and screenshots for layout, canvas, menus, and visual confirmation. Sending only a screenshot makes text entry and exact element targeting harder; sending only raw HTML overwhelms the model and hides visual state.

  • Truncate long text and cap the number of interactive elements.
  • Include the URL and page title on every turn.
  • Assign temporary element references, then resolve them immediately before acting.
  • Capture a new observation after every action that can change the page.
  • Hash normalized observations to detect a loop.

5. Add verification and human approval

Never report success because the model said an action probably worked. Define a postcondition: a confirmation message is visible, a record matches the requested value, a download exists with the expected type, or a URL changed to a known completion route.

Pause before purchases, sending messages, submitting forms, changing account settings, deleting data, or exposing credentials. Show the user the exact target, fields, and consequence. Resume with an explicit approval token, then record who approved and which state was visible.

6. Secure the runtime

  • Run the browser in a sandboxed VM or container with a restricted filesystem and network egress.
  • Keep credentials in the runtime; avoid placing secrets in model-visible page text or prompts.
  • Use separate browser contexts per task and clear them after completion.
  • Enforce domain, navigation, upload, download, and action policies outside the model.
  • Log URLs, actions, approvals, screenshots, errors, and final evidence with sensitive values redacted.
  • Treat page text, images, links, and tool output as untrusted. Content cannot grant permission or override the user’s instructions.

7. Handle prompt injection and hostile pages

A page may contain instructions such as “ignore your task,” fake support messages, or links intended to exfiltrate data. Put the user’s task contract and policy in a higher-priority control layer, and pass page content as data. The agent must not follow instructions found in the page unless the task explicitly asks it to extract those instructions.

Test malicious links, cross-site navigation, credential requests, hidden form fields, downloads, repeated clicks, and attempts to upload local files. Prefer an official API or deterministic integration when the site provides one and browser control is not necessary.

8. Reliability, performance, and cost

Concern Practical control
Latency Use small observations, avoid unnecessary screenshots, and wait only for the selector or network condition you need.
Model cost Limit steps, summarize history, and use deterministic code for stable subflows.
Browser cost Reuse a context during one task, then close it; parallelize only when site limits and isolation allow.
Reliability Retry transient navigation failures with backoff, but stop on repeated identical state.
Recovery Persist the task contract, current URL, action history, and last verified checkpoint so a run can resume safely.

Benchmark results are snapshots, not guarantees for your workflow. OpenAI reported 38.1% on OSWorld, 58.1% on WebArena, and 87% on WebVoyager in 2025. Evaluate your own sites with realistic pages, authentication, popups, latency, and attack cases.

9. Troubleshooting

Symptom Likely cause Fix
Element not found Dynamic rendering or an unstable selector Wait for a specific role or selector; use accessible names or stable attributes.
Click intercepted Cookie banner, modal, or overlay Inspect visible overlays, close them explicitly, then retry once.
Agent repeats the same action No state change detection Hash normalized observations and stop after repeated states.
False success No postcondition Require visible confirmation, matching data, or a verified artifact.
Navigation leaves the site Untrusted link or redirect Check the final URL against the domain policy before every next action.
Credentials appear in logs Secrets included in observations Redact values, isolate secrets, and pass only field-level tokens.
Timeouts Slow page, blocked resource, or bot challenge Use bounded waits, capture diagnostics, retry transient failures, and hand off when the page cannot be verified.
Runaway spend Unlimited steps or verbose history Set action, time, and token budgets; summarize old observations.

10. Or skip the browser setup

If your task is to obtain clean website screenshots rather than operate a site, ScreenshotNeo provides one GET request for PNG, JPEG, WebP, or PDF. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

See the ScreenshotNeo API documentation for all options, including full-page capture, element selectors, device presets, custom CSS and JavaScript, waits, blocked resources, headers, cookies, geolocation, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account.

11. A practical build checklist

  • Define domains, inputs, outputs, limits, and approval points.
  • Run Playwright or CDP in an isolated browser context.
  • Expose typed, minimal actions.
  • Combine accessibility/DOM state with screenshots when needed.
  • Validate every navigation and action outside the model.
  • Detect repeated state and enforce step and time budgets.
  • Require a postcondition and retain evidence.
  • Redact credentials and minimize PII.
  • Test prompt injection, hostile links, downloads, and recovery.
  • Use deterministic actors for stable flows and model control only where uncertainty requires it.

FAQ

Should every browser task use an AI model?

No. Use deterministic Playwright code for known flows. Add model-selected actions when layouts or navigation vary.

Is a screenshot enough for grounding?

Usually not. Pair visual context with accessibility or DOM information so the agent can target elements precisely.

How many steps should an operator get?

Set a small limit based on the task, then increase it only after measuring successful completion and repeated-state failures.

When should a human take over?

Before irreversible actions, sensitive disclosures, uncertain identity or payment details, and any state the verifier cannot prove.