ScreenshotNeo

BlogAI agents

How to Use Browser Automation with Any Language Model

Connect any language model to a real browser with a controlled observe–act–verify loop, reliable locators, runnable Playwright code, and safety checks.

By the ScreenshotNeo team29 September 202611 min read

How to Use Browser Automation with Any Language Model

To use browser automation with any language model, let the model decide the next step and let a browser runtime execute it. Give the model a compact view of the current page, expose a small set of allowed actions such as navigate, click, fill, and screenshot, then return the result so it can verify the outcome before acting again. Playwright, Selenium, and Puppeteer can all serve as the automation layer; the right fit depends on your language, browser coverage, and existing tooling.

This separation matters: the model proposes actions, but your code controls what the browser can access and what actions require approval. The examples below use Python and Playwright, with a deliberately small tool surface. The same loop works with other model APIs and browser runtimes.

1. The architecture: model, tools, browser, observation

A useful browser agent has five parts:

A browser agent observes page state, takes one allowed action, and checks the result before continuing.
A browser agent observes page state, takes one allowed action, and checks the result before continuing.
  1. Model layer: interprets the user’s goal and selects a next step.
  2. Tool layer: defines narrow operations, for example navigate(url), click(role, name), fill(label, value), or observe().
  3. Automation layer: Playwright, Selenium, Puppeteer, or another WebDriver-compatible runtime turns those operations into browser actions.
  4. Browser runtime: runs the browser and holds cookies, credentials, and network access under your application’s control.
  5. Observation loop: returns an accessibility snapshot, selected page data, or a screenshot for the model to inspect and verify.

Keep the model-facing tools smaller than the browser API. Avoid giving a model arbitrary code execution, unrestricted shell access, or a generic “do anything” browser function unless you have a separate policy layer that constrains it. The model’s output should be validated against a schema and an allowlist before execution.

For many pages, start with an accessibility snapshot or targeted DOM extraction. It exposes roles, names, and relevant content in a compact, machine-readable form. Use screenshots when layout, canvas content, or visual state is important. Playwright MCP uses structured accessibility snapshots for this kind of interaction: Playwright MCP introduction.

2. Choose a browser automation runtime

Runtime Good fit Considerations
Playwright One API across Chromium, Firefox, and WebKit; JS/TS, Python, Java, and .NET bindings; agent workflows and MCP. Install browser binaries that match the library version. Its locator and auto-waiting model is a good default for new automation.
Selenium Teams already using WebDriver conventions, its language bindings, or its established test ecosystem. Use current APIs and verify generated locators against the live application. Models may suggest removed APIs or brittle patterns.
Puppeteer JavaScript-first automation around Chrome and Firefox, using a high-level API over Chrome DevTools Protocol and WebDriver BiDi. Consider whether its language and browser coverage fit your deployment and maintenance needs.

Compare the tools by language fit, browser coverage, locator and waiting behavior, debugging and tracing, CI parallelism, authentication handling, and how much human approval sensitive actions need. There is no common benchmark in the sources here that supports calling one universally fastest or most reliable.

Primary references: Playwright overview, Playwright language bindings, Selenium guidance for AI agents, and Puppeteer documentation.

3. Build a minimal Playwright agent loop in Python

This runnable example shows the browser side of the loop. It opens a page, observes its accessibility snapshot, and uses a semantic locator to fill a field and click a button. It deliberately does not call a particular model API: model SDKs differ, but the browser tools can be wrapped in your provider’s function/tool-call format.

Install and prepare the browser

python -m venv .venv
source .venv/bin/activate
pip install playwright
playwright install chromium

On a Linux CI machine that lacks browser system packages, Playwright documents installing dependencies with playwright install-deps chromium or the corresponding install option. Re-run browser installation after upgrading the Playwright package so the browser binaries remain compatible. See Playwright browser installation.

Browser tool implementation

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto("https://example.com", wait_until="domcontentloaded")

        # Return compact, structured state to the model.
        snapshot = await page.locator("body").aria_snapshot()
        print("PAGE STATE:\n", snapshot)

        # In an agent, the model would choose an allowed action from this state.
        # These actions are illustrative; example.com has no form to submit.
        title = await page.title()
        print("TITLE:", title)
        await browser.close()

asyncio.run(main())

To make this an agent, define tools with explicit arguments and return values. For example, fill_field(label, value) should find a label, fill only that field, and return the updated snapshot. A click_button(name) tool should locate a button by accessible name, click it, and report the resulting URL and relevant page state. Reject unknown tool names and malformed arguments; do not execute raw model-generated Python.

Model-facing loop

Your model client typically sends the user goal, current observation, tool schemas, and policy. When it returns a tool call, validate it, execute the matching browser function, then send the result back to the model. Continue until the model returns a final response or a step limit is reached:

while not task_done and steps < MAX_STEPS:
    state = await browser.observe()
    proposal = await model.plan(goal, state, allowed_tools, policy)
    action = validate_tool_call(proposal, allowed_tools)

    if action.is_sensitive and not policy_approves(action):
        result = {"error": "Approval required", "state": state}
    else:
        result = await browser.execute(action)

    if result.get("error"):
        state = {"error": result["error"], "page": await browser.observe()}
    else:
        state = result
    task_done = await model.verify(goal, state)

Adapt observe, execute, and the model calls to your chosen SDK. OpenAI’s computer-use guide, for example, documents JavaScript/Playwright and Python/PyAutoGUI approaches in a shared console, with a function tool that can return text or images: computer-use guide. That is one implementation path, not a requirement; the browser tool boundary is portable across model providers.

4. Make actions reliable

Prefer semantic locators

Target controls by role and accessible name, label, placeholder, or a stable test ID. These communicate intent better than long CSS paths or positional XPath selectors:

await page.get_by_role("button", name="Continue").click()
await page.get_by_label("Email address").fill("person@example.com")
await page.get_by_placeholder("Search").fill("browser automation")
await page.get_by_test_id("save-settings").click()

Use framework auto-waiting and web-first assertions rather than inserting fixed sleeps. A delay such as await page.wait_for_timeout(3000) may be too short on a slow run and waste time on a fast one. Wait for a specific state instead:

from playwright.async_api import expect

await page.get_by_role("button", name="Save").click()
await expect(page.get_by_text("Saved")).to_be_visible()

For a page transition, wait for a meaningful URL or element. If the page is a single-page application, the URL may not change, so verify the expected state in the DOM. Playwright’s migration guidance recommends Locator objects and web-first assertions and notes that explicit waits are often unnecessary: Playwright migration guidance.

Return useful errors and verify state

When an action fails, return the actual exception plus a relevant snapshot or URL to the model. Do not ask it to guess from the original prompt. Bound retries: retry transient navigation or timeout failures only when safe, and stop when the page shows an unexpected state. Record the action, result, duration, and a redacted observation so a developer can diagnose the run.

5. Add safety boundaries before connecting real accounts

A browser agent can submit forms and change state. Separate low-risk navigation and reading from actions with lasting effects, such as purchases, sending messages, deleting records, or changing account settings.

  • Restrict navigation to approved domains where practical.
  • Keep credentials and secrets in application-managed storage; never place them in model-visible page text or prompts.
  • Redact tokens, personal data, and sensitive form values from snapshots and logs.
  • Require an explicit approval or policy check before irreversible or externally visible actions.
  • Use an action allowlist, a maximum step count, and a retry cap.
  • Log tool calls and results with enough context for audit, while excluding secrets.

These are implementation guardrails, not a universal policy supplied by a browser library. Your application owns the authorization rules and the consequences of each action.

6. Screenshots and visual observation

Structured state is usually a more compact observation channel for ordinary controls. Screenshots help when the relevant information is visual, the page uses a canvas, or you need a visual confirmation. With Playwright, capture a page or a specific element:

await page.screenshot(path="page.png", full_page=True)
await page.get_by_role("main").screenshot(path="main.png")

Sending a screenshot to a vision-capable model costs additional image processing and may expose information visible on the page. Crop or redact where appropriate, and use a DOM snapshot when that fully answers the question. A screenshot is an observation; it does not itself establish that a click succeeded. Verify the resulting page state.

7. Performance, reliability, and cost

Browser work has several distinct costs: browser startup and memory, page loading and scripts, model calls, and any remote browser or model service charges. The research sources do not establish a universal runtime benchmark or model cost, so measure your own workflow with representative pages and concurrency.

  • Reuse browsers carefully: launching a fresh browser for every step adds startup overhead. Reuse a browser process where appropriate, while keeping each task’s context and cookies isolated.
  • Limit observations: send a focused accessibility snapshot or selected DOM region instead of the whole page on every turn. This reduces model input and avoids returning irrelevant content.
  • Choose readiness conditions: domcontentloaded, a target locator, or network idle have different behavior. Pages with polling or long-lived connections may never become network-idle, so wait for the state your task needs.
  • Control concurrency: parallel browser jobs can improve throughput, but consume more memory and may trigger site rate limits. Set a concurrency cap and respect the target site’s policies.
  • Pin versions: pin the library and browser versions in CI and install matching browser binaries. A version drift can change behavior or prevent launch.
  • Capture traces on failures: use framework debugging and tracing options to inspect what happened, especially when a failure cannot be reproduced locally.

Do not assume retries are free or harmless: a repeated click can submit a form twice. Make actions idempotent where possible, inspect the page after a timeout, and retry only when the state proves the first action did not take effect.

8. Troubleshooting

Symptom Likely cause Fix
Browser executable not found Playwright package installed but its matching browser binary is missing. Run playwright install chromium; install system dependencies in Linux CI as documented by Playwright.
Browser fails after package upgrade Browser binaries and library versions are out of sync. Re-run browser installation after upgrading and pin versions in CI.
Locator times out Wrong accessible name, a hidden control, an unexpected page, or content that has not loaded. Inspect the current accessibility snapshot and URL; update the semantic locator and wait for the relevant state.
Click succeeds but nothing appears to happen The action may have changed application state without navigation, or the click target was wrong. Check a success message, expected field value, or other postcondition. Return that state to the model.
Generated code uses stale APIs The model may have learned an older Selenium or framework pattern. Provide current documentation and runnable examples, execute a small smoke script, and validate APIs against your installed version.
Page load hangs on network idle Analytics, polling, or persistent connections keep network activity open. Wait for DOM content or a task-specific locator instead of network idle.
Agent repeats a destructive action Timeout recovery retried without checking whether the first action completed. Check current state before retrying; require approval for irreversible actions and cap retries.
Model cannot identify a control from a screenshot The control may be easier to locate semantically, or the screenshot lacks enough context. Return an accessibility snapshot or targeted DOM data; use a cropped screenshot only when visual details matter.

Selenium’s AI-agent guidance specifically warns about obsolete APIs, arbitrary sleeps, manual driver downloads, and copied XPath selectors. It recommends current documentation, runnable examples, changelogs, and checking locators against the running application: Selenium AI-agent documentation.

ScreenshotNeo removes supported consent banners, popups, and chat widgets before capture.
ScreenshotNeo removes supported consent banners, popups, and chat widgets before capture.

9. Or skip the browser setup

If the job is to obtain a website screenshot for an agent or a pipeline, ScreenshotNeo provides a one-request screenshot API and an MCP server. This skips running and maintaining a browser for that capture. It does not replace interactive browser automation for clicking through a workflow or filling forms.

Install only what you need from ScreenshotNeo’s API documentation. The following cURL, Python, and Node.js examples request a WebP screenshot of Stripe:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, and failed loads are not billed, and cache hits cost nothing. Responses identify page outcomes and billing in headers. An MCP server exposes screenshot, page-info, and PDF tools to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

10. Frequently asked questions

Can I connect any language model, or only one vendor?

Any model client that can choose tools or return structured output can fit the pattern. Keep provider-specific code in the model adapter and keep browser operations behind your own tool interface.

Should the model see the entire page source?

Usually not. Start with an accessibility snapshot or a targeted extraction, then request more context only when the task needs it. This keeps the observation focused and reduces exposure of unrelated page content.

Can an agent safely make purchases or send messages?

Only if your application explicitly authorizes that action. Add a policy check and human approval for consequential actions, and verify the final state before reporting completion.

When should I use screenshots instead of page structure?

Use screenshots for visual layout, canvas-rendered content, or visual confirmation. Use roles, names, and targeted page data for ordinary controls and text.

Does ScreenshotNeo automate interactive browser workflows?

It provides website screenshots, page information, PDFs, and MCP tools. Use Playwright, Selenium, or Puppeteer when the task needs a sequence of interactive browser actions.