ScreenshotNeo

BlogAI agents

Building Browser Agents on a Free Plan

Build a browser agent locally with Playwright, then decide when a hosted free tier is worth the tradeoffs. Includes runnable code, safety checks, and troubleshooting.

By the ScreenshotNeo team29 September 202610 min read

Building Browser Agents on a Free Plan

You can build a useful browser agent for free by running Playwright locally. Your code supplies the plan and guardrails; Playwright opens a real browser, reads the page, interacts with it, and checks the outcome. You pay no hosted-browser usage fee, though your machine, network, and time still have costs.

Start with local Chromium for an unauthenticated, narrowly scoped task. Add a hosted browser only when deployment, uptime, or concurrent sessions are the actual constraint. A free hosted plan is an allowance, not an unlimited browser farm: Cloudflare Browser Run currently lists 10 browser minutes per day and three concurrent browsers on Workers Free. See its pricing details.

1. What a browser agent needs

A browser agent is not just a language model with a click tool. The reliable unit is a loop that observes, acts, and verifies. Keep the components explicit so you can inspect why a task succeeded, failed, or stopped.

A dependable browser agent observes the page, takes a bounded action, and verifies the resulting state.
A dependable browser agent observes the page, takes a bounded action, and verifies the resulting state.
  1. Planner: turn the request into a short sequence of allowed actions and a clear stop condition.
  2. Observer: collect the current URL, visible text, accessible controls, and relevant page state.
  3. Executor: use browser actions such as navigation, clicks, typing, uploads, and waits.
  4. Verifier: re-read the page and check that the expected state appeared. Capture a screenshot when visual confirmation matters.
  5. Recovery: stop when the page is ambiguous; retry only actions that are safe to repeat.

This is close to the capabilities in VS Code’s browser-tool documentation: navigate, understand content and accessible elements, interact, verify with screenshots, and use focused Playwright code for complex flows.

2. Install Playwright locally

The Playwright CLI requires Node.js 20 or newer. Use the official initializer and install only the browser engine you need. The initializer creates a project workspace and downloads the configured browser if needed; the exact prompts can vary by CLI version.

node --version
npm init playwright@latest

Choose JavaScript if prompted, and accept the default test setup or select the option that best fits your project. To install Chromium explicitly:

npx playwright install chromium

For a minimal script project without the test scaffolding, you can install the library directly instead:

mkdir browser-agent
cd browser-agent
npm init -y
npm install playwright
npx playwright install chromium

Playwright supports Chromium, Firefox, and WebKit. Keep the Playwright package and browser binaries together: each package version expects compatible browser revisions. Use bundled Chromium for a low-friction prototype; add Firefox or WebKit if cross-engine behavior is part of the task. You can also control installed Chrome or Edge channels, but branded browsers are not installed by default and enterprise policies can interfere. See the browser documentation.

3. A complete local browser-agent example

The following script is runnable with Node.js and Playwright. It demonstrates the browser loop against a public search page: navigate, observe, act, and verify. It does not call a paid model API. Replace the fixed plan with a model-generated plan only after you have defined an allowed-domain list, action limit, and stop rules.

// agent.mjs
import { chromium } from 'playwright';

const startUrl = 'https://example.com/';
const allowedHosts = new Set(['example.com', 'www.example.com']);
const maxActions = 3;

function assertAllowed(urlString) {
  const url = new URL(urlString);
  if (!['http:', 'https:'].includes(url.protocol) || !allowedHosts.has(url.hostname)) {
    throw new Error(`Navigation outside allowed hosts: ${url.hostname}`);
  }
}

const browser = await chromium.launch({ headless: true });
const context = await browser.newContext(); // isolated, in-memory session
const page = await context.newPage();

try {
  assertAllowed(startUrl);
  await page.goto(startUrl, { waitUntil: 'domcontentloaded', timeout: 30000 });

  // Observe: collect a small amount of relevant page state.
  const before = {
    url: page.url(),
    title: await page.title(),
    text: (await page.locator('body').innerText()).slice(0, 1500),
  };
  console.log('OBSERVE', JSON.stringify(before));

  // Plan: this fixed plan is intentionally narrow and predictable.
  const plan = [{ action: 'read-title' }];
  if (plan.length > maxActions) throw new Error('Action limit exceeded');

  // Act: perform only the approved operation.
  for (const step of plan) {
    if (step.action === 'read-title') console.log('TITLE', await page.title());
    else throw new Error(`Unsupported action: ${step.action}`);
  }

  // Verify: check a simple, deterministic postcondition.
  const heading = await page.locator('h1').first().textContent();
  if (!heading?.trim()) throw new Error('Verification failed: no h1 found');
  console.log('VERIFIED', heading.trim());
  await page.screenshot({ path: 'result.png', fullPage: true });
} catch (error) {
  console.error('Agent stopped:', error.message);
  // Preserve a diagnostic artifact where possible; it may contain sensitive data.
  if (!page.isClosed()) await page.screenshot({ path: 'failure.png' }).catch(() => {});
  process.exitCode = 1;
} finally {
  await browser.close();
}

Save it as agent.mjs, then run node agent.mjs. The example does not ask a model to decide arbitrary selectors or URLs. For a real agent, define a narrow task contract first: allowed domains, allowed actions, maximum actions, timeouts, and what counts as success. Prefer role- and label-based locators over brittle CSS selectors when interacting with controls.

Adding model planning safely

If you add an LLM, pass it the observed page summary and ask for structured actions from a small schema, such as click_role, fill_label, and stop. Validate every proposed action in ordinary code before execution. Reject unknown actions, out-of-scope URLs, oversized text input, and steps beyond the action limit. A model response is a suggestion, not authorization to perform a payment, delete data, or bypass an anti-bot check.

4. Sessions, credentials, and human checkpoints

A fresh Playwright context has its own session state. Do not assume it has the cookies or sign-in state of your regular browser. VS Code documents that an agent-opened page uses an isolated in-memory session; explicitly sharing an existing tab is a different choice and can expose that tab’s cookies, storage, and signed-in state. Make this distinction visible in your agent’s design.

An agent’s isolated session should be a deliberate choice, especially when authentication is involved.
An agent’s isolated session should be a deliberate choice, especially when authentication is involved.
  • Keep passwords, API keys, and session tokens outside prompts and logs.
  • Do not print cookies, authorization headers, full form contents, or sensitive page text to routine logs.
  • Pause for a person at payments, account recovery, destructive actions, and anti-bot checkpoints.
  • Treat screenshots and downloaded page content as potentially sensitive. Restrict where artifacts are written and how long they are retained.
  • Use a separate context per user or task when isolation matters, and close it when the task ends.

For authenticated work, provide credentials through a controlled secret mechanism and decide explicitly whether the agent should receive a fresh login session or a deliberately shared session. Never silently inherit a developer’s personal browser state.

5. When to use a free hosted browser

Local Playwright is a good starting point when you can run the task on your own machine or a machine you control. A hosted browser becomes useful when the agent must run on a schedule, remain available while your laptop is off, or serve multiple jobs. Cloudflare Browser Run offers hosted headless Chrome and documents Playwright, Puppeteer, and CDP connections; for agent browsing it also lists Playwright MCP or CDP with MCP clients, and Stagehand for intent-based element discovery. Start at the Browser Rendering documentation.

Cloudflare’s current pricing page lists 10 browser minutes per day and three concurrent browsers on Workers Free. Browser sessions consume browser time and concurrency, so measure actual use before relying on this allowance. The service announced Browser Rendering availability on Workers Free in its April 7, 2025 changelog.

Decision factor Local Playwright Hosted Browser Run
Recurring browser charge No hosted-browser meter; uses your machine and network Free allowance is limited; paid usage follows plan terms
Concurrency Bound by your machine’s CPU and memory Workers Free lists three concurrent browsers
Deployment You manage runtime, browser install, and availability Browser is hosted; your code still needs deployment and credentials
Browser version You choose and update package and binaries together Follow provider-supported runtime and connection options
Network and data Requests originate from your environment Requests originate from the hosted environment; assess data and egress needs
Best first use Prototype, development, low-volume jobs Scheduled work or a need for remotely available execution

Compare the total cost, not just the browser line item: deployment, retries, network egress, logs, storage, and engineering time all count. Check the live provider terms before production use because quotas and policies can change.

6. Reliability, performance, and cost controls

Wait for the condition you need

Choose navigation waits deliberately. domcontentloaded is often a faster starting point than waiting for every network connection to finish, but a dynamic application may need a specific locator to become visible. Prefer waiting for a meaningful selector or state over fixed sleeps. Set navigation and action timeouts, and make the overall task deadline shorter than the job runner’s deadline.

Retry only safe steps

Navigation and reading are usually safe to retry. Submitting a purchase, sending a message, creating a record, or deleting an item may not be. Before retrying a state-changing action, verify whether it already succeeded. If the state is unclear, stop and ask for human review rather than blindly repeating it.

Keep the browser lightweight

Run headless for unattended jobs, reuse a browser process where appropriate, and create separate contexts for isolation. Avoid loading extra tabs or launching more concurrent contexts than the machine can support. Capture screenshots only when they help verify or diagnose the task; full-page images and verbose DOM snapshots increase processing and storage.

Budget the hosted allowance

For a hosted free tier, record browser duration and concurrency per task. Include failed attempts and retries in the estimate, then compare the measured daily total with the published allowance. If jobs queue or exceed the quota, reduce unnecessary navigation and repeated work, schedule tasks across the day, or evaluate a paid plan. Local execution avoids a hosted browser-minute charge but still consumes compute and network resources.

7. Troubleshooting common failures

Symptom Likely cause Fix
Browser executable missing Browser binaries were not installed for the current Playwright version Run npx playwright install chromium; keep package and browser versions in sync.
Navigation timeout Slow site, blocked request, or waiting for a load event that never settles Set a bounded timeout, use domcontentloaded where suitable, then wait for the specific content needed.
Locator finds nothing Page changed, control is inside a frame, or the selector is brittle Re-observe accessible roles and labels; inspect frame structure; wait for the expected state.
Agent is signed out New context is isolated and has no existing cookies Use an explicit, approved login or session handoff. Do not assume another tab’s state is shared.
Works locally, fails in deployment Different browser binaries, missing OS dependencies, network policy, or environment variables Install the matching browser and dependencies in the deployed runtime; compare versions and network access.
Task repeats an operation Retry logic cannot tell whether a prior state-changing action completed Check the postcondition before retrying; require idempotency or human confirmation.
Hosted jobs queue or stop Concurrency or daily browser-time allowance reached Measure session duration and concurrency, reduce overlapping work, and review the current plan allowance.
Screenshot contains private data Artifacts were treated as harmless diagnostics Limit capture, access, retention, and log output; remove sensitive artifacts from shared storage.

8. Or skip the browser setup

If your agent only needs a screenshot or PDF, ScreenshotNeo can return one with a GET request instead of making you install and run a browser. It is a website screenshot API and MCP server from ScreenshotNeo. See the API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, no card required.

9. A practical build checklist

  1. Install Node.js 20 or newer, initialize Playwright, and install only the browser engines you need.
  2. Write down the task’s allowed domains, action budget, timeout, and success condition.
  3. Implement observe → plan → act → verify, and log action names and results without secrets.
  4. Test unauthenticated tasks first; add a deliberate session handoff for signed-in tasks.
  5. Use retries only for safe repeatable steps; stop at payments, destructive actions, and anti-bot checks.
  6. Measure runtime and concurrency locally. Consider Browser Run when hosting is the constraint, and compare usage with its current free allowance.

FAQ

Can I build a browser agent without paying for an AI model?

Yes. The example uses a fixed plan and Playwright only. You can also use a locally available model, but model hosting and usage have separate costs and setup.

Does Playwright use my normal Chrome profile?

Not by default. Playwright contexts are isolated. Treat sharing an existing signed-in tab as an explicit security choice.

Is Browser Run unlimited on the free plan?

No. The current Cloudflare pricing page lists daily browser minutes and concurrent-browser limits. Recheck the provider’s pricing before depending on a quota.

Can a screenshot API replace a browser agent?

Only for screenshot and page-capture work. A screenshot endpoint does not perform arbitrary multi-step interactions such as completing a workflow or verifying a state change.