ScreenshotNeo

BlogAI agents

How to Build Browser Agents with Playwright and Claude

Build a browser agent with Playwright and Claude using MCP, then choose when Test Agents or the CLI fit better.

By the ScreenshotNeo team4 October 20269 min read

To let Claude operate a browser, connect the Playwright MCP server to an MCP client and give Claude a specific task to complete and verify. For a coding-agent workflow that creates or repairs Playwright tests, use Playwright Test Agents instead. Playwright also documents a CLI workflow for coding agents that need browser automation as one part of a larger repository task.

This guide walks through the MCP route first, then explains the Test Agents and CLI choices, browser state, safety, troubleshooting, and a browser-free option for capturing page screenshots.

1. Choose the browser-agent workflow

Route Use it for Interaction model
Playwright MCP Having Claude explore or operate a website through browser actions Claude calls structured MCP tools and uses page snapshots to identify controls
Playwright Test Agents Planning, generating, and repairing Playwright test suites in Claude Code Agent definitions support a coding and test workflow
Playwright CLI Browser work within a broader coding-agent task The agent runs shell commands and loads relevant skills as needed

Playwright positions MCP for specialized agentic loops and exploratory automation. It positions the CLI for coding-agent work in larger codebases, and says its concise output has lower token overhead than carrying MCP tool schemas and snapshots in context. That is Playwright’s workflow guidance, not an independent performance benchmark. Pick the route based on what Claude needs to do.

2. Connect Playwright MCP to Claude Code

The documented setup requires Node.js 20 or newer and an MCP client. Playwright’s quick setup for Claude Code is:

claude mcp add playwright npx @playwright/mcp@latest

The command registers an MCP server named playwright. When the server runs, npx invokes the Playwright MCP package. The browser downloads automatically on first use according to Playwright’s installation documentation. Check the current Playwright MCP guide and installation documentation if your environment or the command has changed.

  1. Install Node.js 20 or newer if it is not already available.
  2. Run the registration command in a terminal where Claude Code is installed.
  3. Start or reopen the MCP client as needed for it to load the server configuration.
  4. Ask Claude to perform a small, reversible task on a page you are allowed to access.
  5. Review the observed page state and the actions before using the agent on a sensitive account or consequential workflow.

For example, give Claude a bounded task such as: “Open the public demo page, inspect the available controls, add two todo items, and verify that both appear in the list. Do not submit forms or make purchases.” The official Playwright walkthrough uses a demo-page task of this kind to illustrate navigation followed by interaction using an element reference from the page snapshot. Adapt the target to a page you are authorized to use.

Playwright MCP gives the model structured information such as roles, labels, and text in accessibility snapshots. That lets the model target a control by what it is, rather than guessing pixel coordinates. Its documented browser actions include navigation, clicking, typing, filling forms, keyboard and mouse input, dialogs, tabs, and screenshots. It also documents areas such as network monitoring and mocking and browser storage state; an agent does not need every capability for every task.

3. Give the agent a bounded task and verification step

A browser agent is easier to inspect when its goal, permitted actions, and expected result are explicit. Anthropic’s prompting guidance recommends clear, direct instructions and notes that tools for verifying UI work can help Claude. Treat the prompt as task guidance, not a guarantee that the task will succeed.

A useful task prompt has four parts:

  • Target: identify the page or app and the intended starting state.
  • Allowed actions: say which interactions are permitted and which are off limits.
  • Expected result: describe what the completed page should show.
  • Verification: ask Claude to inspect the resulting UI and report what it can actually confirm.

For instance: “Open the public task-list demo. Add ‘Review the release notes’ and ‘Update the sample plan’. Do not delete existing items or navigate away. Inspect the list after adding them and report whether each new item is visible.” This gives Claude an observable completion condition and limits the task’s scope.

4. Choose browser mode and state deliberately

Playwright documents browser-engine and operating-mode choices, along with persistent, isolated, and extension-attached profiles. Choose based on whether the task needs a signed-in session and whether you need repeatable fresh state.

Profile choice What it means Consideration
Isolated A fresh browser context for the task Useful when the task should start without prior cookies or login state
Persistent Retains browser state such as cookies and login state Useful when a task needs an existing session; that session is sensitive
Extension-attached Connects through a browser profile with extensions Consider which extensions and account state the agent can access

Use a headed browser when you need to observe interactions directly and a headless mode where the documented setup and task permit it. The available browser choices and options can change; consult the MCP configuration documentation for current settings. Avoid placing broad access to personal or production sessions in a general-purpose agent profile.

5. Use Playwright Test Agents for test creation and repair

Playwright Test Agents provide three roles:

  • Planner: explores an application and produces a Markdown test plan.
  • Generator: turns a plan into Playwright test files.
  • Healer: replays failing tests, inspects the UI, proposes a patch, and reruns until the tests pass or guardrails stop the loop.

For Claude Code, Playwright documents this initialization command:

npx playwright init-agents --loop=claude

Use this workflow when Claude is working on a test suite, rather than when you want a conversational assistant to operate a browser for an arbitrary task. Playwright says to regenerate the agent definitions when you update Playwright so the definitions receive current tools and instructions. See the Playwright Test Agents documentation for setup and workflow details.

6. Consider the CLI for browser work inside a coding task

When the main job is editing a codebase and browser automation is only one step, Playwright’s CLI workflow may fit better than exposing a full set of MCP tools and snapshots. The agent can invoke shell commands and load the relevant skills as needed. Playwright describes this output as more token-efficient than the MCP interaction model; that is its stated comparison, not a measured result for every project.

Use MCP when structured browser-tool calls are the central interaction surface. Consider the CLI when browser work is one part of a larger coding task. For details and current commands, see Playwright’s CLI documentation.

7. Keep arbitrary code behind a trust boundary

Ordinary structured browser actions and arbitrary JavaScript execution have different risk profiles. Playwright explicitly warns that browser_run_code_unsafe runs arbitrary JavaScript in the Playwright server process and is RCE-equivalent; it says to enable this capability only for trusted MCP clients. Do not enable it just to make a basic navigation or form interaction work. See the warning in the Playwright MCP documentation.

Browser access can expose authenticated sessions and allow actions on accounts. Keep the agent’s task and profile scoped to the work, and require review before irreversible actions such as deleting data, submitting consequential changes, or making purchases. Anthropic’s guidance supports balancing autonomy with safety for actions that are difficult to reverse; it does not establish a security guarantee for a particular setup.

8. Remote browser execution

Playwright documents attaching to an existing browser through a Chrome DevTools Protocol (CDP) endpoint, including for cloud browser services. This can be useful when the browser should run outside the machine hosting Claude. The documentation covered here does not compare providers, latency, cost, geographic coverage, or service terms, so evaluate those separately. See Playwright’s CDP connection documentation.

9. Or skip the browser setup

If the job is to capture a page as an image or PDF rather than interact with its controls, ScreenshotNeo provides a website screenshot API and MCP server. Its API accepts a URL in one GET request and returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

10. Troubleshoot common setup problems

Symptom Likely cause What to do
Claude does not show or call the Playwright tools The MCP server was not registered, the client has not loaded the configuration, or the server failed to start Check the command and server configuration, then restart or reload the client as appropriate. Review the MCP client’s server output for startup errors.
npx or the server cannot run Node.js is missing, too old, or not available on the process path Confirm Node.js 20 or newer is installed and that the terminal running the client can find it.
The first browser action is slow or fails before a page opens The browser may need its first-use download, or the environment may not permit the required download or launch Check the Playwright installation guidance and environment access, then retry after the required browser is available.
A control cannot be found The page has not loaded the expected content, the accessible name differs, or the target is not exposed in the current snapshot Ask Claude to inspect the current page snapshot again, identify the control by its role and label, and verify the page before interacting.
The browser is already signed in or shows unexpected data A persistent profile retained cookies or other browser state Use an isolated context for a fresh run, or explicitly confirm the account and state required before proceeding.
A generated test agent does not reflect the current Playwright tools The Playwright version changed after agent definitions were generated Regenerate the definitions using the documented initialization workflow.
A task stops after an unsafe-code warning The workflow requested arbitrary JavaScript execution, which has a separate trust boundary Use structured actions if they cover the task. Only enable arbitrary-code execution for a trusted MCP client and according to the current documentation.

11. Performance, reliability, and cost

Browser-agent time and reliability depend on the page, browser startup, network, current UI state, and the number of interactions. The research sources do not establish task-success rates, latency benchmarks, or a universally best Claude model, so there is no supported numeric estimate to give. For a bounded task, reduce avoidable work by stating the desired result clearly, using an appropriate browser profile, and asking for verification of the relevant final state.

For recurring test work, distinguish agent exploration and test generation from the test suite’s own execution and maintenance. The Playwright documentation describes the planner, generator, and healer roles, but does not establish a fixed runtime or cost for a project. MCP also carries tool schemas and page snapshots in context; Playwright says the CLI has lower token overhead in its intended coding-agent workflow, without providing a universal benchmark.

Costs depend on the model and execution environment selected by the user; the sources here do not provide a comparable price schedule. If the task is only to capture pages, ScreenshotNeo offers a free allowance of 1,000 shots per month and paid plans starting at $5 for 3,000. A screenshot API does not replace browser interaction when the task requires clicking controls or changing application state.

Frequently asked questions

Does Playwright MCP require Claude Code?

No. Playwright describes it as an MCP server for an MCP client. Claude Code is the client used in the documented setup example.

Can I use the test agents to operate any website task?

The planner, generator, and healer are documented for planning, creating, and repairing Playwright tests. Use the general MCP route for conversational browser operation.

Does an accessibility snapshot guarantee the agent will choose the right control?

No. It gives the model structured page information to work from. Ask it to inspect the page and verify the resulting state, and review consequential actions.

Can I run the browser remotely?

Playwright documents connecting to an existing browser through a CDP endpoint, including cloud browser services. Provider selection requires separate research.

When is a screenshot API a better fit?

Use one when you need a rendered page image or PDF and do not need the agent to interact with page controls. For interaction, a browser agent remains the relevant tool.

Sources