ScreenshotNeo

BlogAI agents

How to Make an MCP Server for Browser Automation

Build a browser automation MCP server with Playwright MCP, understand tools and sessions, and keep browser access controlled.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: The fastest reliable way to make an MCP server for browser automation is to configure Microsoft’s maintained @playwright/mcp server in your MCP client. The client starts the server, the server controls Playwright, and each browser action returns a structured accessibility snapshot that the model can use for the next action. You can then design a smaller custom server around the same principles: narrow tools, explicit arguments, stable semantic targets, clear session ownership, and strict permission boundaries.

This guide covers the documented Playwright MCP setup, the interaction loop, browser and profile choices, standalone HTTP transport, custom-server design, security, troubleshooting, and an alternative for jobs that only need screenshots.

1. Install the prerequisites

The documented quick start requires:

  • Node.js 20 or newer.
  • An MCP-compatible client such as Claude Desktop, Cursor, or another client that can launch MCP servers.

The browser is downloaded on first use. The exact configuration file location depends on your MCP client, so use that client’s MCP setup instructions when adding the server.

2. Configure the Playwright MCP server

Add this server entry to your MCP client’s configuration:

{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest"]
    }
  }
}

This tells the client to launch the published Playwright MCP package as a child process. The package communicates with the client over the MCP transport selected by the client.

Primary references: Playwright MCP getting started and the Playwright MCP introduction.

3. Run a first browser automation task

After restarting or refreshing the MCP client, give it a task such as:

Navigate to https://demo.playwright.dev/todomvc and add a few todo items.

The normal interaction sequence is:

  1. The client invokes a navigation tool with the URL.
  2. The server opens or selects a browser page.
  3. The server returns a structured accessibility snapshot containing roles, names, and element references.
  4. The model chooses a target from that snapshot and invokes an action such as click or fill.
  5. The server returns a new snapshot so the model can verify the resulting state.

For ordinary element targeting, this semantic structure is preferable to guessing screen coordinates. Screenshots and vision capabilities can still be useful for visual tasks, but they are not required for every click or text entry.

4. Understand the MCP server’s tool surface

A useful browser automation server exposes a small set of task-focused tools. A practical baseline is:

Tool category Typical inputs Visible effect
Navigation URL, optional wait condition Opens or changes the current page
Snapshot Optional page or frame identifier Returns accessibility structure and element references
Click Element reference Activates a control or link
Text entry Element reference, text Fills or types into a field
Screenshot Optional page, full-page flag Captures visual output when needed
Tab management Tab identifier or index Creates, selects, or closes pages

Keep each tool’s arguments explicit and validate them before execution. Return enough structured state after every action for the model to decide what happened. Element references can become stale after navigation or a major DOM update, so document that clients should request a fresh snapshot and select a new reference.

5. Choose browser, headless, and profile options

The reference implementation supports Chromium-based Chrome, Firefox, WebKit, and Edge. The important runtime choices are:

Choice Use it when Tradeoff
Headed You are developing locally or need to watch the browser More visible and easier to debug; requires a display
Headless The client runs unattended or on a server Better suited to automation; visual debugging is less immediate
Persistent profile A repeat workflow needs existing cookies or login state Convenient, but saved state is sensitive and persists between runs
Isolated profile Each run should start clean Reduces cross-run carryover, but requires signing in again

Choose persistence deliberately. Cookies, local storage, and cached sessions can contain credentials or account data. For multi-user services, assign ownership explicitly and isolate profiles between tenants.

6. Select additional capabilities only when needed

Playwright MCP keeps basic browser automation in its core tools and makes additional groups opt-in. Documented capability groups include:

  • Vision for workflows that need visual understanding.
  • PDF for generating or inspecting PDF output.
  • DevTools for browser diagnostics.
  • Network for request-level inspection or control.
  • Storage for browser storage operations.
  • Testing for test-oriented workflows.

Enable only the groups required by your use case. Every additional capability expands the actions an agent can request and the amount of state it can inspect. See the capabilities documentation for the documented groups and configuration patterns.

7. Run the server as a separate HTTP process

For local client-launched use, the child-process configuration is simplest. The documentation also shows running the server with a port and connecting an MCP client to its HTTP endpoint:

npx @playwright/mcp@latest --port 8931

Configure the client to connect to:

http://localhost:8931/mcp

This example is a transport setup, not a complete production deployment. A networked service needs authentication, authorization, tenant isolation, network controls, logging, resource limits, and a policy for browser permissions before it is exposed beyond a trusted machine.

8. Design a smaller custom browser MCP server

The official material documents and maintains the Playwright MCP implementation; it does not provide a complete from-scratch custom-server tutorial. If you implement your own server, keep the design narrow and inherit these principles:

  1. Expose task-level tools. Prefer navigate, snapshot, click, and fill over an unrestricted command tool.
  2. Validate every argument. Check URL schemes, string lengths, selector or reference formats, and allowed browser operations before touching the page.
  3. Return structured observations. Include the current URL, page title, relevant accessibility nodes, and a clear error object when an action fails.
  4. Define session ownership. Decide whether one MCP connection owns one browser context, one user owns a persistent profile, or every task receives an isolated context.
  5. Handle stale references. After navigation or a DOM-changing action, return a new snapshot and require the caller to use current references.
  6. Make side effects visible. State whether a tool submits a form, downloads a file, changes account data, or opens a new tab.
  7. Set limits. Bound navigation time, action count, page size, downloads, and concurrent contexts.

A minimal request/response contract can look like this conceptually:

{
  "tool": "click",
  "arguments": {
    "element_ref": "button[Save]"
  }
}

{
  "ok": true,
  "url": "https://example.test/settings",
  "snapshot": {
    "role": "main",
    "children": []
  },
  "warnings": []
}

The exact schema is your implementation decision. The important property is that the model receives a dependable observation after each action instead of an opaque success string.

9. Secure browser control

Browser automation is a high-trust permission. Playwright’s documentation warns: This tool runs arbitrary JavaScript in the Playwright server process and is RCE-equivalent — only enable it for trusted MCP clients. Treat this as an execution boundary, not as an ordinary read-only integration.

Security checklist

  • Allow only trusted MCP clients to start or reach the server.
  • Restrict network access at the host, container, or service layer.
  • Isolate browser contexts and saved profiles between users.
  • Do not treat origin lists or file-access guardrails as complete security boundaries; redirects can work around convenience restrictions.
  • Keep arbitrary JavaScript execution disabled unless the workflow requires it.
  • Handle cookies, local storage, downloaded files, and session directories as credential-bearing data.
  • Apply least-privilege browser and filesystem permissions.
  • Log tool names and outcomes without recording secrets or page contents unnecessarily.

Saved-state and secret-redaction features improve handling and output hygiene, but they do not replace access control. The deployment boundary must enforce who can invoke the server and what the browser can reach.

10. Troubleshooting

Symptom Likely cause Fix
The client shows no Playwright tools Invalid MCP configuration or the client has not reloaded it Validate the JSON, confirm the command is npx with @playwright/mcp@latest, then restart or refresh the client.
Server fails before opening a page Node.js is older than the documented requirement Install Node.js 20 or newer and verify the version used by the MCP client.
Browser download or launch fails First-use browser installation is blocked, or the host lacks required runtime dependencies Run the client in an environment that permits the browser download and install the dependencies required by the selected browser.
An action cannot find its element The reference is stale or the page has not finished changing Request a fresh accessibility snapshot, then use the current reference. Add a page-specific wait where appropriate.
Clicks work locally but not unattended The workflow depends on a headed display or timing Use headless-compatible waits and state checks; run headed during diagnosis.
Login state disappears An isolated profile is being used Use a persistent profile only when retaining that state is acceptable, and protect its storage.
A remote client cannot connect Wrong endpoint, port, firewall rule, or transport configuration Confirm the server is listening on the configured port and that the client targets /mcp. Add authentication before any non-local exposure.
A page is blocked by an allowlist Origin or file-access restrictions are active Review the configured origins, but do not assume an allowlist is a security boundary. Enforce network policy separately.

11. Performance, reliability, and cost considerations

Performance

  • Reuse a browser process when your trust model permits it, while keeping contexts isolated between users.
  • Use headless mode for unattended workloads.
  • Keep snapshots focused on the current page state instead of repeatedly requesting unnecessary visual data.
  • Wait for meaningful conditions such as a selector or state change rather than using long fixed delays everywhere.
  • Limit concurrent pages and downloads so one task cannot exhaust the host.

Reliability

  • After navigation, redirects, or DOM updates, reacquire element references.
  • Return structured errors that distinguish timeout, navigation failure, missing element, permission denial, and browser crash.
  • Make retries idempotent where possible; a repeated click or form submission may not be safe.
  • Record the URL, tool, duration, and outcome for diagnosis while excluding secrets.
  • For long workflows, checkpoint progress outside the browser so a crashed context can be replaced.

Cost

The Playwright MCP package, Node.js runtime, browser process, and MCP client run on infrastructure you control. Your costs come from that infrastructure and any external services used by the workflow. The dossier does not provide benchmark figures, so choose capacity from your page complexity, concurrency, browser choice, and retention requirements rather than an assumed speed number.

12. Or skip the browser setup

If the job is to produce a clean website screenshot rather than interact with controls, ScreenshotNeo gives you a single HTTP request. It removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options. The basic request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const bytes = await res.arrayBuffer();

It supports full-page captures with lazy images loaded, CSS element capture, dark mode, device presets or custom viewports, retina scale, PDF output, custom CSS and JavaScript, clicks before capture, selector hiding, selector or network-idle waits, request blocking, custom headers and cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

Plans include 1,000 free shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account and start with the 1,000 included screenshots.

13. FAQ

Do I need to write an MCP server from scratch?

No. Configure the maintained @playwright/mcp package first. Build a custom server only when you need a narrower tool surface, different session model, or specialized policy.

Are screenshots required for browser automation?

No. Playwright MCP’s documented interaction pattern uses accessibility snapshots and element references for ordinary actions. Add visual capabilities when the task depends on visual layout or pixels.

Should every deployment use a persistent profile?

No. Persistent profiles retain login state and cookies, while isolated profiles reduce cross-run carryover. Select based on whether convenience or clean separation matters more.

Can I expose the HTTP endpoint to the internet?

The documented port option shows how to connect locally. Internet exposure requires authentication, authorization, isolation, network restrictions, and operational controls beyond the transport setting.

When is ScreenshotNeo a better fit?

Use it when you need rendered images or PDFs and do not need an agent to click through a live browser session. Its MCP tools let an AI client request captures without you managing Playwright browser setup.