MCP Server for Browser Control: Setup, Sessions, Security, and Screenshots
Learn how an MCP browser-control server connects AI clients to Playwright, manages sessions, and captures reliable screenshots safely.
Direct answer: An MCP server for browser control exposes browser automation tools to an AI client such as Claude, Cursor, VS Code, or another MCP-compatible application. The client sends structured actions to the server; the server drives a browser with Playwright and returns page state, accessibility snapshots, or action results. Playwright MCP is a documented example. It normally uses accessibility snapshots for page understanding instead of requiring screenshots for every step. Read the Playwright MCP documentation.
What an MCP browser-control server does
The Model Context Protocol (MCP) standardizes how an AI application discovers and calls tools. A browser-control server can expose tools for opening URLs, inspecting page content, clicking controls, filling forms, selecting options, taking screenshots, and reading browser state.
A typical request path looks like this:
- Your MCP client sends a tool call, such as “open this URL” or “click the Sign in button.”
- The MCP server translates that call into Playwright browser operations.
- The browser loads the page and performs the action.
- The server returns structured page information, commonly an accessibility snapshot, plus the result of the action.
Playwright MCP keeps ordinary page understanding structured. Screenshots remain useful when you need visual verification, pixel output, layout inspection, or an image artifact, but they do not have to be the only representation sent to the model.
Prerequisites
- Node.js 20 or newer.
- An MCP client that supports custom servers. The Playwright guide lists VS Code, Cursor, Windsurf, Claude Code, Claude Desktop, and other clients.
- A browser that the server can launch or reach.
Client configuration changes over time, so check your client’s current MCP configuration format before publishing a shared setup.
Run Playwright MCP locally
The documented standard launch command uses npx:
npx @playwright/mcp@latest
Add that command to your MCP client’s server configuration. A generic JSON shape looks like this; the exact file location and property names depend on the client:
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
Playwright MCP starts in headed mode by default. For a worker, CI job, or server without a display, add --headless:
npx @playwright/mcp@latest --headless
Browser selection can be changed when you need a specific engine. The documented choices include Chromium-based Chrome, Firefox, WebKit, and Microsoft Edge. Confirm the current option spelling in the version you install before putting it in automation.
Expose only the capabilities you need
Playwright MCP’s capabilities setting controls which tools are exposed to the model. Basic browser automation remains available. Restricting optional capabilities makes the tool list easier for an agent to reason about and reduces accidental access to operations your workflow does not require.
Choose capabilities by task:
- Read and navigate: page inspection, links, and navigation.
- Interact: clicks, typing, form submission, and selection.
- Validate visually: screenshots and viewport checks.
- Debug: console, network, or tracing features when your client and server version support them.
Choose a browser and session mode
Persistent profile
A persistent profile retains cookies, local storage, and login state between sessions. It is useful for repeated work in the same account. The project documentation says a persistent profile is restricted to one browser instance at a time, so do not start competing processes against the same profile directory.
Isolated sessions
An isolated context starts clean. It is appropriate for reproducible tests and tasks that must not inherit a user’s cookies. In-memory cookies and storage disappear when the context closes unless you explicitly persist state or use storage state.
Extension mode
Extension mode attaches to an existing Chrome or Edge profile. It can reuse an already-open tab, existing cookies, browser extensions, SSO state, or a completed two-factor-authentication flow. Treat that attached profile as sensitive: the model can act with the permissions available to the browser.
Connect to a browser that is already running
The server does not always need to launch a browser. Playwright MCP documents several connection paths:
- Browser channel: connect to a locally installed browser channel.
- Chromium CDP endpoint: connect to a browser exposing the Chrome DevTools Protocol.
- Playwright server endpoint: connect to a separate Playwright server process.
- Browser extension: attach to an existing desktop browser.
CDP endpoints can also point to hosted browser services. This separates the MCP process from the machine where the browser runs, but adds network, authentication, and lifecycle configuration you must operate.
Run the MCP server over HTTP
For a headed browser on a machine without a display, or for an IDE worker process, the guide shows an HTTP server on port 8931:
npx @playwright/mcp@latest --port 8931
Configure the client to connect to:
http://localhost:8931/mcp
HTTP sessions use a five-second heartbeat timeout. If a client or proxy does not answer server-initiated pings, set PLAYWRIGHT_MCP_PING_TIMEOUT_MS to a larger value, or set it to 0 to disable the heartbeat:
PLAYWRIGHT_MCP_PING_TIMEOUT_MS=15000 npx @playwright/mcp@latest --port 8931
A heartbeat setting only controls liveness detection. It does not authenticate the endpoint or isolate the browser.
Security boundaries and deployment hygiene
The Playwright MCP documentation states: “Playwright MCP is not a security boundary.” Shared browser context is also described as a convenience rather than a security boundary.
Before attaching an account or exposing an HTTP endpoint, review:
- Which websites the browser can reach, including internal networks.
- Which cookies, saved sessions, extensions, and credentials are present.
- Whether the model can submit forms, upload files, delete data, or make purchases.
- Which optional MCP capabilities are enabled.
- How the endpoint is authenticated and protected from other users.
- Whether logs or tool results could contain secrets or personal data.
Use a dedicated profile for automation, isolate high-risk accounts, and require human review for irreversible actions. Network rules and context options can reduce exposure, but they do not turn a browser context into a complete security boundary.
Browser-control workflow for screenshots
- Start an isolated context when the page must be anonymous and reproducible.
- Navigate to the target URL.
- Wait for the page state your capture needs, such as a selector or network idle.
- Dismiss consent dialogs or other overlays that block the content.
- Capture the viewport or full page.
- Inspect the result and save it with a deterministic filename.
For full-page captures, account for lazy-loaded images and content that appears only after scrolling. For repeatable output, fix the viewport, browser engine, locale, timezone, and authentication state.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF output. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for the complete option list. The same endpoint supports full-page and element captures, dark mode, device presets, custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data, and an OpenAPI specification.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo has an MCP server with take_screenshot, get_page_info, and capture_pdf tools, so Claude, Cursor, and other MCP clients can request captures directly. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
npx cannot find the package |
Old Node.js, blocked registry access, or a transient install failure | Use Node.js 20 or newer, verify npm connectivity, and retry with the documented package command. |
| Browser opens but the client sees no tools | Invalid client configuration or the server process exited | Run the command in a terminal first, inspect stderr, then validate the client’s server name, command, and arguments. |
| Headless launch fails on a server | No display or missing browser dependencies | Use --headless, install the browser dependencies required by your environment, and check the server logs. |
| Login state disappeared | An isolated context was used | Use a persistent profile or explicitly save and restore storage state. |
| Profile is locked | Another process is using the persistent profile | Stop the competing browser instance or use a separate profile directory. |
| HTTP client disconnects after inactivity | The five-second heartbeat was not answered | Fix proxy ping handling or set PLAYWRIGHT_MCP_PING_TIMEOUT_MS to a suitable value. |
| Screenshot contains a consent banner | The page requires an interaction your workflow did not perform | Locate the banner in the accessibility snapshot, click the appropriate control, then capture after it disappears. |
| Screenshot is blank or incomplete | Capture occurred before rendering, lazy loading, or navigation finished | Wait for a selector, a delay, or network idle; scroll when lazy content requires it; verify the final URL. |
| ScreenshotNeo response is not an image | The target failed, timed out, triggered a bot check, or returned another verdict | Inspect the HTTP status and X-Page-Verdict/X-Billed headers before saving or displaying the body. |
Performance, reliability, and cost
- Startup: launching a browser per request adds process and page startup time. Reuse a controlled browser session for batches when your isolation requirements allow it.
- Parallelism: separate contexts can improve throughput, but each consumes CPU and memory. Keep concurrency below the point where pages contend for resources.
- Determinism: fix viewport, browser engine, locale, timezone, geolocation, fonts, authentication state, and wait conditions when comparing images.
- Retries: retry navigation failures with a limit and inspect the final URL. Do not blindly repeat form submissions or other side effects.
- Remote browsers: CDP and hosted endpoints add network latency and another service to monitor, but keep browser compute away from the MCP process.
- Screenshot billing: ScreenshotNeo bills only clean shots. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and verdict headers let you classify responses.
- Plans: Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan.
FAQ
Does MCP itself control a browser?
No. MCP defines the connection between a client and tools. A server such as Playwright MCP implements the browser operations.
Can I use an existing logged-in browser?
Yes. Extension mode and browser connection options can reuse an existing Chrome or Edge profile, cookies, extensions, or tab. Treat that profile as sensitive.
Are screenshots required for browser understanding?
No. Playwright MCP’s documented approach uses structured accessibility snapshots for ordinary inspection. Use screenshots when visual output or layout verification is required.
Should I expose an MCP browser server publicly?
Only with deliberate authentication, network controls, isolated profiles, and a review of the tools and accounts it can reach. The server is not a security boundary.
Can an AI agent capture PDFs without launching Playwright?
Yes. ScreenshotNeo’s MCP server includes capture_pdf, and its API supports PDF paper size, margins, landscape mode, and page ranges.


