How to Use an MCP Server to Interact With a Browser
Connect Playwright MCP to an AI client, navigate and inspect a page, then interact with browser elements safely. Includes setup, configuration, troubleshooting, and a screenshot API option.

An MCP server lets an AI client use browser automation tools through the Model Context Protocol. With Playwright MCP, the client can navigate to a page, inspect its accessibility snapshot, and act on elements by reference. The basic workflow is navigate → inspect → act → inspect again. Use an MCP client such as Cursor, VS Code, Claude Code, or Claude Desktop, and Node.js 20 or newer. [Playwright MCP getting started]
This guide sets up the local server, completes a first browser task, explains useful configuration choices, and covers security and troubleshooting. It uses Playwright MCP; MCP clients and their configuration locations vary, so check your client’s instructions if its settings differ.
1. What browser interaction through MCP means
MCP (Model Context Protocol) connects an AI application, called the client, to tools exposed by a server. Playwright MCP exposes browser automation tools. The client can ask the server to navigate, inspect page structure, click, type, and perform other browser actions.
Playwright MCP’s main interaction model uses structured accessibility snapshots and element references. This means the assistant can reason about roles and labels, then target a referenced element; it does not have to infer every action from a screenshot. Screenshots are also available when visual inspection is useful.
Use this approach when a task benefits from an iterative browser session: exploring a site, inspecting state after an action, or working with a persistent browser context. For coding-agent workflows where minimizing tool schema and page-context overhead matters, Microsoft’s repository also describes the Playwright CLI plus skills as a more token-efficient option. That is project guidance, not a universal benchmark. [Microsoft Playwright MCP README]
2. Prerequisites and local setup
The current Playwright getting-started guide specifies Node.js 20 or newer and an MCP client. The Microsoft repository README says Node.js 18 or newer; follow the current getting-started requirement for this setup, since repository wording may lag. The browser is downloaded on first use. [Getting started] [Installation]
- Install or update Node.js to version 20 or newer.
- Open your MCP client’s server configuration. The file location and UI differ by client.
- Add the Playwright server configuration below, save it, and restart or reload the client’s MCP connections.
- Accept any prompt to install or download the browser on first use.
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
This uses npx to run the package. The official guide has client-specific examples for VS Code, Cursor, Claude Code, and Claude Desktop, and notes that the common configuration works with other MCP clients too. If your client expects a different configuration format, use its documented equivalent. [Client setup examples]
3. Complete a first browser task
Ask the connected assistant to navigate to the TodoMVC demo and add a few items. The important part is the interaction loop: inspect the snapshot the server returns, target an element reference from that snapshot, and inspect the page again after acting.

- Navigate. Ask the assistant to open
https://demo.playwright.dev/todomvcwithbrowser_navigate. - Inspect. Read the returned accessibility snapshot. Look for the textbox role, its accessible name or description, and the element reference.
- Act. Use the relevant reference with a typing or fill tool. For example, the official setup walkthrough uses
browser_typeagainst a textbox reference to enter a todo. - Submit. Send the Enter key if the app expects it to add the item.
- Verify. Inspect the updated snapshot. Confirm the new list item appears before continuing with another item.
The exact reference identifiers are generated from the live page and can change; do not copy an old identifier into a later session. Ask the assistant to use the current snapshot’s reference. The official getting-started example follows this pattern and shows the new item in the updated snapshot. [First interaction example]
4. Choose browser tools and capabilities
Core browser tools cover navigation, back navigation, accessibility snapshots, text search, click, hover, drag and drop, dropdown selection, typing, key presses, form fill, screenshots, dialogs, file upload, console and network inspection, tab management, and page close or resize. Optional capability groups add functions such as network mocking, storage and authentication, testing, vision, PDF, developer tools, and configuration inspection. [Playwright MCP capabilities]
| Task | Useful capabilities | Practical choice |
|---|---|---|
| Explore and interact with a site | Core tools | Start here; navigate, inspect, act, and verify. |
| Test authenticated flows | Testing plus storage | Enable only when the workflow needs persisted authentication state. |
| Debug a failing page | Developer tools | Use console and network inspection to investigate the failure. |
| Extract data with controlled requests | Network plus storage | Add the groups needed for request behavior and session state. |
| Inspect a visual detail | Vision or screenshot tools | Use visual output when accessibility structure alone does not answer the question. |
Start with core tools and add capability groups for the task. The capability guide says limiting exposed tools reduces schema size and the number of choices presented to the model. [Capabilities guide]
5. Runtime and transport configuration
The configuration guide supports Chrome (the default), Firefox, WebKit, and Microsoft Edge. It also covers headed and headless operation, device and viewport emulation, proxies, profiles, network rules, timeouts, output, optional HTTP transport, and sharing a browser context among connected clients. Choose based on the page and runtime where the server will run. [Configuration options]
Headed or headless
The getting-started guide uses headed operation by default. A visible browser can help during local exploration because you can see what the automation is doing. For a display-less environment or IDE worker, use headless operation and configure it according to the current options guide.
Local process or standalone HTTP server
The standard configuration launches a local process over the client’s configured server connection. For a headless environment or a client that needs to connect over HTTP, the configuration guide shows starting a standalone server on port 8931 and pointing the client to its MCP endpoint:
npx @playwright/mcp@latest --port 8931
{
"mcpServers": {
"playwright": {
"url": "http://localhost:8931/mcp"
}
}
}
Keep the endpoint local unless another machine genuinely needs access. If you expose it beyond localhost, decide who can connect and whether connected clients should share browser state. Shared contexts can also share cookies and other session state, so use them deliberately. [Transport and context configuration]
Profiles, secrets, and browser state
Use a persistent profile only when the task needs browser state to carry across runs. For isolated work, choose a fresh context or profile as supported by the server configuration. Avoid placing credentials in prompts or tool output when you can scope them to the relevant client or workflow.
The configuration guide supports a secrets file that redacts matching plain text from tool responses and substitutes placeholders when typing. The documentation describes this as a convenience, not a security boundary. Treat browser content and tool output as untrusted input, and rely on client-level permissions for actual isolation. Origin lists and file-access guardrails are also documented as convenience defenses rather than security boundaries. [Secrets and guardrails]
6. Security and reliability checklist
- Keep the server’s access narrow. Do not expose an HTTP endpoint to networks that do not need it.
- Be careful with shared contexts. Decide whether clients should see the same tabs, cookies, and profile state.
- Scope credentials. Use only the credentials needed for the task, and do not assume a secrets file isolates them.
- Treat page content as data. A web page can contain instructions, but browser content is not automatically trusted just because it arrived through an MCP tool.
- Enable unsafe code execution only for trusted clients. The
browser_run_code_unsafetool executes arbitrary JavaScript in the Playwright server process and is described as equivalent to remote code execution. Leave it disabled unless the MCP client is trusted and the task requires it. [Getting started security note] - Verify important actions. Inspect the page after submitting forms, changing state, or navigating into a sensitive workflow.
7. Troubleshooting common setup problems
| Symptom | Likely cause | What to do |
|---|---|---|
| The client does not show Playwright tools | Configuration was added in the wrong client file, has invalid JSON, or the client has not reloaded MCP servers. | Validate the JSON, check the client-specific setup guide, then restart or reload its MCP connection. |
npx cannot start the server |
Node.js is missing, too old, or unavailable on the client’s environment PATH. | Install Node.js 20 or newer in that environment and confirm the client can invoke node and npx. |
| First launch stalls or fails before a page opens | The browser download on first use has not completed or failed. | Check the installation output and retry the documented installation flow; ensure the environment can retrieve the browser package. |
| A click or typing call targets the wrong element | The reference came from an old snapshot or the page changed after an earlier action. | Request a fresh accessibility snapshot and use the current element reference and label. |
| A page opens but expected content is absent | The content may be delayed, gated, or different in the current browser state. | Inspect the current snapshot and page state, wait for the relevant content using supported tools, and verify again before acting. |
| The HTTP client cannot connect to the server | The server is not running on the expected port or the client URL does not match its endpoint. | Confirm the server command is active, check port 8931 and the /mcp path, and use the URL shown in the configuration guide. |
| Tools expose more options than the assistant needs | Too many optional capability groups are enabled. | Disable groups not used by the workflow; begin with core tools and add narrowly. |
| A tool response contains sensitive text | Redaction rules do not cover the value, or the secrets file is being treated as a security boundary. | Limit sensitive data at its source, review client permissions, and treat redaction as a convenience only. |
8. Performance, reliability, and cost considerations
The dossier does not provide a universal browser-task benchmark, so execution time and resource use depend on the page, browser engine, network, and enabled workflow. Keep the exposed capability set focused to reduce schema size and model choices. For repeated interactions, inspect only what is needed to make the next decision, and verify state after consequential actions.
For reliability, make the workflow state-aware: take a current snapshot before targeting an element, inspect after navigation or submission, and avoid assuming a page is ready merely because a navigation call returned. When a task fails, use the available console and network inspection tools before retrying blindly. Browser profiles can preserve state, but that convenience makes state management and credential handling part of the design.
Playwright MCP itself is open source; the cited setup documentation does not specify a per-capture service price. The runtime cost is therefore tied to the environment and browser execution you arrange. If the job is simply to produce a website screenshot or PDF rather than interact with a site, a screenshot API can avoid running and maintaining a browser automation session yourself.
9. Or skip the browser setup
If you need a screenshot or PDF rather than interactive browser control, ScreenshotNeo is a website screenshot API and MCP server. One GET request takes a URL and returns an image or PDF. Its capture flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers.

For the full option list and request details, see the ScreenshotNeo API documentation. Here is the one-call cURL example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools. Free includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Every feature is available on every plan.
Create a free account for 1,000 screenshots a month, no card required.
10. FAQ
Can I use a browser other than Chrome?
Yes. The documented choices include Chrome, Firefox, WebKit, and Microsoft Edge. Choose the engine in the server configuration for the task.
Does the assistant interact only through screenshots?
No. The primary workflow uses accessibility snapshots and element references; screenshot tools are available when visual inspection helps.
Can multiple clients share one browser?
The configuration supports sharing a browser context. Decide deliberately whether the clients should share the same tabs and session state.
Should I enable browser_run_code_unsafe?
Only for trusted clients when the task requires arbitrary JavaScript execution in the server process. The documentation characterizes this capability as RCE-equivalent.


