What Playwright MCP Includes and How Its Components Work
Learn what Playwright MCP includes, how snapshots and browser tools work, and how to configure secure, reliable AI browser automation.
Playwright MCP is an MCP server that gives compatible AI clients access to Playwright browser automation. The core loop is structured accessibility snapshot → element reference → browser action → updated page state. An MCP client such as VS Code, Cursor, Claude Code, or another compatible application launches @playwright/mcp, and the server controls a browser through Playwright.
The exact tools depend on the package version and your configuration. Documented capability families include navigation, clicks, typing, form filling, keyboard and mouse input, tabs and dialogs, screenshots, network inspection and route mocking, console messages, cookies and storage state, and optional Playwright code execution. The official project also documents tracing, video, and testing-related capabilities. See the Playwright MCP repository, getting-started guide, and configuration reference for version-specific details.
What Playwright MCP includes
| Component | What it does | Decisions you control |
|---|---|---|
| MCP client | Stores server configuration and exposes tools to the AI model. | Client, transport, permissions, and which server commands are allowed. |
| MCP server | Translates MCP tool calls into Playwright browser operations. | Package version, command-line flags, configuration file, and environment variables. |
| Browser | Runs Chromium-based Chrome, Firefox, WebKit, or Microsoft Edge sessions. | Engine, headed or headless mode, viewport, device emulation, proxy, and locale-related settings. |
| Browser context | Holds cookies, local storage, permissions, and session state. | Persistent profile, isolated context, supplied storage state, or an existing-tab connection. |
| Accessibility snapshot | Represents page roles, names, text, and element references in structured form. | When to request a snapshot and which reference to act on. |
| Tool families | Provide navigation, interaction, inspection, capture, storage, and advanced automation operations. | Which capabilities are enabled and visible to the client. |
How the snapshot-based interaction loop works
- Navigate. The agent asks the server to open a URL.
- Inspect. The server returns a structured accessibility snapshot rather than requiring the model to infer every control from pixels.
- Select a reference. The model chooses a button, link, input, or other element reference from the snapshot.
- Act. It calls a click, fill, type, keyboard, mouse, tab, or dialog tool using that reference.
- Observe again. The server returns updated state, and the agent continues until the task is complete.
Screenshots are available for visual verification, but the central documented workflow is snapshot-driven. This can make form and navigation tasks more deterministic because the model acts on roles and names exposed by the page’s accessibility tree.
Install and connect Playwright MCP
Prerequisites
- Node.js 20 or newer.
- An MCP-capable client such as VS Code, Cursor, Claude Code, or another client that supports MCP servers.
- A browser installation supported by your selected Playwright engine.
Run the server directly
npx @playwright/mcp@latest
The getting-started documentation uses this command. Your MCP client normally runs it for you from its server configuration. The browser is headed by default in the documented getting-started flow; add --headless when you need a background process.
Example MCP client entry
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest", "--headless"]
}
}
}
Client configuration formats differ, so adapt the surrounding JSON to your client. Keep the package version and flags aligned with the current official documentation.
First request
A representative first task is: “Navigate to https://demo.playwright.dev/todomvc and add a few todo items.” The assistant should navigate, read the snapshot, identify the input reference, fill it, submit it, and inspect the resulting state.
Browser engines, modes, and contexts
Browser selection
The project documents Chrome, Firefox, WebKit, and Microsoft Edge choices. Select an engine when browser-specific behavior matters, then repeat the workflow in each engine you support. Device emulation and viewport settings are separate configuration decisions.
Headed versus headless
- Headed: shows the browser window, which helps while developing and diagnosing selectors, consent dialogs, and navigation.
- Headless: runs without a visible window and is usually preferable for CI, remote hosts, and unattended agents.
Persistent and isolated profiles
| Mode | State behavior | Use it when |
|---|---|---|
| Persistent profile | Cookies and local storage can survive between sessions. | You need a logged-in workflow or stable browser state. |
| Isolated session | Starts clean and normally loses state when closed unless storage state is supplied. | You need reproducible tests or tenant isolation. |
| Existing-tab extension connection | Connects the server to tabs already open in a browser. | A human has already authenticated or prepared a page. |
Choose the least stateful mode that satisfies the task. Persistent profiles are convenient, but they also retain credentials, consent choices, and data that later tasks may observe.
Tool families and what they are for
Navigation and page interaction
Navigation tools open URLs and move through browser history. Interaction tools click, type, fill forms, press keys, move the mouse, select options, and handle tabs or dialogs. The normal pattern is to obtain a fresh snapshot after an action that changes the page.
Inspection and debugging
- Screenshots: visually verify layout, overlays, and rendering.
- Console messages: inspect page-side errors and warnings.
- Network inspection: identify requests, responses, and failed resources.
- Route mocking: intercept and replace selected network responses.
- Tracing and video: create richer diagnostics when the configured version exposes them.
Storage and authentication
Cookie and storage-state tools let you preserve or inject session data. Treat exported state as sensitive: it can contain reusable login tokens. Use isolated contexts for unrelated accounts and avoid committing storage files to source control.
Advanced code execution
The project documents an unsafe browser code tool. The official warning says: “This tool runs arbitrary JavaScript in the Playwright server process and is RCE-equivalent — only enable it for trusted MCP clients.” Enable it only in a controlled environment where the MCP client and prompts are trusted.
Page-provided WebMCP tools
Some pages can register tools for the current tab. The documentation warns: “Tool names, descriptions, schemas and results are provided by the page, so treat them as untrusted input.” Do not grant page-provided tools automatic authority to access secrets, approve transactions, or change production data.
Configuration and precedence
The configuration guide documents three sources: configuration file, environment variables, and command-line arguments. Follow the documented precedence order for your release, and keep one source authoritative in deployment so operators know which value wins.
| Area | Examples of decisions | Operational effect |
|---|---|---|
| Execution | Headed/headless, browser engine | Visibility, resource use, and browser compatibility. |
| Emulation | Device profile, viewport | Responsive layout, touch behavior, and user-agent characteristics. |
| Networking | Proxy, HTTP transport | Routing, remote access, and deployment topology. |
| Sessions | Persistent profile, isolated mode, storage state | Login persistence and cross-task data isolation. |
| Secrets | Dotenv-based redaction and substitution | Convenience for hiding matching text in responses; it is not a security boundary. |
Complete local workflow example
- Install Node.js 20 or newer.
- Configure your MCP client to run
npx @playwright/mcp@latest --headless. - Ask the client to navigate to a test page.
- Read the returned accessibility snapshot.
- Use the referenced input or button to perform an action.
- Request another snapshot and confirm the resulting state.
- Use a screenshot, console output, or network inspection only when visual or diagnostic evidence is needed.
Navigate to https://demo.playwright.dev/todomvc.
Read the accessibility snapshot.
Add these items using the referenced input: "Buy milk", "Write report", "Review pull request".
Read the updated snapshot and confirm that all three items appear.
This prompt describes the intended interaction loop; the exact tool names and schemas are supplied by the MCP server version and client.
Or skip the browser setup
If your goal is a clean, repeatable screenshot rather than interactive browser control, ScreenshotNeo provides a single GET request for PNG, JPEG, WebP, or PDF output. It accepts consent banners as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports its verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for AI clients.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and custom viewports, retina scale, PDF paper sizes and page ranges, HTML/CSS rendering, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, and a usage API.
There are 1,000 free screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Troubleshooting Playwright MCP
| Symptom | Likely cause | Fix |
|---|---|---|
npx cannot start the server |
Node.js is missing or older than version 20. | Install Node.js 20 or newer and confirm with node --version. |
| No tools appear in the client | Invalid MCP configuration, client not restarted, or server process exited. | Run the command manually, inspect client logs, verify JSON syntax, and restart the client. |
| Browser window is unexpected | Headed mode is the documented default. | Add the release-supported --headless option for unattended runs. |
| Element reference no longer works | The page changed after navigation, a click, or an asynchronous update. | Request a new accessibility snapshot and use the new reference. |
| Login disappears | You are using an isolated context or a profile without supplied storage state. | Use a persistent profile or provide storage state deliberately and securely. |
| Actions are flaky | Race conditions, overlays, slow network, or unstable selectors. | Wait for a meaningful selector or state, inspect the snapshot, and avoid timing-only assumptions. |
| Network requests fail in a private environment | Proxy, DNS, certificate, or firewall configuration. | Verify routing from the server host and configure the documented proxy options. |
| Unsafe code is blocked | The client or server disables the RCE-equivalent tool. | Prefer standard tools; enable code execution only for trusted clients in a controlled environment. |
| Page-provided tools look suspicious | WebMCP definitions come from page content. | Treat names, schemas, and results as untrusted and require explicit approval for sensitive actions. |
Performance, reliability, and cost considerations
Performance
- Headless mode avoids display overhead on servers.
- Reuse a browser process when your client supports it, while keeping unrelated tasks in isolated contexts.
- Use device and viewport settings only when the task needs them.
- Prefer accessibility snapshots and targeted inspection over repeatedly requesting large visual artifacts.
- Use network inspection selectively; tracing, video, and screenshots add work and storage.
Reliability
- Use stable semantic references from fresh snapshots.
- Wait for a selector, navigation state, or application condition instead of arbitrary sleeps.
- Separate authenticated and anonymous profiles.
- Capture console and network evidence when a workflow fails so the next run can distinguish page bugs from automation bugs.
- Pin a tested package version in production and review release notes before changing flags or tool assumptions.
Cost and context
Playwright MCP itself is a software server, so your practical costs are the AI client and model usage, browser CPU and memory, network traffic, and any infrastructure required to run it. The project notes that MCP schemas and snapshots can use more model context than Playwright CLI workflows. Use MCP when an agent benefits from structured, exploratory browser interaction; use shell-oriented automation when a coding agent needs compact commands across a large repository.
Playwright MCP versus Playwright CLI
| Axis | Playwright MCP | Playwright CLI |
|---|---|---|
| Interaction | MCP tool calls and structured snapshots. | Shell commands and scripts. |
| Best fit | Exploratory browser tasks and specialized agent loops. | Coding agents working in larger codebases. |
| Context use | Tool schemas and snapshots can consume more context. | Command-oriented interaction can be more compact. |
| Default mode | Headed in the documented getting-started flow. | Depends on the CLI command and options. |
| Setup | MCP client configuration plus server process. | CLI installation and shell access. |
These are the Playwright project’s own comparison axes, not an independent benchmark.
Security checklist
- Allow only trusted MCP clients to connect to the server.
- Keep
browser_run_code_unsafedisabled unless arbitrary JavaScript execution is required and the environment is trusted. - Never expose persistent profiles or storage-state files to untrusted tasks.
- Redact secrets in logs and remember that dotenv substitution is a convenience, not a security boundary.
- Review page-provided WebMCP tools as untrusted input.
- Run production browser sessions with least-privilege credentials and isolated accounts.
FAQ
Does Playwright MCP use screenshots as its main input?
No. Its central interaction model uses structured accessibility snapshots and element references. Screenshots are available for visual verification.
Which browsers can it control?
The documented choices include Chrome, Firefox, WebKit, and Microsoft Edge. Availability and flags can change with the package version.
Can it keep me logged in?
Yes. Use a persistent profile or deliberately provide storage state. Isolated sessions start clean unless state is supplied.
Is the unsafe code tool safe for arbitrary clients?
No. The official documentation describes it as RCE-equivalent and restricts it to trusted MCP clients.
Why did a page add unexpected tools?
WebMCP tools can be registered by the page. The documentation says their names, schemas, descriptions, and results are untrusted input.
When should I use ScreenshotNeo instead?
Use ScreenshotNeo when you need a clean screenshot or PDF through an API or MCP tool without maintaining browser setup, especially when consent banners, popups, chat widgets, failed loads, or billing predictability matter.


