ScreenshotNeo

BlogAI agents

How to Use Web Search MCP Servers with Browser Automation

Connect web search MCP tools to Playwright so an AI agent can find pages, inspect them, interact with browsers, and capture reliable results.

By the ScreenshotNeo team1 October 20269 min read

Use two MCP servers together: a web-search MCP server finds candidate pages, and Playwright MCP opens those pages and interacts with them through browser automation. The model searches first, navigates to the selected result, reads a structured accessibility snapshot, then uses the returned element references to click, type, scroll, or submit forms.

Playwright MCP is the browser side of this workflow. It is an official Model Context Protocol server built on Playwright and gives an LLM browser automation capabilities through structured accessibility snapshots. The minimum setup is Node.js 20 or newer plus an MCP client configured to launch npx @playwright/mcp@latest. See the Playwright MCP introduction and MCP documentation for the current client and capability references.

1. Understand the architecture

MCP is the interface that exposes tools to the model. Playwright supplies the browser automation implementation. A search MCP server supplies search tools. They can run in the same client configuration, but they solve different problems:

Component Responsibility Typical result
Web-search MCP server Find pages, documents, or answers for a query Search results with URLs, titles, and snippets
Playwright MCP Open a URL and operate a real browser Accessibility snapshot and browser action results
MCP client Connect the model to one or more servers Tool list and tool-call routing
Browser profile Stores cookies, sessions, and permissions New isolated context or an existing Chromium session

The normal loop is:

  1. Call the search server with a focused query.
  2. Choose a result URL and call Playwright browser_navigate.
  3. Call browser_snapshot to obtain the current accessibility tree.
  4. Use the element references in that snapshot, such as e5, in the next click, fill, or check action.
  5. Take another snapshot after every meaningful state change.

2. Install the prerequisites

Node.js

Install Node.js 20 or newer on the machine that will launch the local server. Confirm the version:

node --version
npx --version

An MCP client

The official documentation lists VS Code, Cursor, Windsurf, Claude Code, Claude Desktop, and other MCP-compatible clients. Each client has its own location for MCP configuration, but the server entry is the same.

3. Add Playwright MCP to your client

Add this server entry to the client’s MCP configuration:

{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest"]
    }
  }
}

Restart or reload the client, then ask it to navigate to a page and inspect it. A useful first request is: “Open https://example.com, return the page title, and describe the main links.” The client should invoke navigation, receive a snapshot, and use the snapshot’s references for follow-up actions.

Adding a search MCP server

Search servers use different package names and authentication settings, so copy the installation entry from that server’s documentation. Keep the two servers conceptually separate. A generic combined configuration looks like this:

{
  "mcpServers": {
    "search": {
      "command": "npx",
      "args": ["YOUR_SEARCH_SERVER_PACKAGE"],
      "env": {
        "SEARCH_API_KEY": "YOUR_SEARCH_KEY"
      }
    },
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest"]
    }
  }
}

Replace the search package, environment variable, and key name with the values documented by your chosen search server. Do not put secrets in prompts or commit them to a repository.

4. Run a search-to-browser workflow

Search, select, inspect

Give the model an explicit sequence so it does not browse blindly:

Find three primary sources about WebAssembly component model security.
For each result, record the title and URL.
Open the most authoritative result in the browser.
Take an accessibility snapshot and summarize the headings and links.
Do not submit forms or download files.

The search tool returns candidate URLs. Playwright then loads the chosen URL. The browser snapshot exposes roles, visible text, and element references. An interaction should refer to the latest snapshot reference:

1. browser_navigate(url="https://example.com/docs")
2. browser_snapshot()
3. browser_click(ref="e5")
4. browser_snapshot()

References are state-dependent. After navigation, filtering, or a modal opening, take a fresh snapshot rather than reusing an old reference.

Search result validation

Before trusting a result, ask the agent to verify the final page:

  • Confirm the hostname and expected path.
  • Check that the page heading matches the requested topic.
  • Prefer primary documentation, specifications, and source repositories.
  • Record redirects and stop if the destination is unexpected.

5. Use the browser capabilities safely

Core browser automation is always enabled. Optional capability groups can be enabled with --caps=... or the equivalent client setting:

Capability group Use it for Enable when
Vision Visual page understanding Accessibility data cannot express the needed visual relationship
PDF PDF generation or inspection The workflow explicitly handles PDFs
DevTools Console and debugging information You need runtime diagnostics
Network Requests and responses You need to inspect or control network activity
Storage Cookies and local storage The workflow requires a controlled session
Testing Test-oriented browser actions You are building repeatable checks
Configuration Additional server configuration The deployment needs non-default behavior

Start with core capabilities and add only what a workflow needs. A smaller tool surface makes it easier for an agent to choose the correct action.

6. Connect to an existing Chrome session

By default, Playwright MCP can manage its own browser. To attach to an existing Chromium session, provide a Chrome DevTools Protocol endpoint, for example:

npx @playwright/mcp@latest --cdp-endpoint=chrome

The exact endpoint value depends on how Chromium was started and exposed. Use this mode when the browser already contains a deliberate login or extension state. Treat that profile as sensitive: every tool call can operate with its cookies and permissions.

7. Run a separate HTTP deployment

For a headed or separate deployment, run the server with HTTP transport and point the MCP client at its /mcp endpoint. The precise command-line flags depend on the Playwright MCP version and deployment environment; follow the current transport options in the official documentation. Put authentication, TLS, network access controls, and the browser host inside the same trusted boundary.

Share a browser context between clients

The options documentation describes --shared-browser-context for sharing one browser context. Use it only when clients are intentionally cooperating. Shared state means one agent can see cookies, tabs, and navigation changes made by another.

8. Automate common tasks

Finding and extracting an article

  1. Search for the topic with constraints such as date, domain, or document type.
  2. Open the selected result with browser_navigate.
  3. Take a snapshot and identify the main article region by its heading and landmarks.
  4. Extract the visible text and links.
  5. Follow citations in a new navigation step and repeat validation.

Filling a form

  1. Navigate to the form page.
  2. Snapshot the page.
  3. Use the textbox, checkbox, and button references returned by that snapshot.
  4. Snapshot again after validation messages or page transitions.
  5. Submit only when the user explicitly authorized the action.

Handling infinite scroll and lazy content

Ask the agent to scroll in bounded increments, snapshot after each increment, and stop when the target heading or link appears. Set a maximum number of scrolls and a time limit. Never assume that a page has finished loading because the first snapshot is short.

9. Capture a result without managing a browser

If your goal is a clean image or PDF rather than interactive browser control, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Its consent step accepts cookie banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Or skip the browser setup

Use the API when you need a repeatable capture instead of an interactive browser session. Full API options and parameter names are in the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper sizes and page ranges, HTML/CSS-to-image, custom JavaScript and CSS, clicks, selector or network-idle waits, ad and tracker blocking, custom headers and cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Only clean shots are billed. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

10. Security and trust boundaries

The Playwright MCP documentation warns that the tool can run arbitrary JavaScript in the Playwright server process and is RCE-equivalent. Enable it only for trusted MCP clients. Keep the server, MCP client, browser profile, and target sites inside a trusted boundary.

  • Use a dedicated browser profile for automation.
  • Do not attach an authenticated personal profile unless the workflow requires it.
  • Limit which clients can reach an HTTP MCP endpoint.
  • Keep API keys in environment variables or the client’s secret store.
  • Review navigation and form-submission instructions before allowing side effects.
  • Use capability allowlists and enable optional groups only when required.

11. Reliability and performance

Make actions deterministic

  • Search with a narrow query and explicit source constraints.
  • Validate the hostname after redirects.
  • Snapshot after navigation and after every state-changing action.
  • Use semantic roles and accessible names instead of coordinates.
  • Bound retries, scrolls, downloads, and total runtime.

Reduce latency

  • Keep the search result set small.
  • Reuse one browser context when the workflow needs the same session.
  • Use core capabilities unless an optional group is required.
  • Cache stable research results outside the browser.
  • For repeated static captures, use ScreenshotNeo caching with a TTL you choose.

Plan for failure

Pages can redirect, require a login, render content only after JavaScript runs, or block automation. Record the last successful URL, snapshot, and action. Retry navigation with a bounded policy, then report the failure instead of repeatedly clicking stale references.

12. Troubleshooting

Symptom Likely cause Fix
npx cannot find the package Node.js is missing, too old, or the client cannot access npm Install Node.js 20 or newer, verify node --version, and retry from a network that can reach the package registry.
The Playwright tools do not appear Malformed MCP JSON or the client has not reloaded configuration Validate the JSON, restart the client, and inspect its MCP logs.
An element reference is rejected The page changed after the snapshot Take a new browser_snapshot and use the new reference.
The page is blank Navigation failed, content is blocked, or rendering is still in progress Check the current URL, wait for the expected landmark, inspect console or network capabilities if enabled, and retry once with a bounded timeout.
Search finds a page but the browser reaches another site Redirect or tracking URL Verify the final hostname before extracting content.
Login state disappeared A new isolated browser context was created Use the intended persistent profile or CDP attachment, and protect that profile as sensitive.
HTTP clients cannot connect Wrong transport endpoint, blocked port, or missing network policy Point the client to the server’s /mcp endpoint, confirm reachability, and secure the connection.
ScreenshotNeo response is not billed The page was a bot check, blank page, timeout, failed load, or cache hit Read X-Page-Verdict and X-Billed; adjust waits, headers, or target URL when a real capture is required.

13. A practical checklist

  • Node.js 20 or newer is installed.
  • The search and Playwright servers are separate, clearly named entries.
  • The client was restarted after configuration changes.
  • Search results are validated before navigation.
  • Each browser action uses a reference from the latest snapshot.
  • Optional capabilities are limited to the workflow.
  • Automation runs in a trusted boundary and dedicated profile.
  • Navigation, retries, and side effects have explicit limits.
  • Use ScreenshotNeo when the deliverable is a clean screenshot or PDF rather than interactive browsing.

FAQ

Is Playwright MCP a search engine?

No. It controls a browser. Pair it with a web-search MCP server when the agent must discover URLs.

Does the model receive pixels or page structure?

The default workflow provides a structured accessibility snapshot with roles, text, and element references. Optional vision capabilities can support workflows that require visual understanding.

Can multiple MCP clients use one browser?

Yes, the options include a shared browser context, but shared cookies and navigation state require a deliberate trust boundary.

When should I use an API instead of browser automation?

Use browser automation for interactive research and actions. Use ScreenshotNeo when you need repeatable screenshots or PDFs without maintaining browser setup, especially when consent banners and other overlays must be removed before capture.