How MCP Servers Connect to Web Scraping Actors
Learn how MCP hosts call Playwright browsers and Apify Actors, with architecture, code, security guidance, troubleshooting, and a ScreenshotNeo shortcut.

Direct answer: An MCP server connects an AI host to web scraping Actors by exposing browser, crawler, or hosted-Actor operations as MCP tools. The host creates an MCP client for the server, discovers available tools, sends a typed JSON-RPC tools/call request, and receives extracted content or run metadata over the same MCP connection. Playwright MCP places a browser you operate behind that tool interface. Apify MCP places hosted Apify Actors behind it.
The protocol is the adapter between the model and the scraper. It gives the AI application a stable contract for inputs and outputs while the server handles browser lifecycle, credentials, retries, rate limits, proxy policy, and result storage. That separation is implementation guidance based on the MCP architecture; MCP itself does not require a particular scraper or hosting model.
1. The MCP scraping architecture
The official architecture has three parts:

- MCP host: the AI application, such as Claude, Cursor, or another agent runtime.
- MCP client: a connection created by the host for each MCP server.
- MCP server: a process or service that publishes tools, resources, and prompts.
MCP uses a JSON-RPC data layer and a transport layer. Local servers commonly communicate over standard input/output (stdio). Remote servers commonly use Streamable HTTP, which can support authentication and streaming. See the MCP specification for the protocol details.
Request flow
- A user asks the AI host for information from a website.
- The host’s MCP client discovers tools with the server’s tool-list operation, then selects a suitable scraper.
- The client sends a JSON-RPC
tools/callrequest with typed arguments such as a URL, selector, search query, or Actor input. - The MCP server invokes its execution backend: a Playwright browser, an Apify Actor, or another crawler or API.
- The server normalizes the result into MCP content. The host presents it, stores it, or performs a follow-up action.
{
"jsonrpc": "2.0",
"id": 7,
"method": "tools/call",
"params": {
"name": "scrape_page",
"arguments": {
"url": "https://example.com/products",
"selector": ".product-card"
}
}
}
The exact tool name and argument schema come from the server. A well-designed server publishes a JSON schema so the model can supply valid values instead of guessing.
2. Playwright MCP: a browser as the scraping Actor
Playwright MCP provides browser automation through structured accessibility snapshots. An LLM can identify an element by role, name, text, and reference instead of relying on screenshot coordinates. The documented workflow supports navigation, clicking, typing, forms, screenshots, and JavaScript execution.
In this design, the MCP server is the tool adapter and Playwright is the browser engine. It is a good fit when a target needs JavaScript rendering, clicks, pagination, an authenticated session, or content that appears only after interaction.
Install and run a local server
npx @playwright/mcp@latest --browser chromium
Configure your MCP host with a stdio server entry. The shape varies by host; this generic example shows the important fields:
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest", "--browser", "chromium"]
}
}
}
Playwright MCP supports Chrome, Firefox, WebKit, and Microsoft Edge. You can run headed or headless, preserve login state with a persistent profile, or use isolated sessions. Optional capability groups add network, storage, PDF, DevTools, and testing functionality; enable only what the workflow needs.
Typical agent interaction
- Ask the agent to open the target URL.
- Request an accessibility snapshot so the model can see links, headings, buttons, and form controls.
- Call navigation, click, fill, or press tools using the returned element references.
- Wait for the relevant content, then extract text or take a screenshot.
Playwright’s arbitrary JavaScript execution is equivalent to remote code execution. Run a browser-capable MCP server only for trusted MCP clients, restrict who can connect, and isolate profiles for untrusted jobs.
3. Apify MCP: hosted Actors as callable tools
Apify MCP exposes hosted Apify Actors through a Streamable HTTP endpoint at mcp.apify.com. The server can discover Actors, run them, and access run outputs and storage. Its documented defaults include apify/rag-web-browser and apify/web-fetch; you can configure specific search, social, maps, or e-commerce Actors.
The adapter loads an Actor’s input schema and exposes that Actor as an MCP tool. The model therefore supplies typed Actor inputs without a bespoke integration for every scraper. RAG Web Browser can search and scrape top URLs. Web Fetch retrieves a URL with JavaScript rendering and documented anti-bot support.
Connect to the hosted endpoint
{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com",
"headers": {
"Authorization": "Bearer YOUR_APIFY_TOKEN"
}
}
}
}
Running Actors and reading run data require authentication in the documented service. Limited discovery and documentation tools may be available anonymously. Keep tokens in the MCP host’s secret configuration, never in a prompt or scraped page.
Actor call concept
{
"jsonrpc": "2.0",
"id": 12,
"method": "tools/call",
"params": {
"name": "apify_run_actor",
"arguments": {
"actor": "apify/web-fetch",
"input": {
"url": "https://example.com/docs",
"renderJavaScript": true
}
}
}
}
The actual published tool name and fields are discovered from the server. Treat the example as the request shape, not a guaranteed name.
4. Playwright MCP versus Apify MCP
| Axis | Playwright MCP | Apify MCP and Actors |
|---|---|---|
| Execution | Browser process controlled by your MCP server | Hosted Actor execution behind Apify’s endpoint |
| Best fit | Custom navigation, interaction, authenticated sessions, browser-level control | Reusable scrapers, search, site-specific extraction, managed execution |
| Outputs | Snapshots, extracted text, screenshots, traces, browser state | Actor results, datasets, key-value records, fetched content |
| Operations | Your team manages runtime, concurrency, profiles, and deployment | The provider manages Actor runtime; usage, authentication, and storage remain service concerns |
| Transport | Usually local stdio, or HTTP when separately hosted | Hosted Streamable HTTP, with local stdio also documented |
| Main risk | Browser credentials and arbitrary code execution require strict trust boundaries | Tokens, Actor permissions, target-site terms, and data handling require governance |
The first three rows summarize the documented capabilities. The operations and risk rows are practical deployment guidance, not MCP protocol guarantees.

5. Choosing stdio or HTTP
Use stdio when
- The MCP server runs on the same developer workstation as the host.
- You need a simple process boundary and do not need remote sharing.
- Credentials and browser profiles can stay local.
Use Streamable HTTP when
- A shared service must serve multiple hosts or agents.
- You need centralized authentication, audit logs, or streaming responses.
- The scraper runs in a container or separate network.
For HTTP, authenticate the endpoint, apply request timeouts, limit allowed tools and domains, and make concurrency explicit. For stdio, supervise the child process and capture stderr separately from protocol messages.
6. Build a reliable scraping tool contract
- Define inputs: URL, optional selectors, pagination limits, locale, authentication reference, and output format.
- Validate: accept only approved URL schemes and domains; bound page counts, depth, and response size.
- Return structured output: include extracted records, source URL, timestamp, tool or Actor version, and storage identifiers.
- Make failures typed: distinguish navigation timeout, access denied, selector missing, empty result, and backend failure.
- Make retries safe: retry transient network errors with a limit; avoid repeating side-effecting actions blindly.
Record the target URL, tool name, Actor version or configuration, timestamp, and output storage ID so a result can be reproduced.
7. Security, privacy, and site-access checklist
- Treat a browser-capable MCP server as privileged automation. Playwright explicitly warns that arbitrary JavaScript is RCE-equivalent.
- Use isolated profiles for untrusted jobs. Use persistent profiles only when a workflow genuinely needs cookies or login state.
- Keep Actor and API credentials in server configuration, not prompts or scraped output.
- Constrain allowed domains, tools, and Actor names.
- Respect site terms, robots directives, access controls, and applicable privacy law. MCP standardizes invocation; it does not grant permission to collect data.
8. Complete client examples
These examples call an ordinary HTTP scraping endpoint. For an MCP server, your host normally handles JSON-RPC and tool discovery; the same input validation and timeout practices apply.
cURL
curl --fail-with-body --max-time 90 \
-H "Authorization: Bearer $SCRAPER_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","selector":"main"}' \
https://scraper.internal/v1/extract
Python
import os
import requests
payload = {"url": "https://example.com", "selector": "main"}
r = requests.post(
"https://scraper.internal/v1/extract",
json=payload,
headers={"Authorization": f"Bearer {os.environ['SCRAPER_TOKEN']}"},
timeout=90,
)
r.raise_for_status()
print(r.json())
Node.js
const res = await fetch('https://scraper.internal/v1/extract', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.SCRAPER_TOKEN}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({ url: 'https://example.com', selector: 'main' })
});
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
console.log(await res.json());
9. Or skip the browser setup
If your goal is a clean screenshot or PDF rather than extracted records, ScreenshotNeo provides a single website screenshot API request. Its consent step accepts cookie banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing result. It also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));
Use full-page capture with lazy images loaded, CSS element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, click and wait actions, blocked resource types, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTL, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, PDF controls, and the usage API as your workflow requires. Plans include 1,000 free shots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
10. Performance, reliability, and cost
- Rendering time: JavaScript pages, fonts, images, and bot checks add work. Wait for a selector or network idle instead of using an arbitrary long delay.
- Concurrency: cap parallel browser contexts and Actor runs. Queue excess work and apply exponential backoff to transient failures.
- Caching: cache immutable pages and include a content or configuration key in your cache key. Do not cache private pages across users.
- Result size: return only fields the agent needs; store large datasets externally and return a reference.
- Cost: hosted Actors trade infrastructure work for provider usage and storage charges. Self-hosted Playwright trades those charges for your own compute and operations.
- Reproducibility: pin browser, Actor, and tool versions where possible and log all input parameters.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Tool is missing | Server failed discovery or configuration points at the wrong command or URL | Inspect startup logs, verify the MCP entry, and call tool listing again. |
| Browser opens but content is empty | Page is still rendering, blocked, or requires an interaction | Wait for a meaningful selector, inspect the accessibility snapshot, and complete the required click or form step. |
| Selector not found | Responsive markup, shadow DOM, or a changed page | Use roles and accessible names where possible; verify the current DOM and add a bounded fallback. |
| HTTP MCP calls time out | Long Actor run, proxy delay, or server queue | Use an asynchronous run if supported, increase the client timeout, and poll a run ID. |
| Apify authorization fails | Missing, expired, or under-scoped token | Put a valid token in server configuration and confirm the Actor permission. |
| Repeated duplicate records | Retries replayed a non-idempotent action or pagination cursor was ignored | Use an idempotency key, persist cursors, and deduplicate by a stable source identifier. |
| Playwright server is unsafe to expose | Untrusted client can invoke JavaScript or access a logged-in profile | Keep it local or behind authentication, disable unnecessary capabilities, and use isolated profiles. |
12. FAQ
Can MCP run Playwright?
Yes. Playwright MCP is an MCP server that exposes browser actions and snapshots as tools while Playwright runs the browser.
Does MCP itself scrape websites?
No. MCP defines communication and capability discovery. The connected server chooses Playwright, an Apify Actor, or another backend.
Which option handles login state?
Playwright MCP supports persistent profiles and isolated sessions. Apify Actors can implement authentication according to their own input schema and service configuration.
Is a screenshot the same as scraped data?
No. A screenshot is a visual artifact. Scraping returns text, records, or run data. Choose the tool contract that matches the output your application needs.
How should an agent know which scraper to call?
Publish clear tool descriptions and JSON schemas, constrain domains and limits, and return structured errors so the host can select or retry safely.


