MCP Server Tutorial: Build a Browser Screenshot Tool
Build an MCP browser tool that opens a URL, captures a screenshot, and returns it to an AI client—with Playwright setup, options, errors, and deployment guidance.
An MCP browser screenshot tool exposes one focused capability to an AI client: open a URL, capture the viewport or full page, and return the image. The client calls an MCP tool, the server validates its arguments, Playwright drives a browser, and the result is returned inline or saved to a file.
This tutorial builds a small Node.js server around Playwright. It also explains when screenshots help, when an accessibility snapshot is better, how to configure headed or headless execution, and how to troubleshoot common failures.
What you are building
The request path is:
- An MCP client discovers a
browser_screenshottool. - The client sends a URL and bounded capture options.
- The MCP server validates the URL and options.
- Playwright opens the page, waits for it to be usable, and captures an image.
- The server returns an image content block or a saved-file reference supported by the client.
Playwright MCP is an official reference for this workflow. Its current getting-started guide requires Node.js 20 or newer and shows an MCP client configuration that runs npx with @playwright/mcp@latest. Check the current guide before publishing or deploying because package tags and defaults can change.
Prerequisites and project setup
- Node.js 20 or newer.
- An MCP client that can launch a local server over stdio.
- Playwright and a browser binary.
mkdir mcp-screenshot-server
cd mcp-screenshot-server
npm init -y
npm install @modelcontextprotocol/sdk playwright zod
npx playwright install chromium
Set your package to use ES modules:
npm pkg set type=module
Client configuration locations differ. The common shape is a server entry with command node and the absolute path to your server file. The Playwright MCP guide also demonstrates the equivalent npx configuration:
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
Implement the screenshot tool
Create server.js. The server below accepts a URL, optional target selector, full-page mode, image type, scale, and output filename. It keeps one browser process alive for multiple calls and creates a fresh context for each call so cookies and viewport settings do not leak between requests.
import { Server } from "@modelcontextprotocol/sdk/server/index.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import {
CallToolRequestSchema,
ListToolsRequestSchema
} from "@modelcontextprotocol/sdk/types.js";
import { chromium } from "playwright";
import { z } from "zod";
import fs from "node:fs/promises";
const inputSchema = z.object({
url: z.string().url(),
target: z.string().min(1).optional(),
fullPage: z.boolean().default(false),
type: z.enum(["png", "jpeg", "webp"]).default("png"),
scale: z.enum(["css", "device"]).default("device"),
filename: z.string().min(1).optional(),
timeoutMs: z.number().int().min(1000).max(120000).default(30000),
waitMs: z.number().int().min(0).max(30000).default(0)
}).superRefine((value, ctx) => {
if (value.target && value.fullPage) {
ctx.addIssue({
code: z.ZodIssueCode.custom,
message: "fullPage cannot be combined with target"
});
}
});
let browser;
const server = new Server(
{ name: "mcp-screenshot-server", version: "1.0.0" },
{ capabilities: { tools: {} } }
);
server.setRequestHandler(ListToolsRequestSchema, async () => ({
tools: [{
name: "browser_screenshot",
description: "Open a URL and return a viewport, element, or full-page screenshot.",
inputSchema: {
type: "object",
properties: {
url: { type: "string", format: "uri" },
target: { type: "string", description: "CSS selector for one element" },
fullPage: { type: "boolean", default: false },
type: { type: "string", enum: ["png", "jpeg", "webp"], default: "png" },
scale: { type: "string", enum: ["css", "device"], default: "device" },
filename: { type: "string" },
timeoutMs: { type: "integer", minimum: 1000, maximum: 120000, default: 30000 },
waitMs: { type: "integer", minimum: 0, maximum: 30000, default: 0 }
},
required: ["url"]
}
}]
}));
server.setRequestHandler(CallToolRequestSchema, async (request) => {
if (request.params.name !== "browser_screenshot") {
throw new Error(`Unknown tool: ${request.params.name}`);
}
const parsed = inputSchema.safeParse(request.params.arguments ?? {});
if (!parsed.success) {
return {
isError: true,
content: [{ type: "text", text: parsed.error.issues.map(i => i.message).join("; ") }]
};
}
const options = parsed.data;
const context = await browser.newContext({ deviceScaleFactor: options.scale === "device" ? 1 : 1 });
const page = await context.newPage();
page.setDefaultTimeout(options.timeoutMs);
try {
await page.goto(options.url, { waitUntil: "domcontentloaded", timeout: options.timeoutMs });
if (options.waitMs) await page.waitForTimeout(options.waitMs);
await page.waitForLoadState("networkidle", { timeout: Math.min(options.timeoutMs, 10000) }).catch(() => {});
const screenshotOptions = {
type: options.type,
fullPage: options.fullPage,
animations: "disabled"
};
let buffer;
if (options.target) {
const element = page.locator(options.target).first();
await element.waitFor({ state: "visible", timeout: options.timeoutMs });
buffer = await element.screenshot({ type: options.type, animations: "disabled" });
} else {
buffer = await page.screenshot(screenshotOptions);
}
if (options.filename) {
await fs.writeFile(options.filename, buffer);
return {
content: [{ type: "text", text: `Saved screenshot to ${options.filename}` }]
};
}
return {
content: [
{ type: "image", data: buffer.toString("base64"), mimeType: `image/${options.type}` },
{ type: "text", text: `Captured ${options.url}` }
]
};
} catch (error) {
return {
isError: true,
content: [{ type: "text", text: `Screenshot failed: ${error.message}` }]
};
} finally {
await context.close();
}
});
browser = await chromium.launch({ headless: process.env.HEADED !== "1" });
const transport = new StdioServerTransport();
await server.connect(transport);
const shutdown = async () => {
await browser?.close();
process.exit(0);
};
process.once("SIGINT", shutdown);
process.once("SIGTERM", shutdown);
Run it directly with node server.js, or point your MCP client at that command. The HEADED=1 environment variable opens a visible browser; the default is headless in this custom server.
Screenshot options and their limits
| Option | Use | Important behavior |
|---|---|---|
target |
Capture one CSS-selected element | Wait for the element to be visible. Do not combine with fullPage. |
fullPage |
Capture the full scrollable document | Useful for long pages; can create very large images. |
filename |
Save to disk | Return the path only after the write succeeds. |
type |
png, jpeg, or webp |
JPEG is smaller but does not preserve transparency. |
scale |
Choose CSS-pixel or device-pixel sizing | Higher pixel density increases memory and transfer size. |
timeoutMs |
Bound navigation and selector waits | Use a finite limit so one page cannot hold the MCP connection forever. |
waitMs |
Allow delayed rendering | Prefer a selector or network condition when possible. |
The official Playwright MCP screenshot tool documents viewport, target-element, and full-page captures, PNG/JPEG/WebP output, optional filenames, and scale selection. Full-page capture cannot be combined with a target. See the screenshots documentation for current option names.
Use screenshots and accessibility snapshots together
A screenshot is a visual artifact. It helps an agent inspect layout, charts, visual regressions, and rendered content. It is not the best representation for clicking or reading every control.
Playwright MCP uses structured accessibility snapshots to expose page roles, names, and references for interaction. A robust loop is:
- Navigate to the page.
- Request an accessibility snapshot when you need to identify a button, link, form field, or other control.
- Use the structured reference to perform the action.
- Capture a screenshot after the action to verify the visual result.
The project documentation explicitly describes screenshots as visual inspection output; use an accessibility snapshot for reliable element references.
Headed, headless, browser choice, and HTTP deployment
Playwright MCP currently runs headed by default in its documented configuration, and supports --headless. It also supports selecting Chromium, Firefox, WebKit, or Microsoft Edge. These settings matter when a site renders differently by engine or when a server has no display.
npx @playwright/mcp@latest --headless
npx @playwright/mcp@latest --browser firefox
npx @playwright/mcp@latest --browser webkit
For a separately launched service, the configuration documentation shows an HTTP server with a local /mcp endpoint. Use that mode when the MCP client and browser process run in different environments, then protect the endpoint with your network controls and authentication layer.
Verification checklist
- Use a stable public demo URL.
- Call the tool with only
urland confirm an inline image is returned. - Repeat with
fullPage: trueand inspect the image dimensions. - Repeat with a known selector such as
main. - Use
filenameand confirm the file exists and is nonempty. - Try a missing selector and confirm the error identifies the selector timeout.
- Run once with
HEADED=1to observe navigation and diagnose visual differences.
Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
| Browser executable not found | Playwright browsers were not installed. | Run npx playwright install chromium or install the engine you selected. |
| Navigation timeout | The page is slow, blocked, or waiting on a resource that never finishes. | Keep a finite timeout, try domcontentloaded, increase the limit for known slow pages, and inspect the URL in headed mode. |
| Target selector timed out | The selector is wrong, the element is inside an iframe, or it appears after an interaction. | Use an accessibility snapshot or devtools to find a stable selector; wait for the required state before capture. |
| Blank or incomplete screenshot | Client-rendered content has not finished, lazy content needs scrolling, or a consent dialog covers the page. | Wait for a stable selector, allow a short delay, scroll before capture when needed, and handle the dialog explicitly. |
| Image too large | Full-page and device-scale captures multiply pixel count. | Capture an element, use CSS scale, choose WebP or JPEG, or split a long document into sections. |
| Client shows no image | The MCP client does not render image content blocks. | Pass a filename and return a path, or use a client that supports MCP image content. |
| Server disconnects | An uncaught exception, process exit, or browser crash ended the stdio process. | Catch tool errors, close contexts in finally, keep browser lifetime outside individual calls, and log diagnostics to stderr. |
Performance, reliability, and cost considerations
- Reuse the browser process. Launching Chromium for every request is expensive. Reuse the process while creating isolated contexts.
- Bound every wait. Network-idle can be delayed by analytics, streams, and long polling. Combine it with a timeout and a more specific readiness selector.
- Control image size. Full-page, high-density captures consume memory and increase MCP message size. Prefer element captures for focused tasks.
- Expect dynamic pages. Ads, animations, personalization, cookie banners, and time-based content can make two captures differ. Disable animations where possible and document the readiness condition.
- Handle retries carefully. Retry transient navigation failures with a cap and backoff; do not blindly repeat actions that submit forms or change state.
- Secure remote operation. A standalone HTTP MCP server should not be exposed publicly without authentication, network restrictions, and request limits.
- Account for context. Playwright’s project documentation positions MCP as useful for persistent browser state and rich page introspection, while CLI plus skills may consume less context in coding-agent workflows. Treat that as project guidance, not an independent benchmark.
Or skip the browser setup
ScreenshotNeo provides a hosted screenshot API and MCP server. One request returns a PNG, JPEG, WebP, or PDF, while the service accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Each cleanup step can be turned off.
Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
It also supports full-page capture with lazy images loaded, CSS-element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work to simplify migration.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Should every MCP interaction return a screenshot?
No. Use structured accessibility data for locating and manipulating controls. Capture an image when visual state matters.
Can I combine a target selector and full-page capture?
No. They describe different capture scopes. Choose the element or the full scrollable page.
Why use a fresh browser context per call?
It isolates cookies, local storage, permissions, and viewport settings between requests while allowing the expensive browser process to remain running.
When is an HTTP MCP server useful?
Use it when the MCP client and browser run on separate machines or in a managed environment. Keep the endpoint private and authenticated.
Is a screenshot a substitute for DOM or accessibility data?
No. Images show pixels; accessibility snapshots provide structured references for actions and text-based reasoning.


