How to Have an AI Agent Capture HTML Pages as JPEG Screenshots Through MCP
Connect Playwright MCP to an AI agent, navigate to an HTML page, and save a JPEG screenshot. Choose viewport, full-page, or element capture.
To have an AI agent save an HTML page as a JPEG through MCP, connect a browser MCP server such as Playwright MCP to your MCP client, ask the agent to navigate to the page, then call browser_take_screenshot with type: "jpeg" and a filename ending in .jpeg. Use fullPage: true for the full scrollable page, or omit it for the current viewport.
This guide uses Playwright MCP as the do-it-yourself route. You need an MCP-compatible AI client, Node.js with npx available, and access to the target page. The exact place to add the server configuration depends on your client.
1. Connect Playwright MCP to your agent
Add this standard server entry to your client’s MCP configuration:
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
Save it using the configuration method documented by your MCP client, then restart or refresh the client if required. Client setup varies; use the current instructions for Cursor, Claude, VS Code, Claude Desktop, or whichever MCP client you use. The Playwright MCP getting-started guide has client-specific examples: Playwright MCP setup.
Playwright MCP exposes browser operations to compatible clients using MCP. For a basic screenshot workflow, its core navigation and screenshot tools are sufficient. You do not need to enable optional coordinate-based vision capabilities just to save an image.
2. Navigate to the page and capture a JPEG
Ask the agent to open the complete URL, wait for the page to load, and then save a screenshot. A minimal screenshot tool call is:
{
"type": "jpeg",
"filename": "capture.jpeg"
}
The corresponding MCP tool is browser_take_screenshot. The tool supports png, jpeg, and webp. Specifying type: "jpeg" makes the requested format explicit; a filename extension can also be used for format inference. See the Playwright MCP screenshot tool reference for the current schema.
A practical agent instruction can be as specific as:
Navigate to https://example.com/report, wait for it to finish loading, then use browser_take_screenshot to save the viewport as capture.jpeg with type set to jpeg.
Replace the example URL with a page you are authorized to access. If you need the complete scrollable page, add fullPage: true:
{
"type": "jpeg",
"filename": "full-page.jpeg",
"fullPage": true
}
A supplied filename is saved to that path; relative paths resolve against the workspace root. If you omit the filename, the tool saves to its output directory and returns the image inline as well, which is useful when the agent needs to inspect the result. Confirm the resulting path in the tool response and open the file from that location.
3. Choose viewport, full-page, or element capture
| Goal | Parameters | What to expect |
|---|---|---|
| Capture what is currently visible | Omit fullPage and target |
Captures the viewport. |
| Capture the entire scrollable page | fullPage: true |
Captures beyond the current viewport. |
| Capture one page element | Set target to a snapshot ref or selector |
Captures the selected element; it cannot be combined with full-page capture. |
To capture one element, first have the agent inspect the page’s accessibility snapshot and identify the target. Then pass the element’s snapshot reference or a selector using the tool’s target parameter. For example, the instruction might be: “Find the main report section in the page snapshot and capture that element as JPEG.” Check the connected tool schema for the exact accepted target shape.
Playwright snapshots provide structured page information such as accessible roles, text, and references. Use a snapshot to locate controls or content reliably, then take a screenshot when you need visual evidence, layout, charts, or image-heavy content. Combining both lets the agent act on structure and verify appearance. See the Playwright accessibility guidance and MCP capabilities documentation for current details.
The screenshot tool also documents a scale choice: CSS scale produces dimensions corresponding to CSS pixels, while device scale produces device-pixel resolution. Check the current tool schema to select the scale supported by your installed version and client. A higher pixel count can make an image larger and slower to handle.
4. Handle common capture issues
| Symptom | Likely cause | Fix |
|---|---|---|
| The Playwright tool is unavailable | The server configuration is missing, invalid, or the client has not loaded it. | Recheck the JSON entry, confirm npx is available in the client environment, and reload the MCP server or client. Follow that client’s current setup instructions. |
| The agent navigates but no JPEG is saved | The screenshot call omitted the type and filename was not a JPEG extension, or the tool call failed. | Set both type: "jpeg" and a filename such as capture.jpeg, then inspect the tool response for its output path or error. |
| The image only shows the top of the page | The default capture is the viewport. | Use fullPage: true when the full scrollable page is required. |
| Element capture fails or targets the wrong content | The selector or snapshot reference is stale, ambiguous, or not available. | Take a fresh accessibility snapshot, identify the element again, and pass its current reference or a more specific selector. Do not combine an element target with fullPage: true. |
| The page looks unfinished | Capture happened before the page or its content completed loading. | Have the agent wait for the relevant content or page state before taking the screenshot. For dynamic pages, identify the element or content that must appear first. |
| The file cannot be found | The filename is relative to the workspace root, or the filename was omitted and the tool used its output directory. | Use an explicit path within the workspace and read the path returned by the tool response. |
5. Reliability, speed, and output considerations
- Wait for the page state you need. A navigation completing does not guarantee that every delayed widget or dynamic section is ready. Ask the agent to wait for the key content before capture.
- Prefer stable targets. Accessibility snapshots help the agent find content by role and text. Refresh the snapshot after navigation or major page changes before reusing a reference.
- Choose scope deliberately. Viewport capture is smaller and quick to inspect; full-page capture includes more content and can produce a much taller image. An element capture focuses the result, but requires a valid target.
- Choose scale for the deliverable. CSS scale suits CSS-pixel dimensions; device scale retains device-pixel resolution. Large outputs take more storage and may take longer for an agent to inspect.
- JPEG is a lossy format. It fits requests that specifically require JPEG. For text-heavy pages, charts, or crisp edges, compare the result with PNG if the format is flexible; Playwright MCP documents PNG, JPEG, and WebP as supported types.
- Account for the execution environment. The browser and generated file are associated with the MCP server/client workspace. A file path on the agent machine is not automatically a public URL or a file on your own computer.
The official setup and tool references do not publish capture speed or cost figures. Runtime depends on the page, browser environment, and capture size; avoid assuming a fixed completion time.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf, so an AI agent can request a screenshot without you configuring a browser automation server. For a JPEG image, use the one-call API with the target URL and JPEG output format. See the ScreenshotNeo API documentation for authentication and current parameters.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=jpeg \
-o shot.jpeg
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "jpeg"},
timeout=90,
)
r.raise_for_status()
open("shot.jpeg", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'jpeg'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.jpeg', Buffer.from(await res.arrayBuffer())));
Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. AI agents can use the MCP server. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month, with no card required.
FAQ
Can I use a .jpg filename instead?
The documented output types include JPEG. Use the explicit type: "jpeg" parameter and a filename such as capture.jpeg to make the requested format clear. Check the current tool schema if you want a different extension.
Do I need a vision model to save the screenshot?
No. The screenshot tool can capture and save the page. Optional coordinate-based vision capabilities are separate; snapshots and standard browser tools are enough for the basic workflow.
Can I capture a page that requires a login?
The browser must be able to reach the page in its current session. Whether that session has the required authentication depends on how your MCP client and browser are configured. Do not assume a fresh browser is signed in.
Can the agent return the image for visual analysis?
Yes. If the filename is omitted, the screenshot tool saves the image to its output directory and returns it inline in the response. With a filename, it saves to the requested path; the agent can then inspect the returned image or access the file as supported by the client.


