How to Use an AI Agent to Screenshot a Webpage from a Remote Browser
Capture a webpage with an AI agent using MCP, a screenshot API, or remote Playwright. Choose the right workflow, save the image, and handle dynamic pages safely.
To screenshot a webpage from a remote browser, either call a screenshot endpoint when you already know the URL and capture settings, or connect an AI agent to a browser session when it must navigate, interact, or wait for page state to change. Use MCP if your agent client supports it; use remote Playwright or Puppeteer when you need code-level control. For a one-shot capture, an API call is usually the shortest path.
A screenshot shows pixels. It does not give an agent dependable element references or page text. For interaction, first inspect a structured page snapshot, act on the page, take a fresh snapshot after state changes, and capture a screenshot when you need to verify the visual result.
1. Choose the right remote screenshot workflow
| Approach | Use it when | What you manage |
|---|---|---|
| Direct screenshot API | The URL and capture options are known and no interaction is required. | Request options, credentials, and saving or returning image bytes. |
| MCP screenshot tool | Your assistant or agent client already connects to MCP and you want to request a capture in natural language. | MCP server configuration and the tool’s input/output behavior. |
| Remote Playwright or Puppeteer | The agent needs to navigate, click, type, log in, handle a dialog, or wait for dynamic content. | Browser connection, page state, wait strategy, and session cleanup. |
Browserless documents a direct screenshot endpoint that does not require a WebSocket connection. Its guidance recommends a browser connection when interaction or waiting for dynamic content is necessary. Its examples also cover an MCP server and remote Playwright or Puppeteer connections. These are distinct workflows: use the smallest one that meets the task.
2. Capture a page with an MCP agent
If your MCP-compatible assistant is already configured with a remote browser MCP server, ask it to capture the target page and save or return the image. Browserless documents an agent flow that invokes its browserless_smartscraper tool for a full-page capture and asks the assistant to save the result as an image. The server requires a Browserless API token; configure it in the MCP client or runtime rather than including a real token in prompts or logs.
In your MCP-enabled assistant:
1. Use the remote browser tool to open https://example.com.
2. Capture the full page as a PNG.
3. Save the returned image to screenshot.png.
The exact tool arguments and returned image format depend on the MCP server and client. Check that server’s documentation for its configuration and output handling. MCP is the low-code route when the client already supports it; it does not remove the need to decide whether the page needs interaction or how sensitive screenshot data should be handled.
3. Take a one-shot screenshot through a remote screenshot API
When no browser interaction is needed, send the URL and capture options to a screenshot endpoint. Browserless documents a POST request to its /screenshot endpoint with a JSON body containing url and options. The following runnable examples use a placeholder token; provide your own valid token through an environment variable.
cURL
export BROWSERLESS_TOKEN='YOUR_API_TOKEN'
curl -X POST "https://production-sfo.browserless.io/screenshot?token=${BROWSERLESS_TOKEN}" \
-H 'Content-Type: application/json' \
--data '{"url":"https://example.com","options":{"fullPage":true,"type":"png"}}' \
--output screenshot.png
Python
import os
import requests
url = "https://production-sfo.browserless.io/screenshot"
token = os.environ["BROWSERLESS_TOKEN"]
payload = {
"url": "https://example.com",
"options": {"fullPage": True, "type": "png"},
}
response = requests.post(
url,
params={"token": token},
json=payload,
timeout=90,
)
response.raise_for_status()
with open("screenshot.png", "wb") as image_file:
image_file.write(response.content)
Node.js
const token = process.env.BROWSERLESS_TOKEN;
if (!token) throw new Error("Set BROWSERLESS_TOKEN first");
const response = await fetch(
`https://production-sfo.browserless.io/screenshot?token=${encodeURIComponent(token)}`,
{
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
url: "https://example.com",
options: { fullPage: true, type: "png" },
}),
},
);
if (!response.ok) {
throw new Error(`Screenshot request failed: ${response.status} ${await response.text()}`);
}
const image = Buffer.from(await response.arrayBuffer());
await import("node:fs/promises").then(({ writeFile }) => writeFile("screenshot.png", image));
These examples follow the Browserless documentation’s screenshot request shape and use its documented example endpoint. Verify the current endpoint and account token requirements in the provider documentation before deployment. The endpoint examples return image bytes; handle them as binary data rather than decoding them as text.
4. Connect an AI agent to a remote Playwright browser
Use a connected browser when the agent must establish page state before capture. The cycle is: navigate, inspect a structured snapshot, plan, interact, inspect again after a change, and then take the screenshot. A fresh snapshot helps the agent avoid acting on stale references after a dialog closes, a page navigates, or content updates.
Playwright (Node.js)
Install the Playwright package and configure a remote browser WebSocket endpoint and token according to your hosted browser provider. The URL below is a placeholder; use the provider’s current connection URL and authentication format.
import { chromium } from "playwright";
const endpoint = process.env.REMOTE_BROWSER_WS;
if (!endpoint) throw new Error("Set REMOTE_BROWSER_WS to your provider's WebSocket endpoint");
let browser;
try {
browser = await chromium.connect(endpoint);
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto("https://example.com", { waitUntil: "domcontentloaded", timeout: 60000 });
// Inspect structure or accessibility information before choosing controls.
console.log((await page.locator("body").innerText()).slice(0, 2000));
// Example interaction, if required by the task:
// await page.getByRole("button", { name: "Accept" }).click();
// After an interaction that changes the page, inspect it again before acting.
await page.screenshot({ path: "screenshot.png", fullPage: true });
} finally {
if (browser) await browser.close();
}
For an actual agent, use the browser integration’s structured snapshot or accessibility tree to find controls instead of guessing selectors from a screenshot. Playwright’s guidance distinguishes screenshots, which are for visual inspection, from snapshots, which provide references for interaction. The example prints visible body text only as a minimal runnable inspection step; use the integration’s supported snapshot method for agent-controlled interaction.
Puppeteer alternative
Browserless also documents connecting with Puppeteer, opening a page, navigating, taking a full-page screenshot, and closing the connected browser in a finally block. Configure the WebSocket endpoint for your provider:
import puppeteer from "puppeteer-core";
const endpoint = process.env.REMOTE_BROWSER_WS;
if (!endpoint) throw new Error("Set REMOTE_BROWSER_WS to your provider's WebSocket endpoint");
let browser;
try {
browser = await puppeteer.connect({ browserWSEndpoint: endpoint });
const page = await browser.newPage();
await page.goto("https://example.com", {
waitUntil: "networkidle2",
timeout: 60000,
});
await page.screenshot({ path: "screenshot.png", fullPage: true });
} finally {
if (browser) await browser.close();
}
The vendor examples use networkidle or networkidle2 in some workflows. Treat those as options, not universal guarantees: pages with polling, analytics, streaming, or delayed rendering may never reach a useful idle point. For such pages, wait for a meaningful selector or a bounded delay after the relevant action.
5. Choose viewport, full-page, or element capture
| Scope | What it captures | Good fit |
|---|---|---|
| Viewport | The currently visible browser area. | Visual checks of the initial view, modal, or interaction result. |
| Full page | The full scrollable document. | Long articles, documentation, and page archives. |
| Element or region | A selected element or clipped portion of the page, where supported. | A chart, card, embedded component, or a specific result. |
Browserless documents PNG and JPEG for its screenshot endpoint, and its agent screenshot action supports a full page, selector, or clipped region with those scope choices mutually exclusive. Playwright CLI documents viewport, element, full-page, format selection, and high-resolution capture, including PNG, JPEG, and WebP. In its CLI, PNG is the default when the filename does not imply another format. Check the interface you use: options from a CLI are not necessarily accepted by a hosted API.
For a long page, full-page images may be very tall and harder to inspect or store. Use viewport capture when the agent needs to assess a particular state, and use element capture when only one visual component matters. If a selector is absent or hidden, wait for it or capture the viewport and inspect the page state first.
6. Use screenshots and page snapshots together
Use a structured page or accessibility snapshot to read text and locate controls. Use a screenshot to inspect layout, color, canvas or chart output, and visual regressions. A reliable agent loop is:
- Navigate to the page or requested route.
- Read a structured snapshot and identify the next control or state to inspect.
- Perform one action, such as click, type, select, or scroll.
- Take a new snapshot if the action changed the page or its available controls.
- Capture a screenshot when the task asks for visual evidence or verification.
Playwright explicitly advises using snapshots for interaction references and screenshots for looking at the page. A screenshot can show a button, but it does not provide a robust selector or accessible name to click. If the browser agent cannot determine whether an action succeeded from structure alone, capture the resulting view as a visual check.
7. Protect credentials and captured data
- Put API tokens and browser connection credentials in environment variables or a secrets manager, not source code, prompts, screenshots, or application logs.
- Restrict who can view screenshots: they may contain account details, private messages, customer information, or other sensitive page content.
- Do not log raw image bytes or base64 image payloads. Store captures only where the application has an authorized retention and access policy.
- Close remote browser connections in a
finallyblock after capture so the session is released if navigation or screenshot generation fails. - For a shared or retained agent session, know whether the provider keeps state across tool calls and explicitly end the session when the task is complete.
OpenAI’s computer-use documentation warns that screenshots can contain sensitive page or account data and advises showing them only to authorized users and keeping them out of application logs. Apply the same care to images returned by a remote browser.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 response | Missing, invalid, expired, or incorrectly passed token. | Check the provider’s required token location and confirm the runtime has the intended secret. Avoid printing the token while debugging. |
| HTML or JSON saved with a .png extension | The request failed and the response body is an error message, not image bytes. | Check the HTTP status and content type before saving; inspect the error text without exposing credentials. |
| Blank or partially rendered capture | The screenshot ran before client-side content finished rendering, or the page showed a bot check or access error. | Wait for a page-specific selector or post-interaction state. Confirm the remote browser can access the destination. Do not assume a generic load event means the page is visually ready. |
| Navigation hangs at network idle | Persistent connections, polling, or third-party requests prevent the idle condition. | Use a bounded navigation wait such as domcontentloaded, then wait for a specific element or a deliberate short delay. |
| Click has no effect or targets the wrong control | The agent used stale references, ambiguous text, or guessed from pixels. | Take a fresh structured snapshot after each state change and use a unique accessible name or locator. |
| Element screenshot fails | The selector does not exist yet, is hidden, or matches multiple elements. | Wait for the selector, verify it is visible, and narrow it to a unique element. Fall back to viewport capture to inspect the page. |
| Remote browser session remains open | An exception interrupted normal cleanup or the provider retains sessions by design. | Close the browser in finally; check provider settings for retained-session behavior and explicitly end retained sessions. |
| Image is too large or slow to handle | Full-page capture includes a very long document or large dimensions. | Capture the viewport or a relevant element, and choose a format and resolution supported by the selected interface that fit the downstream use. |
9. Performance, reliability, and cost considerations
No official performance figures are established by the documentation cited here, so do not assume a fixed capture time or success rate. Page behavior, remote browser startup, network conditions, image dimensions, and provider limits all affect completion. Set bounded timeouts, return useful errors to the agent, and avoid retrying indefinitely.
- Reduce unnecessary work: use a direct screenshot request for a known URL; reserve a stateful browser session for tasks that require navigation or interaction.
- Wait for the right signal: a page-specific element is often more useful than waiting for all network activity to stop on dynamic sites.
- Retry carefully: retry transient transport or provider errors with a limit and backoff. Check the result before retrying so a slow completed job does not trigger duplicate work.
- Release sessions: close connected browsers after capture, including on exceptions.
- Budget by provider terms: confirm current quotas, pricing, concurrency, and retention behavior with the service you choose. The research sources do not establish current Browserless pricing or service limits.
- Control output size: choose a suitable scope and format; avoid shipping huge full-page files when a viewport or element capture answers the task.
10. Or skip the browser setup
If you know the URL and do not need an agent to log in or interact first, [ScreenshotNeo](https://screenshotneo.com) can return a screenshot or PDF with one GET request. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
Frequently asked questions
Can an AI agent take a screenshot without controlling a browser?
Yes. If the target URL and settings are known and no interaction is needed, a screenshot API can fetch and return the image directly. Use a browser session when the task depends on page state or actions.
Should an agent use a screenshot or an accessibility snapshot to find a button?
Use a structured or accessibility snapshot to identify and interact with controls. Use the screenshot to inspect the visual appearance or confirm the rendered result.
Does full-page capture include content loaded only after scrolling?
That depends on the browser tool and page behavior. Lazy-loaded content may not exist until the page is scrolled; use the tool’s documented behavior and scroll or wait for the content before capturing when completeness matters.
Which remote browser provider is cheapest or most reliable?
The cited documentation does not establish current prices, quotas, or comparative reliability. Compare current provider terms against your required interaction model, output format, session handling, and data controls.


