How to Capture Web Page Screenshots with an AI Agent Using a Browser Pool
Choose an MCP browser tool, screenshot API, or hosted Playwright session based on whether your agent must interact with the page before capture.
An AI agent can capture a web page through a browser pool in three main ways: use an MCP browser tool when it must navigate or interact first, call a screenshot API when a URL and options are enough, or connect Playwright or Puppeteer to a hosted browser when your application already controls the browser steps. Choose the capture target and wait conditions explicitly: a viewport, a full page, an element, and a clipped region are different outputs.
For a straightforward URL-to-image request, ScreenshotNeo is the first service to consider: cookie banners, popups, and chat widgets are removed before capture; only clean shots are billed; and its lowest paid plan is $5 for 3,000 shots. Its parameter names used by other screenshot APIs also work, which can make switching easier.
1. Choose the browser-pool route
| Route | Use it when | What your code receives |
|---|---|---|
| MCP browser agent | The agent needs to inspect the page, click controls, dismiss a dialog, or change state before capture. | Depending on the tool, an image, a saved file, or an encoded image payload. |
| Screenshot REST API | You know the URL and capture settings; no prior interaction is required. | Usually image bytes in the response. |
| Playwright or Puppeteer over a hosted browser | Your application already owns navigation and page-state logic in code. | Image bytes or a file written by the library. |
| Provider SDK | You prefer a provider’s wrapper and its supported wait and file-output options. | Typically image bytes, optionally saved to a path. |
Browserless documents all of these patterns. Its MCP browser agent is suited to interactive work; its screenshot endpoint is suited to direct captures; and its hosted browser connection can be driven by Playwright or Puppeteer. Provider details can change, so check the linked documentation before deploying.
Do not use a screenshot as the only interface for every agent task. If the agent needs to locate and operate page controls, a DOM snapshot or browser interaction tool may be more useful; capture an image when visual appearance or a visual artifact is the goal.
2. Pick the capture target
| Target | Meaning | Typical use |
|---|---|---|
| Viewport | The visible browser area at the current scroll position. | What a person sees without scrolling. |
| Full page | A capture extending over the page height. | Archiving or reviewing a complete article or landing page. |
| Element | The rendered bounds of a selected element, often addressed by CSS selector. | A card, chart, or component. |
| Clip rectangle | A rectangular region specified by position and dimensions. | A known region of a page or a controlled crop. |
Options vary by integration. Browserless’s screenshot API documents full-page, selector, and clip capture; its browser-agent screenshot tool documents the same target modes and says those modes are mutually exclusive. Playwright has its own page screenshot options. Do not assume a setting for one provider or library works in another.
Browserless Agent Run is a separate response format: its documented screenshot field is a viewport PNG encoded as base64, capped at 5 MiB (5,242,880 bytes). That is a limit for that field, not a general browser-pool limit. If you need a long full-page image, use a route that explicitly supports full-page capture.
3. MCP browser agent: interact, then capture
Use this sequence when the page must be acted on first:
- Configure the browser-pool MCP server in your MCP client, using the provider’s documented endpoint and token. Store the token in the client’s secret or environment-variable facility.
- Ask the agent to navigate to the URL and inspect a page snapshot.
- Have it perform only the required interaction, such as accepting a consent prompt or opening a particular tab.
- Ask it to capture the viewport, full page, element, or clip supported by the tool.
- Save or inspect the returned image or file, and confirm it shows the intended state.
Navigate to https://example.com/pricing, inspect the page, and capture a full-page screenshot. Save the returned image as pricing.png. If a consent dialog blocks the page, handle it before capturing.
The exact MCP configuration and tool invocation depend on the client and browser-pool provider; do not copy credentials into a prompt, source file, screenshot, or logs. Browserless documents one-shot sessions by default and an option to keep a session across calls. Use a persistent session only when a later call needs the state created by an earlier call.
Browserless’s agent documentation exposes navigation, snapshots, interaction, and screenshot operations through an MCP-compatible agent. See Browserless documentation for current setup details. A tool call succeeding does not prove the page reached the intended visual state.
4. Direct REST API with Browserless
For a direct URL capture, Browserless documents a POST request to /screenshot with a URL and Puppeteer-style options. The following examples use placeholders; consult the provider’s current API reference for the exact endpoint host, authentication parameter, and request schema for your account.
cURL: full-page PNG
curl -X POST "https://YOUR_BROWSERLESS_HOST/screenshot?token=$BROWSERLESS_TOKEN" \
-H "Content-Type: application/json" \
--data '{"url":"https://example.com","options":{"fullPage":true,"type":"png"}}' \
--output page.png
Python: save response bytes
import os
import requests
endpoint = "https://YOUR_BROWSERLESS_HOST/screenshot"
response = requests.post(
endpoint,
params={"token": os.environ["BROWSERLESS_TOKEN"]},
json={
"url": "https://example.com",
"options": {"fullPage": True, "type": "png"},
},
timeout=90,
)
response.raise_for_status()
with open("page.png", "wb") as image_file:
image_file.write(response.content)
Node.js: save response bytes
import { writeFile } from "node:fs/promises";
const endpoint = new URL("https://YOUR_BROWSERLESS_HOST/screenshot");
endpoint.searchParams.set("token", process.env.BROWSERLESS_TOKEN);
const response = await fetch(endpoint, {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({
url: "https://example.com",
options: { fullPage: true, type: "png" },
}),
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) throw new Error(`Screenshot request failed: ${response.status}`);
await writeFile("page.png", Buffer.from(await response.arrayBuffer()));
Keep the token out of committed code. The examples demonstrate the request pattern; confirm the provider’s current host and schema before use. Browserless documents PNG, JPEG, and WebP output, along with viewport dimensions, device scale factor, quality, clip, and selector options.
5. Connect Playwright to a hosted browser
This is a useful route when your application must control navigation, wait logic, and page state. Browserless documents remote browser connections with Playwright and Puppeteer. The example below illustrates Playwright’s connection pattern; use the websocket endpoint and authentication form specified by your provider account.
import { chromium } from "playwright";
const browser = await chromium.connectOverCDP(process.env.BROWSER_WS_ENDPOINT);
try {
const context = await browser.newContext({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
});
const page = await context.newPage();
await page.goto("https://example.com", { waitUntil: "domcontentloaded", timeout: 45_000 });
await page.screenshot({ path: "page.png", fullPage: true });
} finally {
await browser.close();
}
Install the library in your project using its official instructions. The connection method and websocket URL are provider-specific. Close the browser in a finally block so the remote session is released even when navigation or capture fails.
For a selector capture with Playwright, locate the element and use its screenshot method:
const chart = page.locator("#revenue-chart");
await chart.waitFor({ state: "visible", timeout: 15_000 });
await chart.screenshot({ path: "revenue-chart.png" });
Playwright’s screenshot guide distinguishes screenshots from snapshots used to locate and interact with controls. Read the official Playwright screenshot guide and Page API reference for current options.
6. Wait for the page state you actually need
A navigation event is not the same as visual readiness. Pages may render content after initial HTML, fetch data later, load images only when scrolled into view, or display a consent layer that covers the content. Choose a wait condition tied to the expected result, then inspect the image.
- Navigation: wait for a meaningful document event such as DOM content loaded when that is enough for the site.
- Specific content: wait for a selector that signals the page section is ready.
- Images or fonts: use the SDK or browser APIs’ documented asset waits when those assets matter.
- Lazy content: scroll through the page before full-page capture so deferred material has an opportunity to load. Browserless recommends scrolling for lazy-loaded material; this is a technique, not a guarantee.
- Network idle: use only when suitable. Long polling, analytics, or streaming can prevent idle, and an idle network does not prove the correct content is visible.
- Short delay: use as a last, site-specific adjustment rather than a universal readiness rule.
One possible Playwright pattern is:
await page.goto(url, { waitUntil: "domcontentloaded", timeout: 45_000 });
await page.locator("main article").waitFor({ state: "visible", timeout: 20_000 });
await page.evaluate(async () => {
const step = Math.max(300, Math.floor(window.innerHeight * 0.75));
for (let y = 0; y < document.body.scrollHeight; y += step) {
window.scrollTo(0, y);
await new Promise(resolve => setTimeout(resolve, 150));
}
window.scrollTo(0, 0);
});
await page.screenshot({ path: "article.png", fullPage: true });
Adjust the selector and scroll behavior to the site. Some pages continuously extend as you scroll, so a fixed pass may not cover them all. Check the resulting capture for missing content, overlays, and unloaded assets.
7. Return the screenshot to the agent
Integrations return images differently. A REST endpoint may return raw image bytes; an SDK may return bytes and optionally write a file; an agent tool may return a file reference or encoded image content. For base64 output, decode to bytes before treating it as a PNG. Do not assume a screenshot is a URL.
For Browserless Agent Run specifically, the documented screenshot field is a base64-encoded viewport PNG with a 5 MiB cap. Check the response schema and handle absent or oversized image fields. For a file-based flow, pass the saved path or file reference using the MCP client’s supported image mechanism instead of pasting a large base64 string into ordinary text.
8. Or skip the browser setup
For a URL-to-image capture, ScreenshotNeo returns a screenshot with one GET request. The parameter names used by other screenshot APIs also work. See the ScreenshotNeo API documentation for options and response details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo free and get 1,000 screenshots a month with no card.
9. Reliability, performance, and cost
Reliability checklist
- Confirm the HTTP status and content type before saving a response as an image.
- Set explicit navigation and request timeouts; report failures with the target URL and stage, but never log access tokens or sensitive headers.
- Check the image dimensions and inspect a sample result for a blank page, challenge screen, consent overlay, or missing section.
- Close remote browser sessions in cleanup code.
- Use a bounded retry for transient connection errors only. Avoid blind retries for authentication errors, invalid selectors, or pages that consistently block automation.
- Keep a stable viewport, device scale factor, and wait rule when captures need to be compared over time.
Performance
Capture time depends on page behavior, browser startup or pool availability, assets, and wait conditions. The reviewed sources do not provide comparable independent performance measurements, so there is no sound basis for a provider speed ranking. Reduce unnecessary work by selecting only the required target, choosing a sensible viewport, waiting for a meaningful condition, and avoiding long fixed delays when a selector can signal readiness. Full-page captures and high device scale factors can increase image size and processing work.
Cost and capacity
Check the provider’s account pricing, session limits, and output limits before setting concurrency. The research does not establish comparable provider prices or success rates. Browserless’s 5 MiB Agent Run image-field cap applies to that field only. For ScreenshotNeo, the free tier is 1,000 shots monthly without a card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed; response headers identify the page verdict and billing status.
10. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| 401 or 403 response | Missing, invalid, or unauthorized provider token. | Check the account’s current authentication format and secret value; do not put the token in the URL you share or in logs. |
| Timeout during navigation | Slow page, stuck resource, or wait condition that never completes. | Set a navigation timeout, wait for a specific selector, and avoid requiring network idle on pages with long-lived requests. |
| Image is blank | The page did not render, automation was blocked, or capture happened too early. | Inspect the page state, try a meaningful selector wait, and check for an automation challenge. A successful transport response alone does not establish successful page rendering. |
| Consent prompt or popup covers content | The overlay is part of the page state and was not handled. | Use an interactive agent to dismiss or accept it when that is appropriate. For direct captures, use a service that removes supported overlays, such as ScreenshotNeo. |
| Missing lower-page images or sections | Lazy loading waits until the area approaches the viewport. | Scroll through the page before capturing and verify the final image; site-specific lazy loading may need additional handling. |
| Selector capture fails | The selector is wrong, the element is hidden, or it is inside a frame or shadow root. | Inspect the rendered page, wait for visibility, and use the integration’s documented selector behavior. |
| Screenshot field missing or too large | Wrong response schema, unsupported capture mode, or Browserless Agent Run’s documented 5 MiB field cap. | Check the endpoint schema; use a dedicated full-page screenshot route when needed and keep the resulting image within the documented field limit. |
| Remote session remains allocated | Browser close was skipped after an exception. | Put browser cleanup in a finally block and review provider session lifecycle options. |
| Image file contains an error document | The endpoint returned an error body that code wrote without checking status. | Check status and content type before saving bytes; inspect the response body safely for diagnostic details. |
11. FAQ
Can an AI agent take a screenshot without a browser pool?
Yes. A direct screenshot API can render a URL for the agent when no interaction is required. A pool is useful for managed browser sessions and interactive workflows.
Should the agent use a screenshot to find buttons?
Usually use the browser’s page structure or snapshot to locate controls, then capture an image when visual review is needed. Screenshots are useful visual evidence, but they can be less direct for identifying actionable elements.
Can a viewport screenshot represent a whole long page?
No. It shows the visible viewport. Request a full-page capture from a route that supports it when the complete document is required.
Does waiting for network idle guarantee the page is ready?
No. It is a signal that may be unsuitable for some sites and does not verify that the intended content or assets are visible.
Sources
- Browserless documentation: screenshot API, browser agent, Agent Run, and BAP guides.
- Playwright screenshots guide.
- Playwright Page API reference.
- ScreenshotNeo documentation.


