How to Use an AI Agent to Capture a Webpage Screenshot and Return It as Base64
Capture a webpage with Playwright, encode the image bytes as base64, and return them in your agent’s expected output format.
Use a browser automation tool to capture the page as image bytes, encode those bytes as base64, and return the encoded string in the format your agent framework expects. With Playwright in Node.js, page.screenshot() returns a Buffer, so you can encode it directly without writing a file:
const imageBytes = await page.screenshot({ type: 'png' });
const imageBase64 = imageBytes.toString('base64');
return { image_base64: imageBase64 };
The important distinction is that capturing an image and returning it as base64 are separate steps. A browser tool may display an image, save a file, or return structured image content without exposing a plain base64 string. Check the receiving agent tool’s output contract before choosing the return shape.
1. Capture and return a screenshot with Playwright
The examples below assume you already have a Playwright page created by your agent runtime. If you are building a standalone script, the Node.js example includes browser setup. Playwright documents the screenshot Buffer and its options in the Page API and screenshots guide.
Node.js: runnable standalone example
Install Playwright and its Chromium browser, then save this as screenshot.mjs and run it with Node.js. It prints a JSON object containing the base64 string:
import { chromium } from 'playwright';
const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('body').waitFor({ state: 'visible', timeout: 10_000 });
const imageBytes = await page.screenshot({ type: 'png' });
const imageBase64 = imageBytes.toString('base64');
process.stdout.write(JSON.stringify({ image_base64: imageBase64 }));
} finally {
await browser.close();
}
Install and run:
npm install playwright
npx playwright install chromium
node screenshot.mjs https://example.com
In an existing agent tool, omit browser launch and navigation if the framework supplies the page and the page is already at the requested URL. Return the encoded value through that framework’s supported mechanism; a JavaScript return works inside a tool handler, while a standalone script writes JSON to stdout.
Python: encode screenshot bytes
When your agent runtime exposes a Playwright Python page, use Python’s standard base64 module:
import base64
image_bytes = await page.screenshot(type="png")
image_base64 = base64.b64encode(image_bytes).decode("ascii")
return {"image_base64": image_base64}
Standalone asynchronous example:
import asyncio
import base64
import json
import sys
from playwright.async_api import async_playwright
async def main():
url = sys.argv[1] if len(sys.argv) > 1 else "https://example.com"
async with async_playwright() as playwright:
browser = await playwright.chromium.launch()
try:
page = await browser.new_page(viewport={"width": 1440, "height": 900})
await page.goto(url, wait_until="domcontentloaded", timeout=30_000)
await page.locator("body").wait_for(state="visible", timeout=10_000)
image_bytes = await page.screenshot(type="png")
image_base64 = base64.b64encode(image_bytes).decode("ascii")
print(json.dumps({"image_base64": image_base64}))
finally:
await browser.close()
asyncio.run(main())
Install the Python package and browser first:
python -m pip install playwright
python -m playwright install chromium
python screenshot.py https://example.com
2. Choose viewport, full-page, or element capture
Choose the smallest capture scope that answers the task. A full-page image can become very tall; an element capture keeps the result focused on one component.
| Scope | Playwright option | Use it for |
|---|---|---|
| Viewport | Default; omit fullPage |
The visible browser screen at the current scroll position. |
| Full page | fullPage: true |
A tall image of the scrollable page. |
| Element | locator.screenshot() |
A specific chart, form, card, or other element. |
Examples using the page from the Node.js code:
// Viewport (default)
const viewportBytes = await page.screenshot({ type: 'png' });
// Full scrollable page
const fullPageBytes = await page.screenshot({ type: 'png', fullPage: true });
// One element; wait for it to exist and be visible first
const chart = page.locator('[data-testid="chart"]');
await chart.waitFor({ state: 'visible', timeout: 10_000 });
const chartBytes = await chart.screenshot({ type: 'png' });
return { image_base64: chartBytes.toString('base64') };
Replace the example selector with a stable selector from your target page. Element screenshots are taken from the locator’s bounds; a locator that matches multiple elements or is not visible can fail, so make it unique and wait for it to appear.
3. Set image format and rendering options
Playwright supports PNG, JPEG, and, where supported by the installed browser, WebP. PNG is a sensible default for sharp text and UI. JPEG or WebP can reduce payload size when lossy compression is acceptable. JPEG supports a quality setting; use a value from 0 to 100. The documented screenshot options also include scale, animation handling, masking, background handling, and path. Consult the current API reference for the exact options available in your Playwright version.
// JPEG at a chosen quality
const jpegBytes = await page.screenshot({ type: 'jpeg', quality: 80 });
// Disable animations for a more stable capture
const stableBytes = await page.screenshot({ animations: 'disabled' });
// Mask a dynamic area in the screenshot
const maskedBytes = await page.screenshot({
mask: [page.locator('.live-counter')],
maskColor: '#999999'
});
Do not set path when you only need the in-memory bytes for encoding. A path is useful when you also need a local artifact, but it does not replace base64 encoding or define what the agent returns.
4. Wait for the content the screenshot needs
No single page-load condition fits every site. domcontentloaded is a practical starting point for pages where the required content appears early; client-rendered content may need a locator wait. Avoid treating “navigation finished” as proof that every image, font, animation, or API-backed widget is ready.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('main h1').waitFor({ state: 'visible', timeout: 10_000 });
// Optional: wait for a known image to finish loading
await page.waitForFunction(() => {
const image = document.querySelector('main img');
return !image || (image.complete && image.naturalWidth > 0);
});
const bytes = await page.screenshot();
Use a selector that represents the content you actually need. For a page that loads below the fold, scroll the relevant element into view or use a full-page capture as appropriate. If the application updates continuously, consider disabling animations or capturing after a known state change.
5. Return raw base64 or a data URL
Base64 is a text encoding of the binary screenshot, not a new image format. Unless the consumer specifically expects a data URL, return the plain base64 string. If it expects a data URL, prepend the media type that matches the screenshot encoding:
const bytes = await page.screenshot({ type: 'png' });
const base64 = bytes.toString('base64');
const rawValue = base64;
const dataUrl = `data:image/png;base64,${base64}`;
return { image_base64: rawValue };
For JPEG, use data:image/jpeg;base64,; for WebP, use data:image/webp;base64,. Do not label PNG bytes as JPEG or WebP. Confirm whether your agent framework accepts strings, objects, or a dedicated image content type; the framework’s schema controls the final response shape.
6. Know what an MCP screenshot tool returns
Playwright MCP provides browser_take_screenshot for viewport, full-page, and element capture, with options including format and scale. Its documentation describes image content in the tool response, but that is not the same contract as a plain base64 field in your agent’s final JSON. If your caller requires a base64 string, use a wrapper that can access the image bytes and encode them, or use an API that returns image bytes. See the Playwright MCP tool documentation and check the output schema of the MCP client you run.
The MCP screenshot tool documents PNG, JPEG, and WebP; fullPage; and CSS-pixel versus device-pixel scale. CSS scale produces a smaller image with dimensions tied to CSS pixels, while device scale accounts for the device pixel ratio and can produce a larger image. Full-page capture cannot be combined with an element screenshot in the documented tool options.
7. Use cURL, Python, or Node.js with ScreenshotNeo
If you do not want to install and operate a browser for a one-off capture, ScreenshotNeo is a website screenshot API and MCP server. Its API returns an image or PDF from a GET request. The examples below fetch a WebP screenshot; see the ScreenshotNeo API documentation for request options and response behavior.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const imageBytes = Buffer.from(await res.arrayBuffer());
const imageBase64 = imageBytes.toString('base64');
console.log(JSON.stringify({ image_base64: imageBase64 }));
The Node.js example encodes the returned image bytes in memory. The cURL and Python examples save the response as a file; if your caller needs base64 instead, read the response bytes and encode them as in the Node.js example.
8. Handle errors, payload size, and cost
Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot is blank or missing page content | Capture happened before the required content rendered, or navigation failed. | Check navigation errors and wait for a content-specific locator before capture. |
| Timeout during navigation or selector wait | The site is slow, the selector never appears, or the chosen wait condition never completes. | Use a suitable readiness condition, verify the selector, and set explicit timeouts. Do not blindly wait for every network request to stop on pages with ongoing connections. |
| Element screenshot fails | The locator is absent, hidden, or ambiguous. | Use a unique selector and wait for the locator to become visible. |
| Returned value is not valid base64 | Text, a file path, or framework image content was returned instead of the encoded bytes. | Encode the screenshot Buffer/bytes and return the base64 string in the agreed field. |
| Image viewer rejects the data URL | The media type prefix does not match the screenshot format, or the receiver expects raw base64. | Match the prefix to PNG, JPEG, or WebP, or omit the prefix if raw base64 is required. |
| Agent response is too large | Base64 expands binary data and the capture may be large, especially for full-page or device-scale images. | Use viewport or element capture, choose an appropriate format and scale, or pass a file/reference if the receiving system supports it. |
| cURL/Python response is HTML instead of an image | The request failed or returned an error response that was saved as if it were an image. | Check the HTTP status and response headers before treating the body as image bytes. In Python, call raise_for_status(). |
Performance and reliability notes
- Keep the browser alive across multiple captures when your agent runtime allows it; launching a browser for every URL adds setup work.
- Choose viewport or element capture when a full-page image is unnecessary. Large images take longer to transfer and encode.
- Wait for a meaningful page state instead of a fixed delay when the page exposes a reliable selector. Use a bounded timeout so a missing element cannot hang the task indefinitely.
- Return only the data the caller needs. A data URL includes a prefix; raw base64 avoids it. A structured MCP image response may be preferable when the next consumer can accept it directly.
- For retries, distinguish transient navigation or service failures from a consistently invalid URL or selector. Apply a bounded retry policy in your agent rather than retrying indefinitely.
Base64 text is larger than its binary source, so returning a full-page high-resolution screenshot can consume a large amount of tool or model context. Keep the bytes in memory while encoding and avoid printing the image string to logs unless you need it. For ScreenshotNeo, only clean shots are billed: bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Check these headers when diagnosing a result.
ScreenshotNeo cleanup, options, and pricing
ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. It also supports full-page and CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF options, HTML/CSS capture, custom CSS and JavaScript, clicks, selector hiding, wait conditions, request and resource blocking, headers, cookies, user agent, Authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed image links, async jobs with signed webhooks, bulk capture up to 100 URLs per call, a usage API, and an OpenAPI spec. Parameter names used by other screenshot APIs also work to make switching easier.
Every feature is available on every plan. Free includes 1,000 shots per month with no card; paid plans are Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free. For agents, its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
9. Or skip the browser setup
Make one GET request to capture a URL with ScreenshotNeo:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
10. Frequently asked questions
Does base64 make the screenshot an image file?
No. It is a text representation of the image bytes. Decode it to recover the original image, or provide a correctly prefixed data URL when that is what the consumer expects.
Should an agent return a data URL or just base64?
Follow the receiving tool’s schema. Use raw base64 for a field that explicitly asks for base64; use a data URL only when the consumer expects the media type prefix.
Can I return an element screenshot instead of the full page?
Yes. Capture a unique visible Playwright locator with locator.screenshot(), then encode those returned bytes the same way.
Does Playwright MCP always give me a base64 string?
Do not assume so. Its screenshot tool returns image content through the MCP response; inspect your MCP client’s response structure or add a wrapper that converts accessible image bytes into the string field your caller requires.


