How to Capture Screenshots of Multiple URLs with an AI Agent
Build an AI-agent workflow that captures a list of URLs with Playwright, saves uniquely named images, and reports failures clearly.
To capture screenshots of multiple URLs with an AI agent, give it a URL list and explicit capture requirements, then have it visit each page in a real browser, save a uniquely named screenshot, and record the result or failure for every URL. Playwright provides the per-page building blocks; looping over a list, naming files, and reporting partial failures are workflow decisions you implement around those calls. Playwright Page API
This guide uses Playwright with Node.js. It covers a runnable batch script, capture choices, reliability and troubleshooting, and an API option for cases where you do not want to manage browser setup.
1. Specify what the agent should capture
Before sending a batch to an agent, decide what each screenshot needs to show. A vague instruction such as “screenshot these pages” leaves important choices unresolved.
- Capture scope: the initial viewport, the full scrollable page, or one element selected with a CSS selector.
- Viewport: width and height in CSS pixels, plus whether device-scale pixels are needed.
- Image format: PNG for lossless images, JPEG for smaller photographic images, or WebP when that format suits the downstream workflow.
- Readiness: a specific selector, a short delay, or network idle, depending on what must appear before capture.
- Output mapping: a unique name for every result and a manifest that keeps the original URL beside its output or error.
- Authentication and state: whether pages need cookies, headers, or a logged-in browser context. Only use credentials the workflow is authorized to access.
Playwright’s screenshot interfaces support viewport, element, and full-page capture, with PNG, JPEG, or WebP output. The Playwright MCP screenshot tool also documents a fullPage option that cannot be combined with a target element, and a scale option for CSS-pixel or device-pixel sizing. Check the interface you use because options can differ. Playwright screenshot commands · Playwright MCP screenshot tools
2. Build a batch capture script with Playwright
This script takes URLs from a JSON file, visits them one at a time, saves a full-page PNG for each successful navigation, and writes a JSON manifest with the outcome for every input. Sequential processing keeps the example simple and avoids opening an unbounded number of browser pages at once.
Install Playwright
mkdir multi-url-capture
cd multi-url-capture
npm init -y
npm install playwright
npx playwright install chromium
Create urls.json:
[
"https://example.com/",
"https://www.wikipedia.org/",
"https://playwright.dev/"
]
Create capture.mjs:
import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { createHash } from 'node:crypto';
const inputPath = process.argv[2] ?? 'urls.json';
const outputDir = process.argv[3] ?? 'screenshots';
const timeoutMs = Number(process.env.CAPTURE_TIMEOUT_MS ?? 30000);
const readinessSelector = process.env.READINESS_SELECTOR;
const urls = JSON.parse(await readFile(inputPath, 'utf8'));
if (!Array.isArray(urls) || urls.some((url) => typeof url !== 'string')) {
throw new TypeError('Input JSON must be an array of URL strings');
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
for (let index = 0; index < urls.length; index += 1) {
const url = urls[index];
const page = await browser.newPage({
viewport: { width: 1440, height: 900 },
deviceScaleFactor: 1
});
const shortHash = createHash('sha256').update(url).digest('hex').slice(0, 10);
const filename = `${String(index + 1).padStart(3, '0')}-${shortHash}.png`;
const outputPath = path.join(outputDir, filename);
try {
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: timeoutMs
});
if (readinessSelector) {
await page.locator(readinessSelector).waitFor({
state: 'visible',
timeout: timeoutMs
});
}
await page.screenshot({ path: outputPath, fullPage: true });
results.push({
inputUrl: url,
finalUrl: page.url(),
status: 'captured',
httpStatus: response?.status() ?? null,
file: filename
});
} catch (error) {
results.push({
inputUrl: url,
status: 'failed',
error: error instanceof Error ? error.message : String(error)
});
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
await writeFile(
path.join(outputDir, 'manifest.json'),
JSON.stringify(results, null, 2) + '\n'
);
console.log(JSON.stringify({ total: urls.length, results }, null, 2));
Run it with:
node capture.mjs urls.json screenshots
To wait for a page-specific element on every URL, set a selector and optionally adjust the per-navigation timeout:
READINESS_SELECTOR='main' CAPTURE_TIMEOUT_MS=45000 node capture.mjs urls.json screenshots
The example treats a non-2xx HTTP response as a captured page if navigation completes; it records the HTTP status so the caller can decide whether that image is useful. A navigation timeout or screenshot exception is recorded as a failure, and the loop continues. The file name combines the input index and a short URL hash, avoiding collisions when different URLs have similar paths. The manifest retains the full input URL and final URL, including redirects.
3. Adapt capture scope, readiness, and output
Viewport, full page, or element
The example uses fullPage: true. For only the visible viewport, remove that option or set it to false. To capture an element instead, use a locator screenshot:
await page.locator('main article').screenshot({ path: outputPath });
Choose a selector that identifies one element. A missing or ambiguous selector can produce a timeout or capture the wrong content. Full-page capture and element capture are different operations; in Playwright MCP, fullPage and a target element cannot be requested together. Playwright MCP screenshot options
Wait for the page state you need
domcontentloaded waits for the initial document to be parsed. It does not guarantee that client-rendered content, images, or data requests have finished. Use an explicit condition when the page needs more time:
// Wait for a known content element to appear
await page.locator('[data-ready="true"]').waitFor({ state: 'visible' });
// Or wait for an application-specific state
await page.waitForFunction(() => window.appReady === true);
// Or use a bounded delay when the page has no reliable readiness signal
await page.waitForTimeout(1500);
A delay is a fallback, not proof that a page is ready. Network-idle conditions can also be unsuitable for pages with persistent requests. Prefer a page-specific selector or application signal when available. The Playwright screenshot command documentation describes waiting for a selector, delay, or network idle as capture controls. Playwright screenshot command options
Choose format and pixel scale
Playwright’s Page screenshot API supports PNG and JPEG output, selected by the file extension or options; its screenshot command documentation also lists WebP. Confirm support in the specific interface and browser version you use before standardizing a format. For Playwright MCP, scale can request CSS-pixel or device-pixel sizing. A higher device scale produces more pixels and can increase storage and transfer requirements. Playwright Page API · Screenshot command formats · MCP scale option
4. Give the agent a clear batch contract
Whether the AI agent writes or runs the script, provide a precise input and output contract. For example:
Capture each URL in urls.json using Chromium at a 1440 by 900 CSS-pixel viewport.
Save one full-page PNG per URL in screenshots/.
Use a unique filename and preserve input order in the manifest.
Wait for the main content selector when one is provided; otherwise wait for DOM content loaded.
Continue after individual failures. Record input URL, final URL, HTTP status, filename, and error.
Do not claim the screenshot proves the page is accessible or that its text is correct.
For visual review, give the agent the resulting images. When it needs to understand page structure or locate controls, use an accessibility snapshot or other structured page data as well. Playwright MCP documentation distinguishes screenshots for visual inspection from snapshots for structure and interaction. Screenshots show rendered appearance at a moment in time; they do not establish that the page is accessible, that its text is accurate, or that another browser or later capture will look identical. Playwright MCP screenshots · Playwright MCP tools
5. Run batches with useful failure handling
A production batch should make partial success visible instead of treating the entire list as one all-or-nothing operation.
- Keep a manifest: store each source URL beside its output filename, final URL, status, and error.
- Continue per URL: catch navigation and capture errors inside the loop so one failure does not discard completed work.
- Bound concurrency: sequential processing is predictable. If you add parallel pages, set a fixed concurrency limit based on available memory and the target sites’ behavior.
- Retry selectively: consider retrying transient navigation failures with a small bounded retry count. Do not retry a stable 404 or a selector that is known to be wrong.
- Use stable inputs: preserve the original URL and order; record redirects rather than silently replacing the source URL.
- Protect credentials: avoid putting secrets in source files, logs, or manifests. Use a controlled browser context for authorized authenticated pages.
These are batch design practices built around Playwright’s documented single-page navigation and screenshot operations; the Page API example establishes the primitives, not a universal batch policy. Playwright Page API
6. When a hosted browser makes sense
A local Playwright process gives you control over the runtime and browser setup. If an agent needs a hosted browser for inspection or interactive, multi-step automation, Cloudflare documents browser tools for agents as a beta feature. Check its current availability and limits before making it a dependency. Cloudflare Browser tools
For an image-only batch, a screenshot API is another execution path. ScreenshotNeo is a website screenshot API and MCP server for developers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf for AI-agent workflows.
Or skip the browser setup
ScreenshotNeo accepts one GET request per URL and returns a screenshot image or PDF. Run the request once for each URL in your list, saving each response under a distinct filename. The API supports full-page capture with lazy images loaded, viewport and device presets, PNG/JPEG/WebP, and other capture controls. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Replace the example URL with one of your input URLs and repeat the request with a unique output name. In the Python example, install the dependency with python -m pip install requests. The Node.js example uses the built-in fetch and assumes writeFile is imported from node:fs/promises.
- Cookie banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include
X-Page-VerdictandX-Billedheaders. - An MCP server lets AI agents use screenshot, page-info, and PDF tools.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.
7. Performance, reliability, and cost
Performance
Batch duration depends on navigation, readiness waits, page size, and screenshot work. The script’s sequential loop uses one page at a time, limiting simultaneous browser work but waiting for each URL in turn. A fixed concurrency limit can reduce elapsed time when there are many independent URLs, at the cost of more memory and simultaneous traffic to target sites. Measure the workload in its actual runtime before choosing a limit; no universal throughput figure follows from the screenshot API documentation.
Full-page screenshots can be larger and take more work than viewport captures. Capture only the scope needed, choose an appropriate image format, and avoid unnecessary waits. A high device scale can increase pixel count. Keep page timeouts bounded so one unresponsive site does not stall the batch indefinitely.
Reliability and interpretation
Web pages change, redirect, personalize content, and depend on network or client-side state. Save the final URL and capture conditions so someone can interpret each file later. A screenshot is a record of the rendered state at capture time, not proof of accessibility or textual correctness. For repeatable comparisons, keep browser, viewport, scale, authentication state, and wait conditions consistent, while expecting some dynamic content to vary.
Cost
With local Playwright, plan for the compute and storage used by your own environment; the cited Playwright documentation describes the browser APIs, not a hosted execution price. For a hosted browser or API, check the provider’s current limits and pricing before relying on a batch. ScreenshotNeo’s listed plans are Free at 1,000 shots per month, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free. Its billing rule counts only clean shots; bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable is missing | Playwright package is installed, but Chromium was not installed for it. | Run npx playwright install chromium in the project environment. |
| Navigation times out | The site is slow, unresponsive, or waiting for a stronger page-load condition. | Use a bounded larger timeout if appropriate; choose domcontentloaded and then wait for a specific readiness selector. Keep per-URL failures in the manifest. |
| Screenshot is blank or missing content | Capture ran before client-rendered content appeared, or the site served a challenge or empty state. | Wait for an application-specific selector or readiness signal. Inspect the rendered page and record the outcome; a screenshot alone does not explain why content is absent. |
| Element capture times out | The selector does not match, the element is hidden, or it appears later. | Check the selector against the page, wait for the intended visible state, and use viewport or full-page capture if no stable target element exists. |
| Wrong page was captured | The URL redirected, required authentication, or returned an error page. | Compare the input and final URLs, record HTTP status, and provide the authorized cookies or headers needed by the workflow. |
| Some output files overwrite others | Names were derived only from a path or a repeated label. | Include the input index and a URL hash or another unique identifier in the filename. |
| Batch slows down or exhausts memory | Too many browser pages are open at once, or full-page images are large. | Run sequentially or cap concurrency, close each page in a finally block, and capture a smaller scope or scale when suitable. |
| One failure stops the whole run | Error handling surrounds the entire batch rather than each URL. | Catch errors inside the per-URL loop and write one success or failure record per input. |
9. Frequently asked questions
Can an AI agent inspect the screenshots after capturing them?
Yes. Provide the image files to an agent that can inspect images. If it must identify controls or understand semantic structure, pair the screenshots with accessibility snapshots or other structured page data.
Does a successful screenshot mean the page works for users?
No. It shows a rendered state at a particular time and browser configuration. It does not by itself verify accessibility, correctness, or behavior across other environments.
Should every URL use the same wait condition?
Only when the pages share a reliable readiness signal. For mixed sites, use per-URL wait rules or a conservative common condition and record exceptions.
Can I capture authenticated pages?
Yes, when you are authorized to access them. Supply the required session state through the browser context or approved credentials, and keep secrets out of source files and output manifests.


