How to Bulk Screenshot URLs with a Browser Farm
Process URL lists with bounded browser concurrency, resumable results, and reliable captures. Compare Playwright workers, managed browsers, and screenshot APIs.
To bulk screenshot URLs with a browser farm, put the URLs in a manifest, run a bounded number of browser sessions in parallel, save each result under a deterministic filename, and record per-URL status so failed work can be resumed. A browser farm supplies concurrent browser capacity; your job still needs to define what a complete capture means, wait for the right page state, and validate the output.
For full browser control, use Playwright workers on infrastructure you manage or connect your existing script to managed browser sessions. For straightforward captures that need no browser-side interaction, a screenshot REST API can be simpler. There is no universal safe concurrency setting: set it from your provider’s current limits, available memory, target site behavior, and retry budget.
1. Choose the browser farm approach
| Approach | Choose it when | What you operate or verify |
|---|---|---|
| Self-managed Playwright workers | You need browser-level control and already have an automation codebase. | Browser versions, worker scaling, memory, storage, retries, and observability. |
| Managed browser sessions | You want to keep a Playwright or Puppeteer flow while delegating browser infrastructure. | Supported browsers, session and concurrency limits, region, data handling, debugging, reliability, and current price. |
| Screenshot REST API | Each job is a stateless URL-to-image capture with no custom browser interaction. | Wait controls, output formats, capture limits, and how blocked pages are represented. |
Browserless documents a REST screenshot endpoint that accepts a URL or HTML and returns image bytes, managed browser connections over WebSocket, and self-hosting options. These are implementation choices, not a verified head-to-head performance ranking. Check current plan limits, data handling, regional availability, and pricing before selecting a provider. [Browserless Screenshot API] [Browserless overview]
For captures that do not need browser automation, ScreenshotNeo is the first screenshot API to try: cookie banners, popups, and chat widgets are removed before capture, only clean shots are billed, and the lowest paid plan is $5. See ScreenshotNeo and its API documentation.
2. Prepare a manifest that can be resumed
Store a stable ID with every URL. A JSON-lines manifest is convenient for worker queues because each row is independent:
{"id":"pricing","url":"https://example.com/pricing"}
{"id":"docs-home","url":"https://example.com/docs"}
Before dispatching work:
- Validate that every URL has an allowed scheme, usually
httpsorhttp, and normalize it consistently. - Decide whether duplicate URLs should produce separate records or be deduplicated. Keep the original URL even if you normalize a copy for naming or deduplication.
- Use stable IDs or a URL hash for filenames. Sanitize IDs and prevent path separators or collisions.
- Keep the manifest and a result record for each row: original URL, final URL if known, output path, timestamp, status, attempt count, and error detail.
- Write completed results incrementally. A process restart should pick up pending and failed rows without recapturing successful ones.
For example, a result record can be {"id":"pricing","url":"https://example.com/pricing","file":"shots/pricing.png","status":"ok","attempts":1}. Avoid using the URL itself as a raw path: query strings, long URLs, and encoded characters make unreliable filenames.
3. Build a bounded Playwright worker
The following Node.js example uses Playwright, a JSON-lines input file, a concurrency limit, deterministic filenames, per-URL result records, and capped exponential backoff. Install the package with npm install playwright and install a browser with npx playwright install chromium. Save the manifest as urls.jsonl, then run the script as node capture.mjs. The worker pool is bounded by CONCURRENCY; tune it rather than launching a browser for every URL at once.
import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
const CONCURRENCY = Number(process.env.CONCURRENCY ?? 3);
const MAX_ATTEMPTS = Number(process.env.MAX_ATTEMPTS ?? 3);
const NAVIGATION_TIMEOUT_MS = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);
const VIEWPORT = { width: 1365, height: 900 };
if (!Number.isInteger(CONCURRENCY) || CONCURRENCY < 1) {
throw new Error('CONCURRENCY must be a positive integer');
}
const rows = (await readFile('urls.jsonl', 'utf8'))
.split(/\r?\n/)
.filter(Boolean)
.map((line, index) => {
const row = JSON.parse(line);
if (!row.id || !row.url) throw new Error(`Row ${index + 1} needs id and url`);
const parsed = new URL(row.url);
if (!['http:', 'https:'].includes(parsed.protocol)) {
throw new Error(`Unsupported URL scheme in row ${index + 1}`);
}
return { id: String(row.id), url: parsed.href };
});
await mkdir('shots', { recursive: true });
const results = [];
let next = 0;
async function capture(browser, row) {
const hash = createHash('sha256').update(row.id).digest('hex').slice(0, 12);
const safeId = row.id.replace(/[^a-zA-Z0-9_-]/g, '_').slice(0, 60) || 'page';
const path = `shots/${safeId}-${hash}.png`;
let lastError;
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
const context = await browser.newContext({ viewport: VIEWPORT, deviceScaleFactor: 1 });
const page = await context.newPage();
try {
const response = await page.goto(row.url, {
waitUntil: 'domcontentloaded',
timeout: NAVIGATION_TIMEOUT_MS
});
// Replace this with a meaningful selector when the page has one.
await page.waitForTimeout(500);
await page.screenshot({ path, fullPage: true, type: 'png', timeout: NAVIGATION_TIMEOUT_MS });
const status = response?.status() ?? null;
return {
id: row.id, url: row.url, finalUrl: page.url(), file: path,
status: status !== null && status >= 400 ? 'http-error' : 'ok',
httpStatus: status, attempts: attempt, capturedAt: new Date().toISOString()
};
} catch (error) {
lastError = String(error?.message ?? error);
if (attempt < MAX_ATTEMPTS) {
const delay = Math.min(8000, 500 * (2 ** (attempt - 1)));
await new Promise(resolve => setTimeout(resolve, delay));
}
} finally {
await context.close();
}
}
return {
id: row.id, url: row.url, file: null, status: 'failed',
attempts: MAX_ATTEMPTS, error: lastError, capturedAt: new Date().toISOString()
};
}
const browser = await chromium.launch({ headless: true });
try {
async function worker() {
while (true) {
const index = next++;
if (index >= rows.length) return;
results[index] = await capture(browser, rows[index]);
// Persist after each URL so a later process failure does not lose the batch record.
await writeFile('results.json', JSON.stringify(results.filter(Boolean), null, 2));
}
}
await Promise.all(Array.from({ length: Math.min(CONCURRENCY, rows.length) }, worker));
} finally {
await browser.close();
}
The sample uses domcontentloaded plus a short delay as a portable fallback, not proof that every image or client-rendered component is ready. If the page exposes a stable readiness signal, wait for it instead, such as await page.locator('main article').waitFor({ state: 'visible' }). A fixed delay can waste time on fast pages and still miss slow content.
This minimal sample rewrites the current results file after each completion. For a large or multi-process farm, use a durable queue or append-only result store and an atomic write strategy; otherwise, simultaneous workers or a crash during a write can lose metadata. To resume, load prior results and enqueue only rows not marked successful. Treat HTTP errors as outcomes to inspect, not automatically successful content.
4. Set capture behavior before adding workers
Make capture semantics consistent across the batch. Playwright’s page screenshot supports full-page capture and options for type, quality, scale, style, and timeout. fullPage captures the full scrollable page rather than only the viewport. See the official Playwright Page screenshot API.
| Need | Configuration choice | Edge to consider |
|---|---|---|
| Visible screen only | Use a fixed viewport and omit fullPage. |
Viewport dimensions and device scale affect wrapping and pixel dimensions. |
| Entire scrollable page | Set fullPage: true. |
Very long pages can use substantial memory; lazy-loaded content may not appear until scrolled. |
| Specific component | Use a locator screenshot, for example await page.locator('.hero').screenshot({ path }). |
Missing or hidden selectors should be recorded as capture errors, not silently replaced by a whole-page shot. |
| Consistent visual comparisons | Fix viewport, device scale, color scheme, locale, timezone, and fonts where possible. | Dynamic ads, timestamps, personalization, and animation can still change output. |
| Image format | PNG preserves lossless detail; JPEG can reduce storage; Playwright also supports WebP where configured. | Quality applies to lossy formats; ensure downstream consumers accept the chosen type. |
| Hide irrelevant moving UI | Apply a screenshot style or injected CSS only for elements irrelevant to the evidence. | Do not hide page state that the capture is meant to document. |
For lazy images, scroll in steps before capture and allow loading to settle. Browserless documents that full-page capture can miss lazy content unless the page is scrolled. For a managed screenshot API, use its documented full-page and wait options; option names vary by provider. [Browserless capture options]
5. Set concurrency, retries, and recovery
Concurrency is a resource and policy setting, not a universal constant. Increase it gradually while watching worker memory, browser launch failures, provider session limits, target response behavior, queue age, and timeout frequency. Reduce it if pages slow down, sessions are rejected, or the machine approaches its memory limit.
- Start with a small configurable worker count and a fixed timeout.
- Record duration, navigation status, retry count, output size, and final URL per job.
- Retry only plausibly transient failures, such as temporary network errors or navigation timeouts. Use capped exponential backoff and add jitter in a distributed queue.
- Do not retry invalid URLs, persistent authorization errors, or a stable access-denied response indefinitely.
- Persist each completion and make the queue idempotent so a repeated job has a known output path and does not corrupt another result.
Browserless’s examples include concurrent sessions and exponential-backoff retries. They demonstrate patterns, not a recommended concurrency number or speed guarantee. [Browserless examples]
6. Validate captures and handle blocked pages
A successful navigation or HTTP response does not guarantee a useful screenshot. Browser automation can encounter blank pages, CAPTCHA challenges, 403/access-denied pages, or missing elements. Keep the original response status and final URL, reject zero-byte outputs, and inspect a sample of images from every batch. For critical captures, add an application-specific check such as the presence of a page heading or expected element.
When a site blocks automation, use an authorized API or export where available, or ask the site owner for access. Do not assume an unblock feature will work or that bypassing a protection is allowed. Browserless documents these blocking symptoms in its screenshot guidance. [Browserless troubleshooting guidance]
7. Use a REST API for stateless captures
If each task only maps one URL to one image, a REST API can avoid browser lifecycle code. Browserless documents POST /screenshot with a URL or HTML and screenshot options; authentication and endpoint host depend on deployment. Use the current provider documentation for request shape and limits. [Browserless Screenshot API]
For a batch, place API calls behind the same bounded queue and result ledger used for browser workers. Check HTTP status, response content type, and whether the response is actually an image before writing it. Apply retries to transient status codes and network failures only; preserve permanent errors for review.
Or skip the browser setup
For ordinary URL-to-image jobs, ScreenshotNeo takes one GET request per capture. The same request can be placed in a bounded worker queue for bulk work. The ScreenshotNeo API documentation describes supported request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. For bulk work, use a bounded queue and inspect the response’s X-Page-Verdict and X-Billed headers to see the outcome and billing status. See ScreenshotNeo and the API docs, then sign up for 1,000 free screenshots a month with no card.
Performance, reliability, and cost
- Performance: Total time depends on page load behavior, browser startup, capture size, and concurrency. Bound sessions to control memory and tune with observed queue and error data; no cross-provider speed figure is established here.
- Reliability: Persist results incrementally, retry transient failures with a cap, keep stable IDs, and validate a sample of image output. Separate navigation success from visual usefulness.
- Cost: For a self-managed farm, account for compute, storage, and engineering maintenance. For managed browsers or APIs, verify current session quotas, request limits, retention, and prices directly with the provider. Avoid recapturing completed rows and choose output formats that fit the quality and storage requirement.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation times out | Slow site, long-lived requests, overloaded worker, or too-short timeout. | Use a readiness condition such as domcontentloaded or a known selector, check site behavior, lower concurrency, and set a realistic timeout. Do not blindly extend every timeout. |
| Screenshot is blank or mostly white | Page did not render, automation was blocked, or capture happened before content appeared. | Check final URL, response status, console/navigation errors, and an expected element. Wait for the relevant content and use an authorized access route if blocked. |
| CAPTCHA or access-denied screen | The target site restricts automated access. | Respect the site’s access rules; request permission or use its supported API/export. Do not treat retries as a solution to a persistent block. |
| Lazy images are missing | Images load only after their region enters the viewport. | Scroll through the page in steps, wait for image loading, then capture; confirm the page’s intended behavior. |
| Some jobs fail only at high concurrency | Memory pressure, provider session caps, or target-site rate limiting. | Lower the worker count, inspect resource use and provider limits, and increase gradually only while errors remain acceptable. |
| Wrong or colliding output filenames | Names came directly from URLs or IDs were not unique. | Use sanitized stable IDs plus a short hash, and retain the original URL in the result record. |
| Batch is lost after a worker crash | Results were only held in memory or the output file was interrupted during a write. | Persist each completion in a durable queue or append-only store and use atomic writes for snapshots. |
| HTTP 200 but unusable image | The server successfully returned a challenge, blank page, or error page as content. | Inspect image samples and validate expected page elements; do not equate transport success with a valid capture. |
FAQ
Should I use one browser per URL?
No. Reuse a browser process and create a bounded number of isolated contexts or sessions. Keep concurrency configurable and within machine and provider limits.
Should I screenshot the viewport or the full page?
Use viewport capture for consistent screen-sized comparisons. Use full-page capture when the complete scrollable document matters, and account for lazy-loaded content and large image outputs.
Can I resume only failed URLs?
Yes. Keep a manifest and a durable result record keyed by stable ID, then enqueue rows that are pending or failed under your retry policy.
Is a browser farm always faster than a screenshot API?
No universal comparison is supported. The right choice depends on browser interaction needs, current service limits, page behavior, and the cost of operating workers.


