How to Run Parallel Bulk Screenshots with Puppeteer Without Overloading a Site
Capture batches with Puppeteer using a bounded queue, per-host limits, and clear signals for slowing down or pausing when a site shows strain.
Use Page.screenshot() for each capture, put URL jobs behind a bounded queue, and limit simultaneous jobs per hostname. Start conservatively, watch navigation latency and HTTP status codes, and reduce concurrency or pause when response times rise or the site returns 429 or 5xx errors. Puppeteer does not provide a universally safe concurrency number: choose one based on authorization, the site’s capacity, and observed behavior.
This guide builds a runnable Node.js batch worker with per-host concurrency limits, a configurable pause between jobs to the same host, cleanup, and a circuit breaker for server pressure. These controls are operational choices, not a rate-limit formula provided by Puppeteer.
1. Check access and plan the batch
Capture only sites you are allowed to access, and review their applicable terms and operational guidance. RFC 9309 defines the Robots Exclusion Protocol for automated clients, but explicitly says its rules “are not a form of access authorization.” A robots.txt allowance does not grant permission or establish a safe request rate. See the IETF RFC 9309.
Before running a batch, decide:
- Which URLs are in scope, and whether you have permission to capture them.
- Whether pages on the same host may run at the same time. Set a per-host cap; a global cap alone can direct too many workers at one site.
- What evidence should cause the batch to slow down or stop. At minimum, observe navigation duration, timeouts, 429 responses, and 5xx responses.
- Whether screenshots need a viewport, full page, or a clipped region.
- Whether jobs need shared cookies and local storage. Choose browser contexts to control task state, not as a site-load control.
2. Install Puppeteer
The following example uses Node.js and Puppeteer. Install Puppeteer in a new project; its package includes a compatible browser installation workflow. See the Puppeteer screenshots guide.
mkdir screenshot-batch
cd screenshot-batch
npm init -y
npm install puppeteer
Save the next example as capture-batch.mjs. It uses only Node.js built-ins and Puppeteer.
3. Run a bounded per-host screenshot queue
Set MAX_PER_HOST and MIN_HOST_GAP_MS for your authorized workload. The defaults are deliberately configurable examples, not universally safe values. A host is derived from each URL’s origin, so different schemes or ports count as different hosts.
import puppeteer from 'puppeteer';
import { mkdir, writeFile } from 'node:fs/promises';
import { setTimeout as sleep } from 'node:timers/promises';
const urls = [
'https://example.com/',
'https://example.com/about',
'https://www.example.org/',
];
const OUTPUT_DIR = 'screenshots';
const MAX_PER_HOST = Number(process.env.MAX_PER_HOST ?? 1);
const MIN_HOST_GAP_MS = Number(process.env.MIN_HOST_GAP_MS ?? 1500);
const NAVIGATION_TIMEOUT_MS = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);
const MAX_SERVER_ERRORS = Number(process.env.MAX_SERVER_ERRORS ?? 2);
if (!Number.isInteger(MAX_PER_HOST) || MAX_PER_HOST < 1) {
throw new Error('MAX_PER_HOST must be a positive integer');
}
if (MIN_HOST_GAP_MS < 0) throw new Error('MIN_HOST_GAP_MS must be nonnegative');
const jobs = urls.map((rawUrl, index) => {
const url = new URL(rawUrl);
if (url.protocol !== 'http:' && url.protocol !== 'https:') {
throw new Error(`Unsupported URL protocol: ${url.protocol}`);
}
return {
url: url.href,
host: url.host.toLowerCase(),
index,
filename: `${String(index + 1).padStart(4, '0')}.png`,
};
});
await mkdir(OUTPUT_DIR, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
const hostState = new Map();
let nextJob = 0;
let serverErrors = 0;
let stopped = false;
function stateFor(host) {
if (!hostState.has(host)) {
hostState.set(host, { active: 0, nextAllowedAt: 0 });
}
return hostState.get(host);
}
async function acquire(job) {
const state = stateFor(job.host);
while (!stopped) {
const waitMs = Math.max(0, state.nextAllowedAt - Date.now());
if (state.active < MAX_PER_HOST && waitMs === 0) {
state.active += 1;
state.nextAllowedAt = Date.now() + MIN_HOST_GAP_MS;
return state;
}
await sleep(Math.max(50, waitMs));
}
return null;
}
async function capture(job) {
const state = await acquire(job);
if (!state) return;
let page;
const startedAt = Date.now();
try {
page = await browser.newPage();
page.setDefaultNavigationTimeout(NAVIGATION_TIMEOUT_MS);
const response = await page.goto(job.url, { waitUntil: 'networkidle2' });
const status = response?.status() ?? null;
const elapsedMs = Date.now() - startedAt;
console.log(JSON.stringify({ event: 'navigation', url: job.url, status, elapsedMs }));
if (status === 429 || (status !== null && status >= 500)) {
serverErrors += 1;
// Stop new work after repeated pressure signals. Workers already in flight finish cleanup.
if (status === 429 || serverErrors >= MAX_SERVER_ERRORS) stopped = true;
throw new Error(`Pressure response: HTTP ${status}; batch ${stopped ? 'paused' : 'continuing'}`);
}
if (status !== null && status >= 400) {
throw new Error(`Page returned HTTP ${status}`);
}
const outputPath = `${OUTPUT_DIR}/${job.filename}`;
await page.screenshot({ path: outputPath, type: 'png', fullPage: true });
console.log(JSON.stringify({ event: 'saved', url: job.url, outputPath, elapsedMs: Date.now() - startedAt }));
} catch (error) {
console.error(JSON.stringify({ event: 'failed', url: job.url, error: String(error) }));
} finally {
if (page) await page.close().catch(() => {});
state.active -= 1;
}
}
async function worker() {
while (!stopped) {
const jobIndex = nextJob++;
if (jobIndex >= jobs.length) return;
await capture(jobs[jobIndex]);
}
}
try {
// One worker per configured slot is enough; the per-host gate enforces each host's cap.
await Promise.all(Array.from({ length: Math.max(1, MAX_PER_HOST * 2) }, () => worker()));
if (stopped) console.error('Batch stopped after a rate-limit or repeated server-error signal. Investigate before resuming.');
} finally {
await browser.close();
}
Run it with the defaults:
node capture-batch.mjs
Or set the controls for a workload you are authorized to run:
MAX_PER_HOST=1 MIN_HOST_GAP_MS=2000 NAVIGATION_TIMEOUT_MS=45000 node capture-batch.mjs
The code reserves a per-host slot before navigating and holds it through screenshot and page cleanup. It stops assigning new jobs when it sees a 429 or repeated 5xx responses. Jobs already in progress finish their cleanup. Review the logs and investigate before restarting a stopped batch; do not automatically retry a pressured host in a tight loop.
4. Tune limits using site behavior
There is no official Puppeteer setting that translates a concurrency value into a safe load for every site. Puppeteer documents how to capture screenshots, but leaves workload policy to the caller. Google’s crawler guidance describes how Google responds to server errors and rate limits and how response time affects its own crawl capacity. That is useful context about site-health signals, not a Puppeteer retry specification or a rate policy for other clients. See Google’s guidance on reducing crawl rate and its crawl budget documentation.
- Begin with one active job per host and a modest gap between starts.
- Record host, status, navigation duration, and failure type for every job.
- Keep the limit unchanged until the batch completes reliably and the site shows no signs of pressure.
- If latency rises, timeouts appear, or 429/5xx responses occur, stop increasing demand. Reduce concurrency or pause and investigate.
- Only resume after reviewing the cause and ensuring the workload remains authorized.
The sample’s circuit breaker is intentionally simple: any 429 stops new work, and repeated 5xx responses stop the batch. Production policy may need a manual pause, alerts, persisted job state, and an operator-controlled resume. Avoid automatic retries that can amplify load while the site is unhealthy.
5. Choose capture and browser-state options
Viewport, full page, or clip
Page.screenshot() returns image bytes by default and accepts options including a file path, full-page capture, clipping, image type, and quality where supported. This example writes PNG files with fullPage: true. For a viewport screenshot, omit fullPage; for a specific region, use the documented clip option. JPEG and WebP quality options depend on the selected type. See the Page.screenshot() API reference.
Navigation completion
The example uses networkidle2, which can be a poor fit for pages that keep network connections open or load continuously. If it times out, use an appropriate navigation condition such as domcontentloaded, then wait for the specific content you need before capture. A screenshot taken too early may omit client-rendered content; waiting for every network request may never complete on some pages. Set a navigation timeout and choose per-site behavior deliberately.
Contexts and isolation
For a public page batch, a fresh page per job can be sufficient. If each task needs isolated cookies and local storage, create separate BrowserContexts. Puppeteer documents contexts as a way to isolate automation tasks; cookies and local storage are not shared between contexts, and closing a context closes its pages. Context isolation does not reduce the requests a page makes to the target. See Puppeteer browser management.
Request interception
Do not enable request interception just to claim a load optimization. Puppeteer requires intercepted requests to be resolved, continued, or aborted; unresolved interception can stall page loading. If interception is genuinely required for your workflow, resolve every request on every code path and test its effect on page behavior. See Page.setRequestInterception().
6. cURL, Python, and ScreenshotNeo option
Puppeteer is a Node.js browser automation library, so cURL and Python do not invoke Puppeteer itself. They can call a screenshot API when you want a managed capture request instead of running Chromium and maintaining browser workers.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
For credentials, response headers, options, and runnable API examples, see the ScreenshotNeo API documentation.
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its clean-shot workflow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Every feature is on every plan. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Higher tiers are Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
7. Performance, reliability, and cost
Performance
- Bounded workers prevent an unbounded number of pages from opening at once. A per-host gate keeps a batch spread across hosts from accidentally concentrating all work on one target.
- Browser launch and page creation have overhead. Reusing a browser for a batch avoids launching Chromium for each URL; close each page in a
finallypath and close the browser even if the batch stops. - Full-page captures can use more memory and produce larger files than viewport captures, especially on long pages. Capture only the needed area and format.
- Do not increase concurrency solely to shorten a batch. The target site’s observed response and your authorization determine whether a change is appropriate.
Reliability
- Persist completed URLs and output paths if batches must resume after process interruption; the short example keeps state only in memory.
- Use deterministic filenames or a mapping from URL to output so reruns do not silently overwrite unrelated captures.
- Log response status and timing separately from screenshot success. Navigation may return an HTTP error page that Puppeteer can still render.
- Do not confuse browser-context isolation or a delay between jobs with authorization or a guarantee of low impact.
Cost
Self-hosted Puppeteer has no per-screenshot API charge in this example, but it uses compute, memory, storage, and engineering time to maintain Chromium, queue state, and failure handling. A hosted API trades that browser operations work for a service plan; check its current documented limits and billing behavior before sending a batch. ScreenshotNeo documents that only clean shots are billed and lists its plan quotas and prices above. Do not infer a cost per successful capture for another provider without its pricing and billing rules.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| HTTP 429 | The server or an intermediary is rate limiting requests. | Stop new work, inspect the authorized request pattern, and contact the site operator if needed. Do not loop immediate retries. |
| HTTP 5xx or rising latency | The origin or an upstream service may be under strain or failing. | Pause or lower per-host concurrency, investigate, and resume only when appropriate. |
Navigation timeout |
The page is slow, the chosen wait condition never settles, or the configured timeout is too short. | Inspect the page behavior. Consider domcontentloaded plus a targeted content wait, and set a reasoned timeout; do not respond by blindly raising concurrency. |
| Screenshot is blank or incomplete | Capture happened before client-rendered content or images appeared, or navigation returned an error page. | Check status and page content, then wait for a relevant selector or application-ready condition before capture. |
| Batch never finishes | A page may keep connections open, a worker may be waiting for a navigation condition, or a queue-control bug may have left work pending. | Use bounded timeouts, log job state, and choose an appropriate readiness condition. Avoid unresolved request interception. |
| Browser or machine runs out of memory | Too many pages are open, or full-page images are large. | Reduce concurrency, close pages promptly, and capture a viewport or clip when sufficient. |
| Some URLs are missing after a stop | The sample stops assigning jobs after a pressure signal and does not persist or retry pending jobs. | Inspect the log and resume deliberately after investigation; add a durable queue if resumability is required. |
| 429/5xx page still produced a screenshot | Browser navigation can render an HTTP error response. | Check response status before saving and treat the status as an operational signal, as the example does. |
9. FAQ
Is there a safe number of Puppeteer pages to run in parallel?
No universal number is established by Puppeteer’s screenshot documentation. Set a per-host limit for your authorized workload and adjust from observed site behavior.
Does robots.txt authorize a screenshot batch?
No. RFC 9309 says robots rules are not access authorization. Follow applicable terms and obtain permission where appropriate.
Do BrowserContexts protect a site from load?
No. Contexts isolate browser state such as cookies and local storage. They do not limit the network activity generated by page navigations.
Should every failed job be retried?
No. A retry policy should distinguish transient local failures from site pressure signals. Pause and investigate 429 or repeated 5xx responses instead of retrying in a loop.


