Puppeteer Screenshot Service with a Queue for Bulk URL Captures
Build a Puppeteer screenshot service that queues bulk URL captures, controls browser workers, and handles failures—or use a hosted screenshot API.
To capture screenshots for many URLs, accept each URL and its capture options as a job, put jobs in a queue, and let a bounded number of Puppeteer workers navigate pages and save results. The queue separates request intake from browser work; it does not by itself guarantee retries, durability, or a particular throughput. For a persistent service, use a durable queue, define retry and timeout behavior, validate URLs, and store results separately from job state.
This guide builds a small Node.js service using Puppeteer and BullMQ. It shows bulk submission, a worker, result storage, failure handling, capacity decisions, troubleshooting, and when a hosted screenshot API is a better fit.
1. How the service fits together
A practical architecture has four parts:
- Producer/API: validates a request and creates one job per URL, including viewport, format, and readiness options.
- Queue: buffers jobs so browser workers can process them at a controlled rate.
- Worker: opens a page, navigates, waits for a useful readiness signal, captures the screenshot, and records success or failure.
- Result store or delivery step: retains the image and makes it available through an object store, database reference, or callback.
Puppeteer captures a page with Page.screenshot(); its guide demonstrates waiting for networkidle2 and also documents element screenshots. Choose readiness based on the target site: network idle can be unsuitable for pages with persistent connections or ongoing requests. Puppeteer screenshots guide · Page.screenshot API.
2. Install and configure the service
Use Node.js with Puppeteer and BullMQ. BullMQ uses Redis for queue coordination. The example stores image files in a local directory to keep the flow runnable; for a multi-worker or production deployment, replace that local result path with shared object storage and persist the resulting object key.
npm init -y
npm install bullmq puppeteer
Start a Redis instance according to its official installation instructions, then set REDIS_HOST and REDIS_PORT if it is not on the defaults. Save the following as service.mjs. Run it in one terminal to start the worker and HTTP producer. The endpoint accepts a JSON array of URLs at POST /captures.
import http from 'node:http';
import { mkdir, writeFile } from 'node:fs/promises';
import { Queue, Worker } from 'bullmq';
import puppeteer from 'puppeteer';
import { createHash } from 'node:crypto';
const connection = {
host: process.env.REDIS_HOST ?? '127.0.0.1',
port: Number(process.env.REDIS_PORT ?? 6379),
};
const queueName = 'captures';
const queue = new Queue(queueName, { connection });
const outputDir = process.env.OUTPUT_DIR ?? './captures';
await mkdir(outputDir, { recursive: true });
function validUrl(value) {
try {
const u = new URL(value);
return u.protocol === 'http:' || u.protocol === 'https:';
} catch {
return false;
}
}
const worker = new Worker(queueName, async job => {
const { url, viewport, format, waitUntil, fullPage } = job.data;
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport(viewport);
await page.goto(url, { waitUntil, timeout: 45_000 });
const bytes = await page.screenshot({ type: format, fullPage });
const name = createHash('sha256').update(job.id).digest('hex') + `.${format}`;
await writeFile(`${outputDir}/${name}`, bytes);
return { file: `${outputDir}/${name}`, finalUrl: page.url() };
} finally {
await browser.close();
}
}, {
connection,
concurrency: Number(process.env.WORKER_CONCURRENCY ?? 2),
limiter: { max: Number(process.env.STARTS_PER_WINDOW ?? 10), duration: 1000 },
});
worker.on('failed', (job, error) => {
console.error('capture failed', job?.id, error.message);
});
const server = http.createServer(async (req, res) => {
if (req.method !== 'POST' || req.url !== '/captures') {
res.writeHead(404).end('Not found');
return;
}
let body = '';
for await (const chunk of req) body += chunk;
try {
const input = JSON.parse(body);
if (!Array.isArray(input.urls) || input.urls.length < 1 || input.urls.length > 100) {
res.writeHead(400).end('urls must contain 1 to 100 entries');
return;
}
const jobs = input.urls.map((item, index) => {
const url = typeof item === 'string' ? item : item.url;
if (!validUrl(url)) throw new Error(`Invalid HTTP(S) URL at index ${index}`);
const opts = typeof item === 'string' ? {} : item;
const format = opts.format ?? 'png';
if (!['png', 'jpeg', 'webp'].includes(format)) throw new Error(`Invalid format at index ${index}`);
const viewport = opts.viewport ?? { width: 1280, height: 800 };
if (!Number.isInteger(viewport.width) || !Number.isInteger(viewport.height) || viewport.width < 1 || viewport.height < 1) {
throw new Error(`Invalid viewport at index ${index}`);
}
return {
name: 'capture',
data: { url, format, viewport, waitUntil: opts.waitUntil ?? 'networkidle2', fullPage: opts.fullPage ?? true },
opts: { attempts: 3, backoff: { type: 'exponential', delay: 1000 }, removeOnComplete: 1000, removeOnFail: 5000 },
};
});
const added = await queue.addBulk(jobs);
res.writeHead(202, { 'content-type': 'application/json' });
res.end(JSON.stringify({ jobIds: added.map(job => job.id) }));
} catch (error) {
res.writeHead(400, { 'content-type': 'application/json' });
res.end(JSON.stringify({ error: error.message }));
}
});
server.listen(Number(process.env.PORT ?? 3000), () => console.log('Listening on :3000'));
Submit a batch from another terminal:
curl -X POST http://localhost:3000/captures \
-H 'content-type: application/json' \
-d '{"urls":["https://example.com",{"url":"https://stripe.com","format":"webp","fullPage":true,"viewport":{"width":1440,"height":1000}}]}'
The response contains job IDs because submission is asynchronous. The sample logs failures, but it does not expose a status endpoint or result download route. Add an authenticated status/result endpoint or callback for clients, and do not treat queue completion as proof that a local file is accessible to another machine.
3. Queue many URLs safely
BullMQ’s Queue.addBulk() adds an array of jobs and may be faster than issuing individual additions. Keep batches bounded: this example caps an HTTP request at 100 URLs, but the right limit depends on request size, Redis capacity, and how much work you want to enqueue at once. See the BullMQ Queue API.
Each job should carry only what the worker needs: URL, viewport, format, full-page choice, readiness condition, and a caller-provided idempotency key if duplicate submissions should map to the same logical capture. Avoid putting large binary data in Redis; store the image elsewhere and keep a reference in job results.
Bulk submission checklist
- Validate URL scheme and allowed hosts before enqueueing. A service that visits arbitrary URLs can become a path to internal network resources; enforce a host allowlist or block private and link-local destinations, including after redirects.
- Set a maximum batch size and request body size. Return job IDs promptly rather than holding the HTTP request open for rendering.
- Use stable IDs or deduplication rules when clients may retry submission after a network timeout.
- Set attempt count and backoff deliberately. Retrying a permanently invalid URL wastes worker time, so classify failures where possible.
- Track queue age, waiting and failed jobs, worker health, and output retention. Queueing creates a place to manage backlog; it does not make jobs durable unless the backing queue and deployment are configured for recovery.
4. Worker capacity and capture options
The example uses worker concurrency and a limiter as separate controls. Concurrency bounds how many jobs that worker processes at once; the limiter constrains how many jobs start in a time window. Neither setting is a universal browser-capacity number. Measure memory and CPU use in your own environment, then tune gradually while watching timeouts and queue age.
Opening a fresh browser for every job is easy to reason about but has startup overhead. A long-lived browser with a fresh page per job can reduce that overhead, but requires lifecycle management: close every page, recycle unhealthy browsers, and ensure one job’s cookies or storage do not leak to another. Whichever model you choose, cap browser processes and pages, set navigation and overall job timeouts, and close resources in finally blocks.
| Need | Typical choice | Trade-off |
|---|---|---|
| Fast first implementation | One browser per job | Simple isolation; browser startup cost for each capture. |
| Repeated captures at volume | Long-lived browser, isolated pages | Less repeated startup work; needs browser recycling and strict cleanup. |
| Long or unpredictable pages | Explicit navigation timeout and job timeout | Some captures fail instead of occupying a worker indefinitely. |
| Sites with persistent network activity | Wait for a selector or a bounded delay after navigation | Site-specific readiness logic; network idle may never occur. |
Useful Puppeteer screenshot controls include image type, full-page capture, clipping, and element capture. For a specific element, select it and capture its bounding box or use the element screenshot capability documented by Puppeteer. Validate that the selector exists and is visible; otherwise return a clear job error rather than silently producing an unrelated full-page image. Set viewport before navigation when responsive layout matters. Retina output can be produced with an appropriate device scale factor, with larger files and more rendering memory as a consequence.
For reliable captures, define whether redirects are accepted, what happens on HTTP error responses, whether authentication cookies or headers are permitted, and how sensitive screenshots are retained. Add per-tenant quotas if clients share the service. Treat page content as untrusted: isolate browser workers, restrict outbound network access, and avoid passing privileged credentials to arbitrary sites.
5. Retries, durability, and results
For work that must survive process restarts, use a durable external queue such as Redis with BullMQ or SQS, and configure the backing service and workers for the recovery behavior you need. ScreenshotOne’s bulk guide recommends retries, respecting request buckets, and a durable queue for persistent workloads; it also cautions that provider concurrency fields can mean starts allowed in a time bucket rather than simultaneously active renders. Read each provider’s exact limit semantics before tuning. ScreenshotOne bulk screenshots guide.
A retry policy should distinguish transient navigation or infrastructure failures from deterministic validation errors. Exponential backoff can reduce repeated pressure on a failing destination, but retries increase total work and delay completion. Make writes idempotent: a retried job should overwrite or version a predictable result rather than create an unbounded set of duplicate files. Keep job metadata and image retention policies separate, since queue cleanup does not automatically delete stored output.
For clients, return a job identifier at submission and offer a status mechanism: pending, active, completed with a result reference, or failed with a sanitized error. Consider signed, expiring result links if images should not be public. Protect webhook callbacks with signatures and retry delivery independently from capture execution.
6. Should you build a Puppeteer service or use a screenshot API?
Build when you need control over queue policy, browser behavior, deployment, and data flow and can own browser operations. A hosted API can reduce the work of running browser workers. For straightforward bulk jobs, compare documented batch endpoints; Screenshot API documents a batch endpoint that returns a tracking ID. When a workflow needs stateful browser interaction, Capture describes hosted browser sessions with a CDP connection URL that can be used with Puppeteer or Playwright. Compare durability, retries, rate limits, capture controls, interaction needs, infrastructure ownership, and integration effort. Current comparative prices, quotas, and performance are not established by the sources cited here, so check each provider’s current documentation before choosing. Screenshot API documentation · Capture.
ScreenshotNeo is the first alternative to try for API-based captures: cookie banners, newsletter popups, and chat widgets are removed before the shot, and only clean shots are billed.
7. Or skip the browser setup
ScreenshotNeo accepts one GET request with a URL and returns a PNG, JPEG, WebP, or PDF. The following cURL, Python, and Node.js examples capture the target URL. See the ScreenshotNeo API documentation for the API options and request details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners, popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, and failed loads are never billed; response headers identify the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account to get 1,000 screenshots per month with no card.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Jobs remain waiting | Worker is not running, cannot reach Redis, or concurrency is exhausted. | Check worker logs and Redis connectivity; verify the configured queue name matches and inspect active jobs and worker capacity. |
| Navigation times out | Slow site, blocked automation, or a readiness condition that never occurs. | Use a bounded timeout; select a site-appropriate readiness signal, such as a specific selector or a bounded post-navigation delay. Record the final URL and failure class. |
| Network-idle wait never completes | Persistent requests, analytics, streaming, or long polling keep the network active. | Use a more targeted readiness condition. Block unnecessary request types only when that is compatible with the capture’s purpose. |
| Screenshot is blank or incomplete | Capture ran before app rendering or lazy content loaded; page may have failed or returned a bot check. | Wait for a meaningful selector or application-ready signal; for lazy images, scroll or use the product’s full-page behavior and verify the content before capture. |
| Element capture fails | Selector does not match, element is hidden, or layout has not settled. | Wait for the selector, verify visibility and dimensions, then capture that element. Report selector errors in the job result. |
| Browser crashes or worker runs out of memory | Too many concurrent pages, very large full-page screenshots, or unclosed browser resources. | Lower concurrency, cap page dimensions or full-page work, close pages and browsers reliably, and monitor process memory. |
| Duplicate images after client retries | Job submission or result writing is not idempotent. | Accept a caller idempotency key and use it consistently for the job and output object name. |
| Local result file is missing from another worker | Workers use separate local disks. | Write to shared object storage or route result retrieval to the worker that owns the file. |
| Queue grows faster than it drains | Arrival rate exceeds worker capacity, destinations are slow, or retries multiply work. | Limit intake, add measured worker capacity, apply per-host limits, and inspect retry and queue-age metrics before increasing concurrency. |
9. Performance, reliability, and cost
There is no universal throughput figure for a Puppeteer queue: page weight, destination latency, browser resources, screenshot dimensions, and readiness rules all affect how long a job occupies a worker. Benchmark representative pages in the deployment environment. Track queue wait time separately from render time so scaling decisions address the actual bottleneck.
- Performance: reuse browser processes only with page isolation and cleanup; avoid unnecessary full-page captures and oversized device scale factors; use request blocking carefully; set batch limits to prevent bursts from overwhelming workers.
- Reliability: use a durable queue for restart recovery, bounded retries with backoff, timeouts, health checks, idempotent writes, result retention rules, and observable failure categories.
- Cost: self-hosting consumes compute, memory, Redis or queue capacity, storage, and engineering time. Hosted pricing and quotas vary and are not established by this research; check current vendor terms. Do not infer that a queue automatically reduces cost.
10. FAQ
How do I take screenshots of multiple URLs?
Create one job per URL, submit them together, and let workers capture them asynchronously. The sample uses BullMQ’s bulk add method.
Does a queue make screenshot jobs reliable?
A queue provides buffering and job state. Recovery across restarts depends on durable queue configuration, retry policy, idempotency, and result storage.
Should I use network idle for every page?
No. Use it when it reflects readiness for the target site. Pages with persistent traffic often need a selector or another bounded readiness condition.
Can I use a hosted browser and keep Puppeteer?
Yes, some hosted browser session services expose a CDP connection for Puppeteer or Playwright. Confirm session behavior and limits in the provider documentation.


