ScreenshotNeo

BlogHow-to

How to Capture Screenshots from Many URLs Efficiently with Puppeteer

Capture a list of URLs with Puppeteer using a reusable browser, bounded workers, deliberate readiness checks, and per-URL error handling.

By the ScreenshotNeo team30 September 202610 min read

How to Capture Screenshots from Many URLs Efficiently with Puppeteer

To capture screenshots from many URLs efficiently with Puppeteer, launch one browser for the batch, create a page for each active job, and limit how many jobs run at once. Use a sequential loop for a short list; use a bounded worker pool for larger lists. Choose a navigation readiness condition that fits each site, save each result under a unique filename, record failures per URL, and close every page even when a capture fails.

Puppeteer’s capture primitive is page.screenshot(). Its official guide shows navigating to a page and saving an image with { path }; a Browser can contain multiple Page instances. Reusing one browser for a batch is a practical design recommendation based on that model, not an officially benchmarked speed claim. See the Puppeteer screenshot guide and Page API.

1. Install Puppeteer and prepare your URL list

Use a Node.js project. Puppeteer’s package installation and browser setup can vary by environment; its getting started guide explains the supported installation paths. The example below imports the package as an ES module. Save it as capture.mjs, create an urls.txt file with one complete URL per line, then run node capture.mjs.

https://example.com
https://developer.mozilla.org/
https://pptr.dev/

The script creates an output directory and names screenshots by list index plus a sanitized hostname. The index keeps filenames unique even when the input contains the same host more than once. Validate that inputs include a scheme such as https://; Puppeteer’s page.goto() expects a URL with a scheme.

2. Choose sequential capture or bounded concurrency

A sequential loop is simplest and is often the right first implementation: only one page is active, failures are easy to trace, and the script puts less simultaneous load on your machine and destination sites. It also takes roughly the sum of the time each navigation and screenshot needs.

A worker pool processes several URLs at a time. This can reduce total elapsed time when pages spend time waiting on navigation or remote assets, while putting more concurrent work on the browser, host machine, and target sites. Puppeteer documentation confirms that a browser can have multiple pages, but the reviewed sources do not publish an optimal concurrency number or throughput benchmark. Begin with a small configurable limit, measure your own workload, and respect each site’s operational limits.

3. Runnable Puppeteer batch script

This complete script launches one browser, runs a bounded number of workers, assigns each URL its own page, checks navigation status, optionally waits for a site-specific selector, saves a screenshot, and writes a JSON report. Each worker catches errors separately so one bad URL does not discard successful captures. Change the viewport, output format, readiness mode, selector, and worker count to suit the job.

import puppeteer from 'puppeteer';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const INPUT = 'urls.txt';
const OUTPUT_DIR = 'screenshots';
const REPORT = 'screenshot-results.json';
const CONCURRENCY = 3; // Tune against your workload and machine.
const VIEWPORT = { width: 1440, height: 900, deviceScaleFactor: 1 };
const WAIT_UNTIL = 'domcontentloaded'; // Or 'load' / 'networkidle2'.
const READY_SELECTOR = ''; // Example: 'main article'; empty disables selector wait.
const FULL_PAGE = true;
const FORMAT = 'png'; // png, jpeg, or webp where supported by the installed version.

function safePart(value) {
  return value.replace(/[^a-z0-9.-]+/gi, '_').replace(/^\.+/, '').slice(0, 100) || 'page';
}

const urls = (await readFile(INPUT, 'utf8'))
  .split(/\r?\n/).map(line => line.trim()).filter(Boolean);
await mkdir(OUTPUT_DIR, { recursive: true });

const results = new Array(urls.length);
let nextIndex = 0;
const browser = await puppeteer.launch({ headless: true });

async function worker() {
  while (true) {
    const index = nextIndex++;
    if (index >= urls.length) return;
    const url = urls[index];
    let page;
    try {
      page = await browser.newPage();
      await page.setViewport(VIEWPORT);
      page.setDefaultNavigationTimeout(45000);
      const response = await page.goto(url, { waitUntil: WAIT_UNTIL, timeout: 45000 });
      const status = response?.status() ?? null;
      if (status !== null && status >= 400) {
        throw new Error(`HTTP ${status}`);
      }
      if (READY_SELECTOR) {
        await page.waitForSelector(READY_SELECTOR, { visible: true, timeout: 15000 });
      }
      const host = safePart(new URL(url).hostname);
      const file = path.join(OUTPUT_DIR, `${String(index + 1).padStart(4, '0')}-${host}.${FORMAT}`);
      const screenshotOptions = { path: file, type: FORMAT, fullPage: FULL_PAGE };
      if (FORMAT !== 'png') screenshotOptions.quality = 85;
      await page.screenshot(screenshotOptions);
      results[index] = { url, ok: true, status, file };
    } catch (error) {
      results[index] = { url, ok: false, error: error.message };
    } finally {
      if (page) await page.close().catch(() => {});
    }
  }
}

try {
  await Promise.all(Array.from(
    { length: Math.min(CONCURRENCY, urls.length) },
    () => worker(),
  ));
} finally {
  await browser.close();
}

await writeFile(REPORT, JSON.stringify(results, null, 2));
const failed = results.filter(result => !result.ok).length;
console.log(`Saved ${results.length - failed}/${results.length} screenshots; see ${REPORT}.`);
if (failed) process.exitCode = 1;

For a small batch, set CONCURRENCY to 1. The worker structure still functions, but a plain loop can be clearer when you want stop-on-first-error behavior. The script deliberately treats an HTTP status of 400 or higher as a failed capture. A navigation response and a page load are different signals: Puppeteer may resolve navigation for HTTP error responses, so inspect response.status() when status matters. See Page.goto().

4. Make readiness, viewport, and output choices deliberate

waitUntil controls when goto() considers navigation complete. The common options include domcontentloaded, load, networkidle0, and networkidle2. Puppeteer’s screenshot guide demonstrates networkidle2, but it is an example, not a guarantee that every site’s application data or images are ready. Sites with long polling or analytics can keep network activity going. A known completion element is often a better signal: wait for a selector that appears after the content you need is rendered. Selector waits have a timeout and will fail if the selector never appears.

Viewport and full-page behavior

Set the viewport before navigation if consistent dimensions matter. A fixed viewport makes captures more comparable; a larger deviceScaleFactor produces a higher-density image and can increase pixel dimensions and storage. fullPage: true captures the whole document, which is useful for archival or review, but can produce very tall, larger images. Use the default viewport-sized capture when only the initial screen matters.

Screenshot options

Option What it controls Use it when
path Writes the image to a file; relative paths resolve from the current working directory. You need durable files rather than returned image bytes.
type Image format; PNG is the documented default. You need to choose an output format explicitly.
quality Quality from 0–100; does not apply to PNG. You use a lossy format and want to manage file size.
fullPage Captures the full page rather than only the viewport. You need the whole document in one capture.
clip Limits capture to a specified rectangular region. You need a consistent crop or a particular screen area.
omitBackground Hides the default white background, allowing transparency where supported. You need a transparent capture.
encoding Returns binary image data by default, or base64 when requested. You need image bytes in memory instead of a file.

These settings are documented in Puppeteer’s ScreenshotOptions reference. For a single element, locate it and call elementHandle.screenshot(); Puppeteer attempts to scroll it into view, and capture can throw if the element was detached from the DOM. See the screenshot guide.

5. When to use separate browser contexts

Pages created in the default context share its browser state. If URLs must not share cookies, local storage, or other context storage, create a separate browser context for each isolated unit of work, then create the page from that context and close the context afterward. Contexts are also useful when a batch has distinct authenticated sessions. They add setup and resource overhead, so use them for isolation rather than by default. Puppeteer documents that separate contexts have isolated storage in its BrowserContext API.

For authenticated pages, apply the required cookies or headers before navigation, and avoid putting secrets in filenames or reports. If multiple URLs belong to the same session, keep them in that session’s context. Avoid sharing a mutable page between concurrent jobs; each worker should own and close its own page.

6. Operational checklist for large batches

  • Normalize the input list and decide what to do with duplicate URLs.
  • Use deterministic viewport, format, full-page setting, and filename policy.
  • Set a bounded concurrency value; increase it only after measuring CPU, memory, elapsed time, and failure rate on representative pages.
  • Set navigation and selector timeouts that match the workload. Avoid disabling timeouts for an unbounded batch.
  • Keep a per-URL record for success, HTTP status, output path, or error.
  • Close each page in a finally path and close the browser after workers settle.
  • Retry only failures you have classified as transient, with a retry limit and delay. Repeatedly retrying a blocked or invalid URL wastes time and can burden a site.
  • Respect the target sites’ terms and rate limits. A screenshot workflow does not grant permission to access a page.

Page creation and closure coordinate with an in-progress screenshot in a browser context, according to the Page.screenshot() API. Still, awaiting each screenshot before closing its page keeps the code’s lifecycle clear.

7. Performance, reliability, and cost

There is no documentation-backed universal worker count. More parallel pages can help overlap waiting, but they also increase simultaneous browser work, network requests, and memory pressure. Benchmark on representative pages, not only fast static pages. Track total batch duration, per-URL duration, error categories, and machine resource use. If the machine starts swapping, pages time out more often, or the target sites begin returning errors, reduce concurrency.

Sequential and parallel batches have different failure and load profiles. A worker pool isolates individual errors and allows the rest of the queue to proceed; a sequential script is easier to inspect and limits simultaneous requests. The right choice depends on the batch size, page complexity, target-site limits, and host capacity. Puppeteer itself has no per-screenshot usage charge described in the cited API docs; your costs are operational, such as compute, storage, and network usage in your environment. Browser launch options document a default startup timeout of 30 seconds and warn that an alternate executable path is not guaranteed to work like Puppeteer’s bundled browser. See launch options.

8. Troubleshooting common failures

Symptom Likely cause Fix
Navigation timeout The page never reached the selected lifecycle event, the site is slow, or ongoing requests prevent network idle. Choose a less strict waitUntil, raise the timeout for that workload, or wait for a meaningful selector after navigation.
Screenshot is blank or missing application content Navigation completed before client-side rendering or a required element appeared. Wait for a site-specific selector or application-ready condition; verify the selector is present on every URL.
Screenshot saved despite an error page An HTTP error status may still produce a navigation response. Inspect response.status() and decide which statuses count as failure for your use case.
File overwritten or name collision Names were derived only from hostname or URL paths can sanitize to the same string. Include a stable index or unique ID in every filename, as in the example.
Memory grows or browser becomes unstable Too many simultaneous pages or pages not closed after exceptions. Lower concurrency and keep page cleanup in finally; close the browser when the batch ends.
Element screenshot throws The element was detached or replaced while the page was changing. Wait for a stable selector immediately before capture and reacquire the handle if the page updates.
Browser fails to launch Browser installation, environment dependencies, startup timeout, or a mismatched executable path. Follow Puppeteer’s installation guide, use its bundled browser where possible, and inspect launch options and environment logs.

9. Or skip the browser setup

If you want an API call instead of maintaining a Puppeteer browser, ScreenshotNeo accepts a URL and returns a screenshot or PDF. Its capture options include full-page screenshots, CSS selector captures, format and viewport settings, custom waits, headers and cookies. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed; response headers say which outcome occurred.
  • An MCP server lets Claude, Cursor, and other MCP clients use screenshot tools.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, no card required.

10. FAQ

Can I capture the same URL more than once?

Yes. Keep duplicate entries if you need separate files or captures at different times; the stable input index gives each run entry a distinct output name.

Should I use networkidle2 for every site?

No. It is shown in Puppeteer’s example, but page behavior varies. Use the condition that corresponds to the content you need, and wait for a meaningful selector when that is more reliable.

Can I capture only one component instead of a full page?

Yes. Wait for the target element and use its screenshot() method, or pass a clip rectangle to page capture when you need a fixed region.

What is the fastest safe concurrency setting?

There is no official universal value. Use a small limit and measure on the actual browser, pages, and machine; raise or lower it based on throughput, resource use, and failures.

Does a successful screenshot prove the page is correct?

No. It proves that Puppeteer produced image data. For quality checks, separately validate HTTP status, expected page markers, output existence, and any visual criteria your workflow requires.