ScreenshotNeo

BlogHow-to

How to Take Bulk Screenshots of URLs with Node.js and Puppeteer

Build a reliable Node.js screenshot job with Puppeteer: process URLs safely, handle failures, and tune capture options and concurrency.

By the ScreenshotNeo team4 October 20269 min read

To take bulk screenshots with Node.js and Puppeteer, launch one browser, process URLs sequentially or through a small bounded worker pool, navigate each page, save its screenshot, record its result, and close every page and the browser in finally blocks. Check the HTTP response status as well as navigation errors: a resolved navigation does not necessarily mean the page returned a successful status.

This guide uses Puppeteer’s ES module API. The examples write PNG files into a local directory; change the format, page scope, wait condition, and concurrency to fit the target sites and machine. See the Puppeteer screenshot guide and Page.goto reference for version-specific details.

1. Install Puppeteer and prepare the input

Create a project and install Puppeteer. Its install process normally downloads a compatible browser. Ensure that download and the browser cache are available in the environment where the job will run; a package install that skips the browser download can leave the runtime without an executable.

mkdir bulk-shots
cd bulk-shots
npm init -y
npm install puppeteer

Save the main example below as capture.mjs. Put the URLs you want to capture in the urls array, or replace it with input from a file, database, or queue. Validate externally supplied URLs before visiting them, and decide whether private network addresses are allowed in your environment.

2. Complete bulk capture script

This runnable script creates one page per active task, limits concurrency, checks HTTP status where a response is available, stores a result for every input URL, and closes pages and the browser even when a capture fails. The concurrency value of three is only a conservative starting point, not a universal recommendation.

import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';

const urls = [
  'https://example.com',
  'https://pptr.dev',
];
const outputDir = path.resolve('screenshots');
const concurrency = 3;
const navigationTimeoutMs = 30_000;

function safeName(url, index) {
  const parsed = new URL(url);
  const host = parsed.hostname.replace(/[^a-z0-9.-]/gi, '_');
  return `${String(index + 1).padStart(3, '0')}-${host}.png`;
}

async function mapWithLimit(items, limit, worker) {
  if (!Number.isInteger(limit) || limit < 1) {
    throw new Error('concurrency must be a positive integer');
  }
  const results = new Array(items.length);
  let next = 0;

  async function runWorker() {
    while (true) {
      const index = next++;
      if (index >= items.length) return;
      try {
        results[index] = await worker(items[index], index);
      } catch (error) {
        results[index] = {
          ok: false,
          url: items[index],
          error: error instanceof Error ? error.message : String(error),
        };
      }
    }
  }

  await Promise.all(
    Array.from({ length: Math.min(limit, items.length) }, () => runWorker()),
  );
  return results;
}

await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
let results;

try {
  results = await mapWithLimit(urls, concurrency, async (url, index) => {
    const page = await browser.newPage();
    try {
      const response = await page.goto(url, {
        waitUntil: 'networkidle2',
        timeout: navigationTimeoutMs,
      });
      const status = response?.status();
      if (status !== undefined && status >= 400) {
        throw new Error(`HTTP ${status}`);
      }

      const file = path.join(outputDir, safeName(url, index));
      await page.screenshot({ path: file, fullPage: true });
      return { ok: true, url, file, status, finalUrl: page.url() };
    } finally {
      await page.close();
    }
  });
} finally {
  await browser.close();
}

for (const result of results) {
  console.log(JSON.stringify(result));
}

Run it with node capture.mjs. Each result is logged as JSON, so it can be redirected to a file or consumed by another process. A failed URL does not stop the remaining workers. Browser startup failure, however, prevents the job from starting and should be reported by the calling process.

3. Choose a processing model

Sequential processing

For small batches or initial debugging, set concurrency to 1. This is easiest to reason about and uses fewer simultaneous page resources, at the cost of finishing one URL before opening the next.

Bounded concurrency

A worker pool overlaps navigation and rendering while putting a ceiling on active pages. Puppeteer supports multiple Page instances in a browser, but its documentation does not establish a universally safe concurrency number. Full-page captures, large images, complex scripts, available memory, and destination-site limits all affect the useful setting. Start low, observe memory and failures, and increase gradually.

Reuse one page or create pages per task

The example creates and closes a page for each active URL, which isolates page state and makes cleanup local to a task. Reusing a page can reduce page creation overhead, but be careful about cookies, storage, service workers, and other state carrying between destinations. Do not run concurrent navigations on a single page.

4. Navigation readiness and HTTP status

page.goto() can wait for different navigation milestones. Select one based on what the page needs before capture:

Option Use Trade-off
domcontentloaded The document structure is sufficient or the site has its own readiness signal. Images and later scripts may still be loading.
load The page’s load event is a reasonable baseline. Does not guarantee that client-rendered content is finished.
networkidle0 / networkidle2 Network activity is expected to settle. Long polling, analytics, and streaming requests can keep activity open; idle does not prove application readiness.

For a page that renders content after navigation, wait for a selector or an application-specific condition after goto(), for example:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.waitForSelector('main article', { timeout: 10_000 });

Use the selector that signals the content you actually need. A fixed delay can be useful for a known animation or delayed widget, but it wastes time on fast pages and can still be too short on slow ones.

Keep the response returned by goto(). Navigation completing is separate from the HTTP status: inspect response.status() and decide whether 4xx or 5xx responses count as failed captures. Some headless-shell behavior can return such responses without throwing. Redirects may mean the final page URL differs from the requested URL; log both if that matters.

5. Screenshot scope, format, and file names

The screenshot API offers several useful options. See the Puppeteer screenshot options reference for the complete, version-specific list.

Option Behavior Notes
path Writes the image to a file. Relative paths resolve from the process working directory. The output directory must exist; the example creates it first.
fullPage Captures the full page rather than only the viewport. Can produce very tall and memory-heavy images on long pages.
type Selects PNG, JPEG, or WebP where supported by the installed Puppeteer version. Use a matching filename extension. If omitted, the path extension can determine the type.
quality Sets lossy image quality for supported formats. Does not affect PNG output.
clip Captures a specified rectangle. Useful for a region; it is a different scope from a full-page capture.

Host-only filenames collide if a batch has multiple paths or query variants from the same host. The example adds a zero-padded input index to keep names distinct and predictable. For repeated runs, consider a stable hash of the full URL or an output manifest; avoid putting raw untrusted URLs into filesystem paths. Decide whether reruns overwrite earlier files or use run-specific directories.

6. Reliability, performance, and cost

  • Isolate failures: return a result for each input rather than aborting the whole batch on one navigation or screenshot error.
  • Close resources: close each page in finally and the browser in an outer finally. This limits orphaned Chromium processes after task failures.
  • Bound work: unbounded page creation can exhaust memory or overload destination sites. Tune concurrency for the machine and workload; there is no source-backed universal value.
  • Set timeouts: use an explicit navigation timeout and, where applicable, a separate timeout for application readiness. Retry only errors that are plausibly transient, and cap retries so one URL cannot hold the batch indefinitely.
  • Record useful context: log input URL, final URL, HTTP status, output path, elapsed time, and error. Avoid logging credentials or sensitive query parameters.
  • Expect variable captures: pages may change, personalize content, lazy-load images, or block automation. Use consistent viewport, wait conditions, locale/session settings, and timestamps when comparing runs.
  • Budget for infrastructure: self-hosted Puppeteer has no per-screenshot API charge, but the job still consumes compute, memory, storage, and maintenance time for Chromium and deployment. Managed services trade browser operations for their own usage and price terms; compare those terms directly.

7. Troubleshooting

Symptom Likely cause Fix
TimeoutError from goto() The site is slow, navigation is stuck, or the selected readiness event never occurs. Set an explicit timeout; choose a narrower wait condition and then wait for a page-specific selector. Investigate whether the URL is reachable from the runtime.
A screenshot is saved for an error page Navigation resolved, but the server returned 404 or 500. Inspect the response status from goto() and record or reject status codes according to the job’s policy.
“Could not find Chrome” or executable missing The browser download was skipped, or the deployment image lacks Puppeteer’s compatible browser. Install Puppeteer and its browser in the final runtime environment, preserve the expected browser cache, and check the installation guidance. Puppeteer only guarantees compatibility with its bundled browser.
Memory usage climbs or the browser crashes Too many simultaneous pages or very large full-page captures. Lower concurrency, capture viewport or a clip when sufficient, and split very large batches into smaller jobs.
Content is missing from the image Capture happened before client rendering, lazy images, or a target element was ready. Wait for the relevant selector or application-ready signal; use full-page capture when content below the fold is required.
Different URLs overwrite the same file Filename generation uses a non-unique component such as hostname alone. Add the input index, a stable URL hash, or a collision-resistant identifier and keep a manifest mapping names to URLs.
Automation is blocked or a CAPTCHA appears The destination site challenged the browser or restricted automated access. Respect the site’s access rules and use authorized access where available. Do not treat a challenge page as a successful content capture.

8. Or skip the browser setup

If you would rather not install, operate, and scale Chromium, ScreenshotNeo accepts a URL in one API request and returns an image or PDF. Its screenshot API options include full-page capture, device and viewport settings, selectors, waits, and output controls; see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before capture; known consent platforms, newsletter popups, and chat widgets are removed. Each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.

9. Frequently asked questions

Can one Puppeteer browser capture multiple URLs at once?

Yes. Create multiple pages in one browser and limit how many are active. More concurrency is not automatically faster if rendering or memory becomes the bottleneck.

Does networkidle2 guarantee all page content is ready?

No. It describes network activity, not whether a particular application has finished rendering. Wait for a selector or application-specific signal when capture completeness depends on it.

Can Puppeteer save screenshots without writing directly to disk?

Yes. The screenshot API can return image data when no path is supplied, which is useful when uploading bytes to object storage or passing them to another process. Check the return type for your installed version.

Should I retry failed URLs?

Retry transient navigation or infrastructure failures with a small capped policy. A repeated 404, invalid URL, or access denial is unlikely to improve with immediate retries; preserve the error and continue the batch.