ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Multiple URLs Using Puppeteer

Capture a list of web pages with one Puppeteer browser, save a separate image for each URL, and handle readiness, failures, and concurrency.

By the ScreenshotNeo team4 October 20269 min read

Use one Puppeteer browser, visit each URL in a loop, and save a screenshot to a unique path. Start sequentially: it keeps browser resource use predictable and lets one failed URL be reported without stopping the rest. The runnable script below creates an output directory, uses a consistent viewport, waits for a page-specific selector when configured, checks HTTP status, and closes each page even when capture fails.

Puppeteer’s screenshot guide demonstrates Page.screenshot() after navigation with waitUntil: 'networkidle2'; it also supports capturing a specific element. A single browser can contain multiple pages. See the Puppeteer Screenshots guide and the Page API.

1. Install Puppeteer and prepare the URL list

Use a current Node.js runtime with ES modules. In a new project, initialize npm and install Puppeteer:

npm init -y
npm install puppeteer

Save the script below as capture.mjs. Puppeteer’s package installs a compatible browser for its standard setup. If your environment supplies its own Chrome or Chromium, configure the executable path for your deployment and ensure that browser version works with your installed Puppeteer.

2. Capture each URL sequentially

import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';

const urls = [
  { url: 'https://example.com' },
  { url: 'https://pptr.dev', readySelector: 'main' },
];

const outputDir = path.resolve('screenshots');
const timeoutMs = 30_000;

await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
const results = [];

try {
  for (const [index, item] of urls.entries()) {
    const page = await browser.newPage();
    const filename = `page-${String(index + 1).padStart(3, '0')}.png`;
    const outputPath = path.join(outputDir, filename);

    try {
      // Apply the viewport before navigation for consistent layout.
      await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
      page.setDefaultNavigationTimeout(timeoutMs);

      const response = await page.goto(item.url, {
        waitUntil: 'networkidle2',
        timeout: timeoutMs,
      });

      if (response && response.status() >= 400) {
        throw new Error(`HTTP ${response.status()} ${response.statusText()}`);
      }

      if (item.readySelector) {
        await page.waitForSelector(item.readySelector, { timeout: timeoutMs });
      }

      await page.screenshot({ path: outputPath, fullPage: true });
      results.push({ url: item.url, ok: true, file: outputPath });
      console.log(`Saved ${item.url} → ${outputPath}`);
    } catch (error) {
      results.push({ url: item.url, ok: false, error: error.message });
      console.error(`Failed ${item.url}: ${error.message}`);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

const failed = results.filter((result) => !result.ok);
console.log(`Finished: ${results.length - failed.length} succeeded, ${failed.length} failed.`);
if (failed.length > 0) process.exitCode = 1;

Run it with node capture.mjs. Output files use list position rather than URL text, so query strings and unusual path characters cannot cause invalid filenames or accidental overwrites. Keep the results records if another process needs a machine-readable summary.

3. Choose when a page is ready

page.goto() accepts a navigation readiness condition. The example uses networkidle2, matching Puppeteer’s screenshot guide. Network idleness is not proof that an application has finished rendering: a page may load data later, or persistent network activity may prevent an idle condition from arriving. When the screenshot depends on a particular component, wait for a selector that represents it, as the script optionally does.

Readiness choice Use when Trade-off
domcontentloaded The document has parsed and you will wait for a specific element or app signal. May be too early for styles, images, or client-rendered content.
load The page’s load event is a useful baseline. Does not guarantee late API data or animations have completed.
networkidle2 You want the guide’s convenient network-idle example. Persistent requests can make it slow or unsuitable; it may still miss later rendering.
Selector or app-ready condition A known element indicates the content is ready. Choose a stable selector, and set a timeout for missing content.

For a fixed delay, use await new Promise(resolve => setTimeout(resolve, milliseconds)) after navigation, but prefer a meaningful condition when one exists. A delay can be wasteful on fast pages and insufficient on slow ones. If you need to click a control that triggers navigation, coordinate the click and navigation wait with Promise.all() to avoid a race; Puppeteer documents this pattern in the Page API.

4. Configure viewport, image, and target

Viewport and device emulation

Set a viewport before navigation when captures should share dimensions or emulate mobile layout. For device emulation, Puppeteer provides known device profiles; emulation sets viewport metrics and user agent, and the API notes that resizing can affect sites, so do it before loading the URL. A full-page screenshot captures beyond the viewport. Without fullPage: true, the screenshot is limited to the visible page area.

Screenshot format and options

The output extension and screenshot options should agree. Puppeteer supports screenshot options for image type and related capture behavior. PNG is a straightforward default; choose JPEG or WebP when your downstream workflow supports those formats and you want smaller image files. Consult the current Page.screenshot() options reference for the exact options accepted by your installed version. The screenshot API returns image bytes when no path is provided; you can save those bytes yourself when integrating with object storage or a queue.

Capture one component instead of the whole page

Wait for the target element, then capture its handle:

const card = await page.waitForSelector('.product-card', { timeout: 15_000 });
if (!card) throw new Error('Product card was not found');
await card.screenshot({ path: outputPath });

Puppeteer’s guide says element screenshotting attempts to scroll a hidden element into view. This is useful for component snapshots; use page.screenshot({ fullPage: true }) when the entire document is the target.

Per-site state and authentication

If pages require a particular locale, authenticated session, cookie, or request header, configure that state before navigation. For example, page.setExtraHTTPHeaders() adds headers to requests, and page.setCookie() is deprecated in favor of browser or browser-context cookie methods in the current API. Avoid putting secrets directly in source code or logs. Do not reuse one user’s authenticated context for unrelated or untrusted URLs.

5. Continue after failures and verify what was captured

The loop catches errors per URL, records the failed address, and moves on. It checks the main navigation response status because a screenshot can still be produced for an HTTP error page. This is important: an image file alone does not prove the destination returned a successful status. Redirects may lead to a different final URL; log page.url() if that distinction matters.

  • Give every input a unique output name. Index-based names are simple; for repeatable builds, combine a sanitized slug with a stable hash.
  • Keep URL, final URL, response status, capture timestamp, and error in your job record when auditing matters.
  • Retry only failures that are plausibly transient, such as a timeout. Use a small retry count and backoff; do not retry invalid URLs or deterministic missing selectors indefinitely.
  • Always close pages in finally and the browser in an outer finally, including when a batch is interrupted by an exception.
  • Validate and constrain input URLs if they come from users. A browser can access internal network addresses; do not expose an unrestricted screenshot worker to arbitrary untrusted URLs.

6. Scale with bounded concurrency

Sequential processing is the safest baseline. A single browser may own multiple pages, but Puppeteer does not specify one concurrency limit that fits every workload. Throughput and memory use depend on the sites, image dimensions, page scripts, and machine. Measure with your own URL mix before increasing parallelism.

A simple worker pool can cap active captures without launching a browser for every URL. Each worker below claims the next list index; each capture still creates and closes its own page:

const concurrency = 3; // Start small and measure on your workload.
let nextIndex = 0;

async function worker() {
  while (true) {
    const index = nextIndex++;
    if (index >= urls.length) return;

    const item = urls[index];
    const page = await browser.newPage();
    try {
      await page.setViewport({ width: 1440, height: 900 });
      const response = await page.goto(item.url, {
        waitUntil: 'networkidle2',
        timeout: timeoutMs,
      });
      if (response && response.status() >= 400) {
        throw new Error(`HTTP ${response.status()}`);
      }
      if (item.readySelector) {
        await page.waitForSelector(item.readySelector, { timeout: timeoutMs });
      }
      await page.screenshot({
        path: path.join(outputDir, `page-${String(index + 1).padStart(3, '0')}.png`),
        fullPage: true,
      });
      results[index] = { url: item.url, ok: true };
    } catch (error) {
      results[index] = { url: item.url, ok: false, error: error.message };
    } finally {
      await page.close();
    }
  }
}

await Promise.all(
  Array.from({ length: Math.min(concurrency, urls.length) }, () => worker()),
);

Keep filenames unique even when jobs run concurrently. Puppeteer documents that creating or closing pages in a browser context waits for an active screenshot to finish, so do not assume those operations are entirely independent. A modest worker cap avoids opening an unbounded number of resource-heavy pages; tune it against completion time, memory, and failure rate.

7. Troubleshooting

Symptom Likely cause Fix
Navigation timeout The site is slow, never becomes network-idle, or the timeout is too short. Use a readiness condition that matches the page, raise the timeout for known slow sites, and record the failure. Avoid waiting for network idle when persistent connections keep it active.
Screenshot is blank or missing content The capture ran before client rendering, a selector was wrong, or an error page loaded. Wait for a content-specific selector, inspect the response status and final URL, and log page errors when diagnosing.
HTTP 403, CAPTCHA, or access-denied page The destination refused automated browsing or requires a visitor/session. Respect the site’s access rules. Do not treat an image of the denial page as a successful content capture; classify and report it.
Output file cannot be written The output directory is missing, the process lacks permission, or the path is invalid. Create the directory recursively, verify write permissions, and use generated filenames rather than raw URLs.
Browser launch fails in a container The runtime lacks compatible browser dependencies or launch configuration. Install the required browser dependencies for that environment and follow Puppeteer’s deployment configuration for the chosen runtime.
Memory rises during a large batch Too many pages are open at once, or a page is unusually heavy. Use sequential capture or lower the worker cap, close every page, and process very large input lists in chunks.
Duplicate or overwritten images Different jobs constructed the same filename. Use an input index plus a stable unique identifier and ensure parallel workers write to distinct paths.
Capture shows the wrong layout Viewport, device metrics, or user agent were applied too late or differ between pages. Set the viewport and emulation consistently before navigation.

8. Performance, reliability, and cost

Self-hosted Puppeteer has no per-screenshot API charge, but the job consumes compute, memory, storage, and engineering time. Browser startup is shared across the sequential batch in the example. Parallel pages can improve throughput when navigation waits dominate, but they also increase resource demand and may encounter destination rate limits; measure rather than assuming a speedup.

For reliability, keep per-URL outcomes, use explicit timeouts, clean up pages and browsers, and retry selectively. Long full-page captures and high device scale factors create larger images and can increase processing and storage needs. Choose the smallest dimensions and format that satisfy the consuming workflow. For recurring jobs, monitor failed URLs and capture duration so changes in the target sites or runtime are visible.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. For each URL, make one request; the response is the image or PDF. Its clean-shot options accept cookie and consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status.

For a multiple-URL workflow, call the API once per URL or use its bulk capture option for up to 100 URLs per call. The ScreenshotNeo API documentation describes the request options. Here is a runnable Node.js example for one URL (repeat it for each item in your list):

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', image));

The API also accepts the common screenshot API parameter names, supports signed links for public image tags and asynchronous jobs with signed webhooks, and offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to get started.

9. FAQ

Can I capture different URLs in one browser?

Yes. Reuse one launched browser and navigate a page for each URL. The simplest batch opens and closes a page per item.

Does a screenshot prove the page loaded successfully?

No. Check the navigation response status and record the final URL if success matters; an error page can also be screenshotted.

Should I use full-page screenshots for every URL?

Only when the full document is needed. Viewport captures are smaller and often sufficient for previews or consistent visual checks.

How many pages should I process at once?

There is no universal number. Start sequentially, then raise a small concurrency cap while measuring memory, duration, and errors on representative URLs.