ScreenshotNeo

BlogComparisons

Puppeteer or Playwright for Screenshotting Thousands of URLs

Choose Puppeteer or Playwright for batch screenshots based on engine needs, team investment, isolation, and repeatability. Here’s a bounded, recoverable Node.js runner.

By the ScreenshotNeo team4 October 202613 min read

Short answer: Both Puppeteer and Playwright can capture screenshots of web pages, and the documentation available does not establish a universal winner or reliable throughput figure for screenshotting thousands of URLs. Choose based on the browser engines your workload needs, your existing code and team experience, how you isolate browser state, and how repeatable the images must be. For either library, use a bounded queue, reuse browser processes across jobs, close each page or context when its job finishes, and record an outcome for every URL.

This guide shows a runnable Node.js batch pattern for both libraries, explains how to scale it safely, and covers the operational choices that matter more than a general speed claim. For the APIs and lifecycle details, see the official Puppeteer screenshots guide, Puppeteer browser management guide, Playwright screenshots guide, and Playwright Page API.

1. How to choose

Decision factor Prefer the library that… Why it matters for a batch
Browser engines Fits your required browser coverage. Playwright’s Page API documents Chromium, Firefox, and WebKit examples; check each project’s current installation documentation for supported engines and versions. Different engines can render a page differently. Don’t assume the screenshot path is interchangeable across engines.
Existing investment Your team already uses, deploys, and understands. Fixtures, language bindings, runtime images, and operational experience affect the work of shipping and maintaining a batch pipeline.
State isolation Supports the context model your jobs need. Separate contexts when cookies or local storage must not carry over between jobs. Puppeteer documents that BrowserContexts isolate cookies and local storage; closing a context closes its pages.
Repeatability Can be pinned to the same runtime and environment used for your expected output. Playwright documents that OS, browser version, settings, hardware, power source, and headless mode can affect rendered output.
Operations You can operate with bounded concurrency and explicit timeouts, retries, and result recording. Neither official documentation set reviewed publishes a safe concurrency limit or reliable jobs-per-second figure for this workload.

Practical rule: If your current library meets your engine, isolation, and repeatability needs, keep it and benchmark your actual targets. If you need to pick from scratch, prototype both against representative URLs and deployment conditions before committing. The evidence here does not support claiming that one is faster for thousands of screenshots.

2. Build a bounded screenshot queue

The runner below accepts one URL per line in urls.txt, reuses a single browser process, limits active pages, writes image files to shots/, and records success or failure per input in results.jsonl. Each job gets its own context so cookies and local storage do not leak from one URL to another. It closes each context after capture and the browser when the batch ends.

Playwright version

Install Playwright and its Chromium browser using the current commands in the official Playwright getting started guide. Save this as batch-playwright.mjs. Set CONCURRENCY and NAV_TIMEOUT_MS in the environment as needed.

import { chromium } from 'playwright';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';

const input = process.argv[2] ?? 'urls.txt';
const concurrency = Math.max(1, Number(process.env.CONCURRENCY ?? 4));
const timeoutMs = Math.max(1, Number(process.env.NAV_TIMEOUT_MS ?? 30000));
const outputDir = process.env.OUTPUT_DIR ?? 'shots';

const urls = (await readFile(input, 'utf8'))
  .split(/\r?\n/).map(s => s.trim()).filter(Boolean);
await mkdir(outputDir, { recursive: true });
await writeFile('results.jsonl', '');
const browser = await chromium.launch({ headless: true });
let next = 0;

async function worker() {
  while (true) {
    const index = next++;
    if (index >= urls.length) return;
    const url = urls[index];
    const started = Date.now();
    const id = createHash('sha256').update(`${index}:${url}`).digest('hex').slice(0, 16);
    let result;
    let context;
    try {
      context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
      const page = await context.newPage();
      page.setDefaultNavigationTimeout(timeoutMs);
      const response = await page.goto(url, { waitUntil: 'load', timeout: timeoutMs });
      await page.screenshot({ path: `${outputDir}/${id}.png`, fullPage: true });
      result = {
        index, url, ok: true, file: `${outputDir}/${id}.png`,
        status: response?.status() ?? null, finalUrl: page.url(),
        elapsedMs: Date.now() - started
      };
    } catch (error) {
      result = {
        index, url, ok: false, error: String(error?.message ?? error),
        elapsedMs: Date.now() - started
      };
    } finally {
      await context?.close().catch(() => {});
    }
    await appendFile('results.jsonl', `${JSON.stringify(result)}\n`);
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
} finally {
  await browser.close();
}

The example uses Chromium and a full-page PNG. You can change the browser import and launch path to evaluate another supported engine, and adjust viewport or screenshot options for your needs. The worker count is a starting point for measurement, not a documented safe maximum. For very long batches, consider writing results through a single queue or a database if multiple processes will append to the same file.

Puppeteer version

Install Puppeteer following its official installation guide. Save as batch-puppeteer.mjs. The same input, environment variables, output naming, and result format apply.

import puppeteer from 'puppeteer';
import { mkdir, readFile, appendFile, writeFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';

const input = process.argv[2] ?? 'urls.txt';
const concurrency = Math.max(1, Number(process.env.CONCURRENCY ?? 4));
const timeoutMs = Math.max(1, Number(process.env.NAV_TIMEOUT_MS ?? 30000));
const outputDir = process.env.OUTPUT_DIR ?? 'shots';

const urls = (await readFile(input, 'utf8'))
  .split(/\r?\n/).map(s => s.trim()).filter(Boolean);
await mkdir(outputDir, { recursive: true });
await writeFile('results.jsonl', '');
const browser = await puppeteer.launch({ headless: true });
let next = 0;

async function worker() {
  while (true) {
    const index = next++;
    if (index >= urls.length) return;
    const url = urls[index];
    const started = Date.now();
    const id = createHash('sha256').update(`${index}:${url}`).digest('hex').slice(0, 16);
    let result;
    let context;
    try {
      context = await browser.createBrowserContext();
      const page = await context.newPage();
      page.setDefaultNavigationTimeout(timeoutMs);
      const response = await page.goto(url, { waitUntil: 'load', timeout: timeoutMs });
      await page.screenshot({ path: `${outputDir}/${id}.png`, fullPage: true });
      result = {
        index, url, ok: true, file: `${outputDir}/${id}.png`,
        status: response?.status() ?? null, finalUrl: page.url(),
        elapsedMs: Date.now() - started
      };
    } catch (error) {
      result = {
        index, url, ok: false, error: String(error?.message ?? error),
        elapsedMs: Date.now() - started
      };
    } finally {
      await context?.close().catch(() => {});
    }
    await appendFile('results.jsonl', `${JSON.stringify(result)}\n`);
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
} finally {
  await browser.close();
}

Puppeteer’s documented BrowserContext behavior makes it a clear fit when each task needs isolated cookies and local storage. The example closes one context per URL. If you intentionally want state shared across a group of pages, change the context lifetime deliberately and document that choice.

3. Run the batch and inspect outcomes

  1. Put one absolute HTTP or HTTPS URL per line in urls.txt. Remove blank lines if they are not meaningful jobs.
  2. Start with a small sample of representative sites. Confirm that the navigation milestone, viewport, full-page behavior, output format, and timeout meet your needs.
  3. Run with a conservative concurrency setting, then increase it in measured steps while tracking failures, elapsed time, image size, memory, and throughput.
  4. Review results.jsonl after each batch. Keep failed records so that retries and reruns are explicit and auditable.
CONCURRENCY=4 NAV_TIMEOUT_MS=30000 node batch-playwright.mjs urls.txt
# or
CONCURRENCY=4 NAV_TIMEOUT_MS=30000 node batch-puppeteer.mjs urls.txt

The code emits a JSON line for each URL. A successful navigation response can still have an HTTP error status, and a navigation timeout does not necessarily tell you whether the target site itself is slow or the chosen milestone was too strict. Inspect the recorded status, final URL, and error. Add a policy for HTTP status handling if your workflow needs to distinguish 4xx/5xx from a successfully captured page.

4. Concurrency, browser reuse, and isolation

  • Bound active work. A queue with a fixed number of workers prevents thousands of pages from opening at once. There is no universal safe concurrency number in the documentation reviewed. Derive a limit from your machine, typical page weight, screenshot size, and target-site behavior.
  • Reuse browser processes. Launching a new browser for every URL adds process startup and teardown work. Reusing a browser for the batch follows the documented browser/page lifecycle and is a practical engineering recommendation, not a published performance guarantee.
  • Close per-job resources. Close the context or page when its capture completes, including on errors. This limits retained page state and makes cleanup predictable. The Puppeteer guide explicitly says closing a BrowserContext closes its pages.
  • Choose state boundaries intentionally. A fresh context for each URL isolates cookies and local storage. If pages must share a login session, use an explicit shared context per account or group and close it when that group ends.
  • Scale across processes only after measuring. Multiple workers can increase capacity, but also multiply browser memory and CPU use. Track resource use and errors before adding workers or machines.

5. Screenshot settings and page readiness

Decide what “the screenshot” means before processing thousands of pages. Both tools expose page screenshot capture; Puppeteer also documents screenshots of selected elements. Playwright documents full-page screenshots. Check each library’s current API documentation for the exact options supported by the installed version.

  • Viewport: Fix width and height for comparable captures. Responsive layouts, menus, and breakpoints change with viewport size.
  • Full page or viewport: Full-page captures include content beyond the initial viewport and can produce much taller files. Use viewport screenshots when only the initial view matters.
  • Element capture: Use an element screenshot when the output should focus on a known component. Confirm the selector resolves and the element is visible before capture.
  • Readiness milestone: The sample waits for the page’s load event. Some pages load content later; others keep network activity open indefinitely. Choose a milestone appropriate to the site and add a targeted selector or short delay only when required. Avoid assuming that all network activity becoming idle means the page is visually complete.
  • Lazy content: Full-page capture does not guarantee every lazy-loaded image has loaded. If complete below-the-fold content matters, use a deliberate scrolling or site-specific readiness strategy and verify representative pages.
  • Dynamic content: Ads, timestamps, rotating content, and personalized state can change between captures. Control the environment and state where possible, and record that perfect pixel equality may not be realistic.
  • Output: PNG is lossless and can be large. Choose another supported format only if your chosen library version and downstream workflow support it.

6. Repeatability and visual comparisons

Playwright’s visual-comparison guide documents that rendered output can vary with host OS, browser version, settings, hardware, power source, and headless mode. For stable comparisons, use the same environment that generated the baseline. Pin the browser/runtime and operating-system image, keep fonts and viewport consistent, and control relevant settings. Recheck screenshots after changing the browser version or deployment image rather than interpreting every difference as a site change.

7. Reliability and retry policy

Thousands of independent jobs will encounter a range of outcomes. Preserve one result per input URL and make retry decisions from recorded failures rather than silently dropping work.

  • Record input URL, final URL, start or elapsed time, navigation status, output path, and failure reason.
  • Retry only failures that your policy classifies as transient, such as a timeout you are willing to try again. Set a retry limit and a delay strategy so a failing destination cannot loop forever.
  • Do not automatically retry every HTTP status. Decide whether an error page is a valid artifact for your use case or a failed job.
  • Make reruns safe. Stable output names or a manifest should let you identify completed jobs without confusing old and new screenshots.
  • Handle process interruption. Persist results as jobs finish, so a crash does not erase the status of earlier URLs.
  • Respect the target sites’ access policies and avoid using automation to bypass access controls. The library documentation does not establish a site-specific permission conclusion.

8. Performance and cost planning

No published concurrency ceiling, stable throughput, or jobs-per-second figure for this workload is established by the official material reviewed. Estimate your own capacity with a representative pilot rather than extrapolating from a small, fast set of pages.

  1. Choose a representative sample that includes typical and unusually slow or long pages.
  2. Run it at a low concurrency and record elapsed time per URL, failure rate, image dimensions and file size, and process memory/CPU.
  3. Increase concurrency gradually. Stop increasing when failures rise, memory pressure becomes unstable, or the throughput improvement no longer justifies the resource use.
  4. Repeat the measurement in the production-like browser and operating-system image. Browser version and headless mode can change rendering as well as operational behavior.
  5. Include retries and failed jobs in capacity planning; they consume time even when they do not produce a usable image.

With self-hosted browser workers, costs depend on the compute, storage, and engineering effort you choose; the sources provided do not quantify them. Full-page PNGs, retained browser processes, and high concurrency can each increase resource needs. Set output retention and storage limits based on how long artifacts must remain available.

9. Common problems and fixes

Symptom Likely cause Fix
Browser executable missing or launch fails The installed browser/runtime does not match the library installation or deployment image. Follow the current official installation instructions for the library, install the intended browser, and keep the runtime image consistent.
Navigation times out The site is slow, the timeout is too short, or the chosen readiness event waits for resources that do not settle. Inspect the URL and error, choose an appropriate readiness milestone, increase the timeout only for justified cases, and record a bounded retry if policy allows.
Blank or incomplete capture Capture ran before the relevant content appeared, or the page depends on delayed/lazy content. Wait for a meaningful selector or application-ready condition; for lazy content, use a tested scroll strategy and inspect sample captures.
Memory use grows during the batch Pages, contexts, or browser processes are not being closed, or concurrency is too high. Ensure cleanup runs in finally, limit workers, and measure whether a browser restart after a bounded number of jobs is needed for your workload.
Different screenshots from one run to another Dynamic page content or rendering environment differs. Pin browser/runtime, OS image, viewport, fonts, settings, and state; compare only like-for-like environments.
Output files overwrite each other Names derive only from the URL or use a lossy transformation. Use stable unique identifiers, such as an input index plus a hash, and preserve the URL-to-file mapping in the results manifest.
Some input lines fail immediately Malformed URLs, unsupported schemes, or blank lines. Validate URLs before enqueueing, accept only expected HTTP(S) schemes, and write invalid inputs as explicit failures.
Batch stops before all URLs finish An uncaught worker error or process interruption aborts the run. Catch per-URL failures, persist each result, and maintain a manifest so unfinished inputs can be resumed safely.

10. ScreenshotNeo: a managed alternative

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. If you want to avoid operating browser workers, it accepts one GET request with a URL and returns a PNG, JPEG, WebP, or PDF. Its parameter names also work with those used by other screenshot APIs, which can make switching easier.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options and response details. Before capture, it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

ScreenshotNeo offers 1,000 screenshots a month free with no card. Paid plans are Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. For a batch, compare the API’s per-request workflow and published plan allowance with the operational work of running your own workers.

Or skip the browser setup

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Sign up for 1,000 free screenshots a month, with no card required.

11. Frequently asked questions

How many pages can Playwright run at once?

There is no universal concurrency number established by the documentation reviewed. Start small and measure on your deployment image and representative sites; capacity depends on the pages and resources available.

Should I reuse one browser for multiple screenshots?

For a batch, reusing a browser process while closing each job’s page or context is a practical pattern. Use fresh contexts when jobs need isolated cookies and local storage.

Which library makes screenshots faster?

The available official documentation does not establish a reliable speed winner for thousands of URLs. Benchmark the same representative workload, browser version, host, viewport, and capture settings.

Does a full-page screenshot include every lazy-loaded image?

Not necessarily. A full-page capture covers page dimensions, but lazy content may require scrolling or a site-specific readiness condition before capture.

Can I get identical screenshots on different machines?

Not reliably by default. Rendering can vary with the operating system, browser version, settings, hardware, power source, and headless mode. Use a pinned environment for comparisons.

Sources