ScreenshotNeo

BlogHow-to

Best Way to Bulk Screenshot Hundreds of URLs with Playwright

Build a reliable Playwright screenshot batch with bounded concurrency, deliberate page readiness, safe filenames, and per-URL error handling.

By the ScreenshotNeo team4 October 202611 min read

The best way to screenshot hundreds of URLs with Playwright is to reuse one browser, process URLs through a bounded queue, and record a separate result for every URL. Start with a small concurrency limit, choose a page-specific readiness signal where possible, and raise concurrency only after measuring the batch on the machine and sites you actually use. There is no universal worker count or published throughput guarantee for arbitrary pages.

This guide uses a custom Node.js queue because it makes input handling, output naming, retries, and per-URL reporting explicit. Playwright Test is another good fit when each URL is naturally a test and you want its runner and reports.

1. Install Playwright and prepare a URL list

Use a current Node.js release supported by your installed Playwright version. In a new project, install Playwright and its Chromium browser:

npm init -y
npm install playwright
npx playwright install chromium

Create urls.txt with one absolute HTTP or HTTPS URL per line. Blank lines and lines starting with # are ignored.

https://example.com/
https://developer.mozilla.org/
# Add more URLs here

The script below creates an output directory, validates URLs, derives collision-resistant filenames from each URL, and writes a JSON report containing each input’s status and output path. Keep the report: it lets you rerun only failures and trace a file back to its source URL.

2. Run a bounded screenshot queue

Save this as bulk-screenshot.mjs. It reuses one browser process and gives each active job its own page in a shared context. Set ISOLATE_CONTEXTS=1 to use a separate context per URL when cookies, local storage, or other browser state must not be shared.

import { createHash } from 'node:crypto';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { chromium } from 'playwright';

const inputFile = process.env.URLS_FILE ?? 'urls.txt';
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const concurrency = positiveInt(process.env.CONCURRENCY, 4);
const timeoutMs = positiveInt(process.env.TIMEOUT_MS, 45_000);
const fullPage = process.env.FULL_PAGE === '1';
const waitUntil = process.env.WAIT_UNTIL ?? 'domcontentloaded';
const isolateContexts = process.env.ISOLATE_CONTEXTS === '1';
const format = (process.env.FORMAT ?? 'png').toLowerCase();
const quality = positiveInt(process.env.QUALITY, 82);

if (!['png', 'jpeg', 'webp'].includes(format)) {
  throw new Error('FORMAT must be png, jpeg, or webp');
}
if (!['commit', 'domcontentloaded', 'load', 'networkidle'].includes(waitUntil)) {
  throw new Error('WAIT_UNTIL must be commit, domcontentloaded, load, or networkidle');
}
if (format !== 'png' && (quality < 0 || quality > 100)) {
  throw new Error('QUALITY must be between 0 and 100');
}

const lines = (await readFile(inputFile, 'utf8')).split(/\r?\n/);
const inputs = lines.map(s => s.trim()).filter(s => s && !s.startsWith('#'));
const jobs = inputs.map((input, index) => {
  try {
    const parsed = new URL(input);
    if (!['http:', 'https:'].includes(parsed.protocol)) throw new Error('URL must use HTTP or HTTPS');
    const safeLabel = parsed.hostname.replace(/[^a-z0-9.-]/gi, '_').slice(0, 60) || 'page';
    const digest = createHash('sha256').update(input).digest('hex').slice(0, 12);
    return { index, input, output: path.join(outputDir, `${String(index + 1).padStart(4, '0')}-${safeLabel}-${digest}.${format}`) };
  } catch (error) {
    return { index, input, validationError: error.message };
  }
});

await mkdir(outputDir, { recursive: true });
const results = new Array(jobs.length);
for (const job of jobs) {
  if (job.validationError) results[job.index] = { url: job.input, status: 'invalid', error: job.validationError };
}

const browser = await chromium.launch({ headless: true });
const sharedContext = isolateContexts ? null : await browser.newContext({ viewport: { width: 1440, height: 1000 } });
let next = 0;

async function capture(job) {
  let context = sharedContext;
  let page;
  try {
    if (!context) context = await browser.newContext({ viewport: { width: 1440, height: 1000 } });
    page = await context.newPage();
    page.setDefaultNavigationTimeout(timeoutMs);
    page.setDefaultTimeout(timeoutMs);
    const response = await page.goto(job.input, { waitUntil, timeout: timeoutMs });
    // A 4xx/5xx response still has a page that may be useful to capture.
    await page.screenshot({
      path: job.output,
      fullPage,
      type: format,
      ...(format === 'png' ? {} : { quality }),
      animations: 'disabled'
    });
    results[job.index] = {
      url: job.input,
      status: 'ok',
      httpStatus: response?.status() ?? null,
      output: job.output
    };
  } catch (error) {
    results[job.index] = { url: job.input, status: 'error', error: String(error), output: job.output };
  } finally {
    if (page) await page.close().catch(() => {});
    if (isolateContexts && context) await context.close().catch(() => {});
  }
}

async function worker() {
  while (true) {
    const index = next++;
    if (index >= jobs.length) return;
    const job = jobs[index];
    if (!job.validationError) await capture(job);
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, jobs.length || 1) }, worker));
} finally {
  if (sharedContext) await sharedContext.close();
  await browser.close();
}

const reportPath = path.join(outputDir, 'results.json');
await writeFile(reportPath, JSON.stringify(results, null, 2) + '\n');
const failed = results.filter(r => r?.status !== 'ok').length;
console.log(`Finished ${jobs.length} URLs: ${jobs.length - failed} succeeded, ${failed} failed. Report: ${reportPath}`);
if (failed) process.exitCode = 1;

function positiveInt(value, fallback) {
  if (value === undefined) return fallback;
  const parsed = Number(value);
  if (!Number.isInteger(parsed) || parsed < 0) throw new Error(`Expected a non-negative integer, got ${value}`);
  return parsed;
}

Run it with the defaults, or set options for your workload:

node bulk-screenshot.mjs
CONCURRENCY=2 FULL_PAGE=1 FORMAT=webp QUALITY=80 node bulk-screenshot.mjs
ISOLATE_CONTEXTS=1 TIMEOUT_MS=60000 node bulk-screenshot.mjs

Configuration: CONCURRENCY limits simultaneous pages; TIMEOUT_MS bounds navigation and page waits; WAIT_UNTIL chooses the navigation event; FULL_PAGE=1 captures the full scrollable document; FORMAT accepts PNG, JPEG, or WebP; and QUALITY applies to JPEG or WebP. Change the viewport in newContext for your target layout. The script uses CSS-pixel scale by default, which avoids device-pixel-sized output; configure context device scale if high-DPI output is required.

3. Choose readiness, capture size, and isolation deliberately

page.goto supports commit, domcontentloaded, load, and networkidle. The default in this example is domcontentloaded, a useful starting point for a mixed list, but it does not guarantee that client-rendered content or lazy images are ready. If a site has a reliable page-specific cue, wait for it before capture:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.locator('[data-page-ready="true"]').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: outputPath, fullPage: true });

Replace the selector with a real signal from your application. A bounded delay can help with a known animation or delayed widget, but it adds time to every URL and does not prove readiness. Playwright discourages using networkidle as a general readiness signal; pages with polling, analytics, or persistent connections may never reach it. Prefer an application-specific state when available. See the official Page API navigation reference.

Viewport, full page, element, and image format

  • Viewport screenshot: the default captures the visible viewport. It is faster and smaller when only the first screen matters.
  • Full page: pass fullPage: true to capture the full scrollable document. Long pages can use substantially more memory and produce larger files.
  • Element: use a locator screenshot when only one component is needed: await page.locator('main article').screenshot({ path: outputPath }).
  • PNG: lossless output, useful for text and visual comparisons, often larger.
  • JPEG or WebP: set type and a quality from 0 to 100 to trade fidelity for smaller output. These formats are lossy.
  • Scale: CSS scale produces one image pixel per CSS pixel. Device scale follows the context’s device scale factor and can increase dimensions and file size.

Playwright can write a screenshot to a path or return bytes for uploading, processing, or custom naming. Its screenshot API also supports masking and other capture options; consult the version-matched Screenshots guide and Page API before depending on version-specific options.

Contexts and state

A browser context contains pages and their browser state. Reusing one context reduces setup overhead, but cookies, storage, and session state can carry from one URL to another. Use one context per URL when isolation matters, or deliberately seed a context with the authentication state the batch needs. Contexts also hold viewport, locale, and other emulation settings for their pages. See Playwright Pages.

4. Tune concurrency and choose a runner

Start with a modest value such as the script’s default of four, then compare a few representative batches at different limits. This is a starting configuration, not a universal recommendation. Measure elapsed time, memory, CPU, failure rate, and whether target sites begin throttling or returning different content. Raise the limit only while the actual workload benefits and the machine and sites remain stable.

Each active page uses resources, and full-page captures can be especially expensive. Very high concurrency can increase memory pressure, timeouts, and remote-site rate limits. A lower limit is appropriate when deterministic order, site politeness, or available memory matters more than parallelism. Playwright’s documentation describes Test runner workers, but does not publish a universal concurrency setting or throughput benchmark for arbitrary screenshot jobs.

Use Playwright Test if each URL is a test-like job and its reporting and retry workflow fit your needs. Set workers explicitly in the configuration or use npx playwright test --workers 4. Test workers are separate OS processes and each starts its own browser, so worker count has more overhead than opening multiple pages in one browser. Tests in a file run in order by default; fully parallel modes are available. Avoid shared output filenames or mutable globals across workers. See Parallelism.

Use a custom queue when inputs come from a file or data feed and you want direct control of naming, per-URL status, and recovery. Neither pattern is fastest for every workload. Compare setup, isolation, resource use, failure reporting, determinism, and the value of test-runner reports.

5. Make batches repeatable and recoverable

  1. Keep output mapping. Store the original URL, status, HTTP response status, output path, and error for each item. Do not build filenames from raw URLs; query strings and fragments can collide or contain unsafe characters.
  2. Retry selectively. Retry likely transient navigation failures with a small retry count and backoff. Do not repeatedly retry invalid URLs, persistent 404s, or deterministic selector failures. Record each attempt so retries do not hide the original problem.
  3. Resume from the report. Read results.json and enqueue only entries that failed, unless the output is missing or the capture settings changed.
  4. Control dynamic content. Disable animations or mask genuinely irrelevant dynamic regions for visual comparisons. Do this only if it preserves the screenshot’s purpose.
  5. Pin the environment. Keep Playwright and browser versions, operating system, headless mode, viewport, fonts, and relevant settings consistent when screenshots need to be compared. Rendering can vary across those factors. See Visual comparisons.
  6. Check the installed version. Options can change between releases. Use the documentation matching your installed Playwright version; release context is available in Release notes.

6. Troubleshoot common batch failures

Symptom Likely cause Fix
Navigation timeout The server is slow, the timeout is too short, or the chosen readiness event waits on resources that never settle. Use a bounded timeout appropriate to the workload. Try domcontentloaded or commit, then wait for a specific content signal. Record and selectively retry transient failures.
Page looks incomplete Client-side rendering, lazy loading, or delayed content continues after navigation. Wait for a meaningful selector or application state. For known cases, use a bounded delay; do not treat networkidle as a universal fix.
Blank or unexpected page The URL redirected, a bot check or error page appeared, or the application did not render. Record response.status(), inspect the final page and URL, and decide whether it should count as a captured error page or a failed job.
Some output files overwrite others Names derived only from host or path collide for URLs with different query strings. Include a stable hash of the full input URL, as the example does, and preserve the URL-to-file report.
Process runs out of memory Too many pages are active, pages remain open, or full-page images are large. Lower concurrency, close each page promptly, capture only the viewport when suitable, or use JPEG/WebP where lossy output is acceptable.
Different screenshots across runs Dynamic content, browser/OS/font differences, timing, or viewport changes affect rendering. Keep the environment and settings consistent, wait on stable page state, and hide or mask only irrelevant dynamic regions.
One failure stops the batch An unhandled exception escapes the job loop. Catch errors per URL, continue processing, and write a final report. The example does this and exits with a nonzero status if any URL failed.
Context state leaks between URLs Pages share a context with shared cookies or storage. Enable ISOLATE_CONTEXTS=1 or create contexts at the isolation boundary you need.

7. Or skip the browser setup

If you need a managed screenshot API instead of running and maintaining a browser batch, try ScreenshotNeo. Its GET endpoint captures one URL per request; the API also supports bulk capture of up to 100 URLs per call. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

The same request in Python and Node.js:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

With ScreenshotNeo, cookie banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.

8. FAQ

How many Playwright pages can I run at once?

There is no documented universal limit for arbitrary screenshot batches. It depends on page weight, image dimensions, machine resources, and remote-site behavior. Measure with a bounded queue on representative URLs.

Should I use one page, one context, or one browser per URL?

Reuse the browser process. Use pages in a shared context when shared state is acceptable; create isolated contexts when cookies or storage must not cross jobs. Launching a browser per URL usually adds avoidable startup work.

How do I screenshot a list of URLs with Playwright Test?

Represent each URL as a test or data-driven case, save to a unique path, and set the runner’s worker count to a measured limit. Use the custom queue above when you need a simple input-file-to-output-directory batch.

Does a successful screenshot mean the page returned HTTP 200?

No. A page can render and be captured after a 4xx or 5xx response. Check and store the navigation response status separately from capture success.