ScreenshotNeo

BlogHow-to

How to Take Bulk Screenshots with Playwright in Node.js

Capture a list of URLs reliably with Playwright in Node.js, with full-page options, stable filenames, bounded concurrency, retries, and CI tips.

By the ScreenshotNeo team29 September 202610 min read

How to Take Bulk Screenshots with Playwright in Node.js

Direct answer: use one Playwright browser, create a reusable context and page, loop over URL records, wait for an intentional readiness condition, and call page.screenshot() for each URL. Give every record a deterministic, filesystem-safe filename. Use fullPage: true when the complete scrollable document is required; omit it for the current viewport.

This pattern is suitable for documentation snapshots, visual regression inputs, sitemap previews, and scheduled archives. The complete example below adds per-URL error handling, retries, stable output paths, screenshot-only CSS, masking, and a bounded worker pool for larger batches.

1. Install Playwright and prepare a batch

Create a Node.js project and install Playwright. The browser binaries are installed separately so deployments and CI environments have the same engine available.

mkdir bulk-shots
cd bulk-shots
npm init -y
npm install playwright
npx playwright install chromium

Represent each target as an object. A separate slug keeps output names predictable and avoids deriving filenames from query strings or unsafe URL characters.

const targets = [
  { url: 'https://example.com', slug: 'example' },
  { url: 'https://playwright.dev', slug: 'playwright' },
  { url: 'https://developer.mozilla.org/en-US/docs/Web/JavaScript', slug: 'mdn-javascript' },
];

For inputs supplied by users or a CSV file, sanitize and deduplicate slugs before starting the browser. A collision can silently overwrite an earlier image.

2. A complete sequential script

This runnable script launches Chromium once, reuses one context and page, creates the output directory, waits for each page, captures a full-page PNG, and continues after an individual failure. The finally block closes resources even when the batch is interrupted.

A reusable browser and deterministic filenames keep bulk captures manageable.
A reusable browser and deterministic filenames keep bulk captures manageable.
import { chromium } from 'playwright';
import fs from 'node:fs/promises';
import path from 'node:path';

const targets = [
  { url: 'https://example.com', slug: 'example' },
  { url: 'https://playwright.dev', slug: 'playwright' },
];

const outputDir = path.resolve('screenshots');
const viewport = { width: 1440, height: 900 };
const navigationTimeout = 45_000;
const retries = 2;

function safeSlug(value) {
  const cleaned = value
    .normalize('NFKC')
    .replace(/[^a-zA-Z0-9._-]+/g, '-')
    .replace(/^-+|-+$/g, '')
    .slice(0, 120);
  return cleaned || 'page';
}

function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}

async function captureOne(page, target) {
  const slug = safeSlug(target.slug);
  const outputPath = path.join(outputDir, `${slug}.png`);

  for (let attempt = 1; attempt <= retries + 1; attempt++) {
    try {
      await page.goto(target.url, {
        waitUntil: 'domcontentloaded',
        timeout: navigationTimeout,
      });

      // Wait for the application state that matters to your page.
      await page.waitForLoadState('networkidle', { timeout: 15_000 }).catch(() => {});

      await page.screenshot({
        path: outputPath,
        fullPage: true,
        type: 'png',
        scale: 'css',
        style: `
          *, *::before, *::after {
            animation: none !important;
            transition: none !important;
            caret-color: transparent !important;
          }
        `,
      });

      return { ok: true, url: target.url, outputPath };
    } catch (error) {
      if (attempt > retries) {
        return {
          ok: false,
          url: target.url,
          error: error instanceof Error ? error.message : String(error),
        };
      }
      await sleep(1_000 * attempt);
    }
  }
}

const browser = await chromium.launch();
const context = await browser.newContext({ viewport });
const page = await context.newPage();
page.setDefaultTimeout(navigationTimeout);

try {
  await fs.mkdir(outputDir, { recursive: true });
  const results = [];

  for (const target of targets) {
    const result = await captureOne(page, target);
    results.push(result);
    console.log(JSON.stringify(result));
  }

  const failures = results.filter(result => !result.ok);
  if (failures.length) {
    process.exitCode = 1;
  }
} finally {
  await page.close().catch(() => {});
  await context.close().catch(() => {});
  await browser.close();
}

Playwright’s page.screenshot() can write directly to a path or return image bytes. fullPage: true captures the full scrollable document; the default captures the current viewport. The official guide describes a full-page image as one in which a very tall screen could contain the entire scrollable page.

3. Screenshot options you will use in bulk jobs

Option Use Notes
fullPage Capture the complete scrollable page. Can create very tall files and expose lazy-loading issues.
type png, jpeg, or webp. PNG is lossless; JPEG requires a quality value; WebP often reduces size.
quality JPEG/WebP quality tuning. Use a value from 0 to 100 where supported; it has no effect on PNG.
path Write the image to disk. Ensure the parent directory exists.
scale css or device pixel output. CSS scale produces predictable dimensions; device scale is higher resolution.
clip Capture a rectangle with x, y, width, and height. Useful for fixed cards or charts.
mask Cover sensitive or changing locators. Pass an array of locators, such as [page.locator('.avatar')].
style Inject CSS only for the screenshot. Disable motion, hide cursors, or remove a volatile widget.
timeout Set the screenshot operation limit. Keep navigation and screenshot timeouts explicit in CI.

PNG, JPEG, or WebP?

Choose PNG for pixel comparison, text-heavy pages, and archival fidelity. Choose JPEG when small files matter and slight artifacts are acceptable. WebP is useful when your downstream system supports it and you want a compact image. Set the extension and type together so consumers do not misinterpret the bytes.

Viewport versus full page

Viewport captures have stable dimensions and are faster to review. Full-page captures include content below the fold, but pages that load content on scroll may need a scroll routine before the screenshot. A simple helper scrolls incrementally and then returns to the top:

async function loadLazyContent(page) {
  await page.evaluate(async () => {
    await new Promise(resolve => {
      let y = 0;
      const step = Math.max(300, window.innerHeight * 0.8);
      const timer = setInterval(() => {
        y += step;
        window.scrollTo(0, y);
        if (y + window.innerHeight >= document.documentElement.scrollHeight) {
          clearInterval(timer);
          window.scrollTo(0, 0);
          resolve();
        }
      }, 100);
    });
  });
}

await loadLazyContent(page);
await page.screenshot({ path: 'screenshots/lazy.png', fullPage: true });

4. Waiting for the right state

networkidle is not a universal definition of “ready.” Analytics, polling, advertisements, and WebSockets can keep connections open indefinitely. Prefer a selector or application signal that means the content you need is rendered.

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('[data-testid="report-ready"]').waitFor({ state: 'visible', timeout: 30_000 });
await page.screenshot({ path, fullPage: true });

For a known fixed delay, use page.waitForTimeout() sparingly. A selector is usually faster and more reliable. If a consent dialog blocks the page, handle it explicitly in your own script or use a screenshot service that can remove common banners before capture.

5. Bounded concurrency for larger URL lists

One page in a loop is simple and gentle on the target host. For hundreds or thousands of URLs, a bounded worker pool can reduce total time while limiting memory, CPU, and request pressure. The official Playwright documentation does not publish a universal concurrency number, so measure your workload and increase workers gradually.

async function worker(browser, jobs, workerId) {
  const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
  const page = await context.newPage();
  try {
    while (true) {
      const job = jobs.shift();
      if (!job) break;
      const file = path.join(outputDir, `${safeSlug(job.slug)}.webp`);
      try {
        await page.goto(job.url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
        await page.screenshot({
          path: file,
          fullPage: true,
          type: 'webp',
          quality: 85,
          scale: 'css',
        });
        console.log(`worker ${workerId}: ${job.url}`);
      } catch (error) {
        console.error(`worker ${workerId}: ${job.url}:`, error);
      }
    }
  } finally {
    await page.close().catch(() => {});
    await context.close().catch(() => {});
  }
}

const browser = await chromium.launch();
const jobs = targets.map(target => ({ ...target }));
const workerCount = 3;
await Promise.all(
  Array.from({ length: workerCount }, (_, i) => worker(browser, jobs, i + 1)),
);
await browser.close();

Use separate contexts for isolation. Do not create an unbounded page or browser for every URL. Keep the worker count below what the machine and target servers can sustain, and honor the target site’s access rules.

6. Deterministic output and CI stability

  • Use a fixed browser engine, viewport, locale, timezone, and color scheme.
  • Sanitize slugs and include an ID or hash when two URLs can share a slug.
  • Disable animations and transitions with style.
  • Mask timestamps, avatars, rotating ads, and user-specific regions.
  • Use a stable data fixture instead of live, changing content for visual regression.
  • Store a manifest containing URL, timestamp, browser version, options, and output path.
  • Write to a temporary filename and rename after success if another process reads the directory.
const manifest = {
  url: target.url,
  path: outputPath,
  capturedAt: new Date().toISOString(),
  viewport,
  fullPage: true,
  type: 'png',
};
await fs.writeFile(`${outputPath}.json`, JSON.stringify(manifest, null, 2));
Scroll for lazy content, then stabilize motion and volatile regions before capture.
Scroll for lazy content, then stabilize motion and volatile regions before capture.

7. The official Playwright CLI for one-off captures

For a single URL or a small shell script, Playwright’s screenshot CLI supports full-page capture, an output filename, image type, and high-resolution mode. Check the installed version’s help because command syntax can vary:

npx playwright screenshot --help
npx playwright screenshot --full-page --filename=shot.png https://example.com

Use the Node.js API for batches because it gives you structured retries, per-URL logs, custom readiness checks, and controlled concurrency.

8. Troubleshooting common failures

Symptom Likely cause Fix
Browser executable not found Playwright package is installed but browsers are not. Run npx playwright install chromium during image or CI setup.
Timeout at goto Slow origin, redirect loop, blocked request, or an overly strict timeout. Log the URL, inspect redirects, raise the timeout selectively, and retry transient failures.
Blank or partial page Capture started before client rendering or required data arrived. Wait for a meaningful selector or application-ready marker.
Lazy images are missing Images load only after scrolling into view. Scroll through the document before calling screenshot.
Different pixels on every run Animations, clocks, ads, random IDs, or personalized data. Inject screenshot CSS, mask volatile locators, freeze fixtures, and set a stable context.
Cookie banner covers content A consent component remains interactive and visible. Accept it in the test flow, hide its selector, or use a capture service that handles consent automatically.
Files overwrite each other Two targets produce the same slug. Deduplicate names and append a stable ID or short hash.
Process runs out of memory Too many workers, very tall pages, or leaked contexts. Lower concurrency, reuse contexts, close resources in finally, and split large batches.
Works locally but fails in CI Missing browser dependencies, fonts, environment variables, or different viewport. Install browser dependencies, pin versions, set fonts and viewport explicitly, and save failure artifacts.

9. Performance, reliability, and cost considerations

Capture time is dominated by page load, JavaScript execution, image decoding, and full-page layout. Measure median and tail duration for your actual URLs; no official Playwright source publishes a universal throughput benchmark. Reusing one browser avoids repeated startup cost. Bounded workers improve throughput until CPU, memory, bandwidth, or the target host becomes the bottleneck.

For reliability, classify errors as navigation, readiness, screenshot, or filesystem failures. Retry transient navigation failures with backoff, but do not endlessly retry deterministic 404s or authentication failures. Persist a result record for every URL so a later run can retry only failures. Use a queue or checkpoint file when batches may be interrupted.

Playwright itself is open source, but your operational cost includes compute, browser storage, bandwidth, and maintenance of consent, authentication, and rendering logic. Full-page images consume more disk and transfer than viewport images. WebP or JPEG can reduce storage when lossless pixels are not required.

10. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when you want one HTTP request instead of maintaining browser workers. Its clean-shot flow accepts common cookie and consent banners, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets each step be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.

See the ScreenshotNeo documentation for all options. A basic call is:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await fs.writeFile('shot.webp', bytes);

For bulk work, ScreenshotNeo supports up to 100 URLs per call, async jobs with signed webhooks, caching with a chosen TTL, and a usage API. It also supports full-page capture with lazy images loaded, CSS element capture, dark mode, custom viewports and device presets, retina scale, PDF output, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, and signed links. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account to try the API.

11. Short FAQ

Should I use one page or one page per URL?

One reused page is simplest for sequential jobs. Use separate pages or contexts in a bounded worker pool when URLs are independent and the machine has capacity.

Can screenshots be returned as buffers?

Yes. Omit path and await the return value from page.screenshot(); the result is image bytes suitable for object storage or further processing.

How do I capture only an element?

Locate it and use its bounding box for clip, or capture the element through the locator screenshot API when your installed Playwright version supports it.

Why does full-page capture differ from scrolling manually?

Full-page capture lays out the document as one tall image. Manual scrolling can trigger lazy loading and intersection observers first, so scroll when those page behaviors are required.

How many concurrent workers should I run?

There is no universal official number. Start with a small pool, measure completion time and resource use, and increase it until CPU, memory, bandwidth, or the target host becomes the limiting factor.