ScreenshotNeo

BlogHow-to

How to Take Bulk Screenshots of Web Pages and Create a PDF Report

Capture a list of web pages with consistent settings, save named screenshots or PDFs, and assemble and check a single report.

By the ScreenshotNeo team4 October 20269 min read

To take bulk screenshots of web pages and create a PDF report, start with a URL list, automate one browser capture per URL, give each output a stable numbered filename, and then assemble the page PDFs into one document if you need a single report. Use screenshots when you need the page’s visible appearance; use browser PDF export when you want printable pages. Those are different outputs: PDFs usually follow print CSS and page breaks, while screenshots capture pixels.

For repeatable work, Playwright is a practical choice because one script can visit every URL, wait for page-specific content, save full-page images or PDFs, and record failures. Chrome Headless is useful for simple one-page command-line captures. Tools such as webshot2 and shot-scraper provide multi-URL workflows. A batch of individual screenshots or PDFs does not automatically mean you have one compiled report.

1. Choose the output and capture plan

Before automating, settle these choices:

  • One image per page: best for visual review, design evidence, or archiving the rendered viewport. Full-page screenshots include content below the fold, but can become very tall and difficult to read in a report.
  • One PDF per page: best when readers need selectable text, print layout, or paper-sized pages. Browser PDF export applies print styles by default in Playwright, so it may not match the screen.
  • One combined PDF: capture one PDF per URL, then merge those PDFs with a PDF tool approved in your environment. Keep merging as an explicit step; browser capture tools do not universally create a single report from a URL list.

Use a fixed viewport, browser version, color scheme, and capture mode when the results will be compared over time. Decide how to handle redirects, authentication, locale, and dynamic content. Use ordered names such as 001-example-home.pdf so filesystem ordering matches report ordering.

2. Run a bulk capture with Playwright and Node.js

This complete example reads URLs from urls.txt, writes one PDF per URL to output/, limits concurrency, waits for an optional selector, retries transient navigation failures, and writes a JSONL result record for every URL. It uses Node.js 20 or newer and Playwright.

npm init -y
npm install playwright
npx playwright install chromium

Create urls.txt with one URL per line, then save this as capture.mjs:

import { chromium } from 'playwright';
import { mkdir, readFile, appendFile } from 'node:fs/promises';

const input = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line && !line.startsWith('#'));

const outputDir = 'output';
const concurrency = 3;
const maxAttempts = 2;
const readySelector = process.env.READY_SELECTOR;
const fullPage = process.env.FULL_PAGE === '1';
const media = process.env.PDF_MEDIA === 'screen' ? 'screen' : 'print';

await mkdir(outputDir, { recursive: true });
await import('node:fs/promises').then(({ rm }) => rm('results.jsonl', { force: true }));

const browser = await chromium.launch({ headless: true });
let next = 0;

async function captureWorker() {
  while (next < input.length) {
    const index = next++;
    const url = input[index];
    const number = String(index + 1).padStart(3, '0');
    let lastError;

    for (let attempt = 1; attempt <= maxAttempts; attempt++) {
      const page = await browser.newPage({
        viewport: { width: 1440, height: 1000 },
        deviceScaleFactor: 1,
      });
      try {
        const response = await page.goto(url, {
          waitUntil: 'domcontentloaded',
          timeout: 45000,
        });
        if (response && response.status() >= 400) {
          throw new Error(`HTTP ${response.status()}`);
        }
        if (readySelector) {
          await page.locator(readySelector).waitFor({
            state: 'visible',
            timeout: 15000,
          });
        } else {
          await page.waitForTimeout(1000);
        }

        await page.emulateMedia({ media });
        await page.pdf({
          path: `${outputDir}/${number}.pdf`,
          format: 'A4',
          printBackground: true,
          preferCSSPageSize: true,
          displayHeaderFooter: false,
          margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
        });

        await appendFile('results.jsonl', JSON.stringify({
          index: index + 1, url, output: `${outputDir}/${number}.pdf`,
          status: 'ok', httpStatus: response?.status() ?? null,
        }) + '\n');
        lastError = null;
        break;
      } catch (error) {
        lastError = error;
        if (attempt < maxAttempts) await new Promise((r) => setTimeout(r, attempt * 1000));
      } finally {
        await page.close();
      }
    }

    if (lastError) {
      await appendFile('results.jsonl', JSON.stringify({
        index: index + 1, url, status: 'error', error: String(lastError),
      }) + '\n');
    }
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, input.length) }, captureWorker));
} finally {
  await browser.close();
}

Run it with node capture.mjs. To wait for a page-specific element instead of using the one-second baseline, set READY_SELECTOR, for example READY_SELECTOR='.report-content' node capture.mjs. The selector should represent content that indicates the page is ready for your report. A fixed delay is only a practical baseline: it does not prove that every asynchronous request, image, or lazy-loaded section has finished.

Change the script for full-page images

For PNG captures, replace the page.pdf(...) call with:

await page.screenshot({
  path: `${outputDir}/${number}.png`,
  fullPage: process.env.FULL_PAGE === '1',
  animations: 'disabled',
  caret: 'hide',
});

Set FULL_PAGE=1 for the full scrollable page. Otherwise, Playwright captures the viewport. Its screenshot API also supports clipping, masking, image quality, and CSS-pixel versus device-pixel scale. Full-page output can be very large; a viewport or an element capture may be more useful for unusually long pages.

Useful Playwright settings

  • waitUntil: domcontentloaded waits for the initial HTML to be parsed. load waits for the load event; neither guarantees that a site-specific client render is complete. Prefer a meaningful selector for dynamic pages.
  • page.pdf(): accepts paper format or dimensions, margins, landscape, page ranges, background printing, and header/footer templates. Playwright uses print CSS by default. Call page.emulateMedia({ media: 'screen' }) first when screen styles are desired.
  • preferCSSPageSize: lets CSS @page size rules determine the page size. Turn it off if the chosen PDF format should take precedence.
  • printBackground: includes background graphics, which may increase output size. For exact color rendering, the page’s print CSS can use -webkit-print-color-adjust: exact.
  • viewport and deviceScaleFactor: control screenshot dimensions and pixel density. Use consistent values to make captures comparable.
  • Browser context: configure locale, timezone, color scheme, user agent, permissions, or authenticated storage state if the pages require them. Keep credentials out of source control and result logs.

3. Capture with Chrome Headless

Chrome’s command-line flags capture one target page per invocation. For a URL list, wrap the command in a loop. The flags do not themselves turn a list into a batch.

mkdir -p output
n=0
while IFS= read -r url; do
  [ -z "$url" ] && continue
  n=$((n + 1))
  name=$(printf '%03d' "$n")
  chrome --headless --disable-gpu --no-pdf-header-footer \
    --timeout=15000 --print-to-pdf="output/${name}.pdf" "$url"
done < urls.txt

Use the executable name installed on your system; it may be google-chrome or chromium. For an image instead of a PDF, use --screenshot="output/001.png" --window-size=1440,1000. Chrome’s documented --timeout bounds the wait, and --virtual-time-budget can advance time-dependent page code before printing. These are timing controls, not proof that every page-specific render is finished. For one report PDF, merge the resulting PDFs in a separate step.

See the Chrome Headless command-line reference for supported flags and current behavior.

4. Other bulk capture options

  • webshot2: accepts multiple URLs, can capture them in parallel, and supports output naming and viewport controls. Its PDF output uses browser printing. The manual notes that selector, clipping, expansion, and zoom options do not apply to PDF output. See the webshot2 reference.
  • shot-scraper: has a multi workflow that reads URL and output entries from configuration. It is useful when a declarative list is more convenient than writing a browser loop. See the shot-scraper documentation.
  • Puppeteer: a JavaScript browser automation alternative for screenshot, PDF, navigation, and interface automation. Check the current Puppeteer documentation for version-specific APIs.

Compare tools by whether they accept a URL list, support parallel capture, produce images or browser-printed PDFs, allow page-specific waits, control filenames and ordering, and include or omit single-file report assembly.

5. Assemble and review the report

  1. Read results.jsonl and resolve failed URLs before treating the batch as complete.
  2. Check that output filenames have a one-to-one mapping to input URLs. Keep the original URL list and result log with the report for traceability.
  3. Merge individual PDFs with a PDF utility available in your environment, preserving the numbered order. Because merger tools differ, follow that tool’s documentation and verify its output.
  4. Open the final PDF and inspect ordering, missing or blank pages, clipping, page breaks, readability, headers, and whether dynamic content appeared.
  5. For screenshot reports, consider whether tall full-page images remain legible when placed on paper-sized pages. A viewport capture or one image per report page may be easier to review.

6. Performance, reliability, and cost

Each URL requires browser navigation and rendering, so total duration depends on the sites, waits, and concurrency. More parallel workers can reduce elapsed time, but consume more memory and CPU and may increase load on target sites. Start with a small concurrency value and increase it only if the machine and sites handle it reliably. Do not send a burst of requests to sites unless you are authorized to capture them.

Use bounded timeouts, per-page readiness checks, limited retries, and a result log. Retry transient navigation or network failures, but do not retry indefinitely or treat an HTTP error page as a successful capture. A wait for network idle can hang on pages with continuous analytics or streaming; a selector tied to the content you need is often a clearer readiness condition. Check representative outputs before running a large batch.

Local browser capture has no per-shot service charge, but uses your compute, storage, maintenance time, and possibly paid infrastructure if run in hosted environments. PDFs and full-page, high-density images can consume substantial disk space. Choose image compression and pixel density based on the report’s actual readability needs.

7. Troubleshooting

Symptom Likely cause Fix
PDF is blank or missing page content Capture occurred before client-side rendering, or the page blocks headless access. Wait for a content-specific selector, inspect the URL manually, and record HTTP status and errors. A longer fixed delay alone may not solve blocked access.
Images or lower sections are missing Lazy loading has not triggered, or the capture only covers the viewport. Use full-page capture for screenshots and test whether scrolling or a page-specific wait is needed to load content. Inspect the output.
PDF looks different from the browser PDF export uses print CSS by default and may change colors, layout, or page breaks. Use screen media before PDF generation if screen styling is wanted; review print CSS, @page, margins, and background printing.
PDF is clipped or has unexpected page breaks Print styles, fixed dimensions, or margins conflict with paper size. Adjust CSS page breaks and margins, select the intended paper size, and inspect representative pages before the full batch.
Capture times out Slow navigation, a stalled request, or an overly short timeout. Set a bounded timeout appropriate to the site, wait for a narrower readiness signal, and retry a limited number of times.
Browser cannot launch Playwright’s browser binary is not installed or system dependencies are unavailable. Run npx playwright install chromium and consult Playwright’s installation guidance for the operating system.
Some outputs overwrite others Names are based on the URL or a reused filename. Use a stable index in every filename and preserve the input-to-output mapping.
Batch is slow or machine runs out of memory Too many pages run at once, or pages/images are unusually large. Lower concurrency, close each page after capture, use viewport captures where appropriate, and avoid unnecessary high-density output.
One PDF is expected but many files exist Capture produced one PDF per URL; compilation was not included. Run a distinct merge step and verify page order and completeness.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. For one page, use this request; for a bulk list, make one request per URL and save each response under its numbered filename. See the ScreenshotNeo API documentation for request parameters and PDF options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, no card required.

FAQ

Can I make one PDF from a URL list in Chrome?

Chrome Headless exports a page per invocation. Capture the pages, then use a separate PDF assembly step and verify the resulting order.

Should I use screenshots or PDFs for a report?

Use screenshots to preserve visual pixels. Use PDFs for print-oriented pages and text that readers may select or search. Inspect PDF print styling before distributing it.

Does a fixed delay guarantee a page is ready?

No. It only waits for a set duration. A selector or other site-specific readiness condition gives a more relevant signal, but the resulting artifact still needs review.

Can I capture authenticated pages?

Yes, when your browser workflow is configured with the required session or credentials and you are authorized to access the pages. Keep secrets out of scripts, logs, and shared report files.