ScreenshotNeo

BlogHow-to

How to Capture Website Screenshots in Bulk from a CSV with Puppeteer

Build a reliable Node.js job that reads URLs from CSV, captures screenshots with Puppeteer, and records failures without stopping the batch.

By the ScreenshotNeo team4 October 20269 min read

Use a CSV parser to read and validate each URL, launch one Puppeteer browser for the job, and process rows with a page lifecycle that records each result independently. For every valid row, navigate using a readiness signal suited to that site, save the screenshot to a unique path, and continue after errors. The example below processes rows sequentially; add bounded concurrency only after measuring resource use on representative pages.

Puppeteer’s documented capture sequence is to launch a browser, create a page, navigate, call Page.screenshot(), and close the browser. The Puppeteer guide demonstrates waitUntil: 'networkidle2', but that is an example rather than a universal readiness setting. See the Puppeteer screenshots guide, Page.screenshot() API, and ScreenshotOptions reference.

1. Install Node.js, Puppeteer, and a CSV parser

This example uses csv-parse as a separate CSV parser. Puppeteer itself does not provide a CSV batch runner or prescribe a parser.

mkdir csv-screenshots
cd csv-screenshots
npm init -y
npm install puppeteer csv-parse

The puppeteer package downloads a compatible Chrome during installation. puppeteer-core does not download a browser and is intended for cases where you manage browser installation and configuration yourself. Some package managers block dependency install scripts; if that prevents the browser download, Puppeteer documents npx puppeteer browsers install as a manual installation path. See the Puppeteer installation and overview documentation.

Create urls.csv with a header named url:

url
https://example.com/
https://pptr.dev/
https://www.wikipedia.org/

2. Create a resilient batch script

Save this as capture.js. It validates URLs, generates filesystem-safe unique names, records one result per CSV row, catches navigation and screenshot errors per row, and closes pages and the browser in cleanup blocks.

const fs = require('node:fs/promises');
const path = require('node:path');
const { parse } = require('csv-parse/sync');
const puppeteer = require('puppeteer');

const inputPath = process.argv[2] || 'urls.csv';
const outputDir = process.argv[3] || 'screenshots';
const resultsPath = path.join(outputDir, 'results.jsonl');
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS || 30_000);
const waitUntil = process.env.WAIT_UNTIL || 'networkidle2';
const fullPage = process.env.FULL_PAGE !== 'false';

function validateUrl(value) {
  const raw = String(value || '').trim();
  if (!raw) throw new Error('Empty URL');
  let parsed;
  try {
    parsed = new URL(raw);
  } catch {
    throw new Error('Invalid URL');
  }
  if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
    throw new Error(`Unsupported URL scheme: ${parsed.protocol}`);
  }
  return parsed.href;
}

function outputName(rowNumber, url) {
  const parsed = new URL(url);
  const host = parsed.hostname.replace(/[^a-zA-Z0-9.-]/g, '_');
  const slug = (parsed.pathname + parsed.search)
    .replace(/[^a-zA-Z0-9.-]+/g, '_')
    .replace(/^_+|_+$/g, '')
    .slice(0, 80) || 'page';
  return `${String(rowNumber).padStart(5, '0')}-${host}-${slug}.png`;
}

async function appendResult(record) {
  await fs.appendFile(resultsPath, `${JSON.stringify(record)}\n`, 'utf8');
}

async function main() {
  await fs.mkdir(outputDir, { recursive: true });
  // Start a fresh log for each run so old and new outcomes are not mixed.
  await fs.writeFile(resultsPath, '', 'utf8');

  const csvText = await fs.readFile(inputPath, 'utf8');
  const rows = parse(csvText, {
    columns: true,
    skip_empty_lines: true,
    bom: true,
    relax_quotes: false,
    trim: true
  });
  if (!rows.length) throw new Error(`No data rows in ${inputPath}`);
  if (!Object.hasOwn(rows[0], 'url')) {
    throw new Error('CSV must have a column named url');
  }

  const browser = await puppeteer.launch();
  try {
    for (let index = 0; index < rows.length; index++) {
      const row = rows[index];
      const csvRow = index + 2; // Header is CSV row 1.
      let url;
      let filePath;
      let page;
      try {
        url = validateUrl(row.url);
        filePath = path.join(outputDir, outputName(csvRow, url));
        page = await browser.newPage();
        await page.goto(url, { waitUntil, timeout: timeoutMs });
        await page.screenshot({ path: filePath, fullPage, type: 'png' });
        await appendResult({ csvRow, url, status: 'success', output: filePath });
      } catch (error) {
        await appendResult({
          csvRow,
          url: url || String(row.url || ''),
          status: 'error',
          output: filePath || null,
          error: error instanceof Error ? error.message : String(error)
        });
      } finally {
        if (page) await page.close().catch(() => {});
      }
    }
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it with:

node capture.js urls.csv screenshots

The default timeout, wait condition, and full-page choice in this script are configurable examples, not Puppeteer requirements or guarantees. The log at screenshots/results.jsonl contains one JSON object per data row. Invalid URLs also receive an error record, so they do not stop later rows.

3. Choose when a page is ready

The browser can finish navigation before a site’s useful content is ready, and some pages keep network requests open. Select a condition based on the pages you capture:

Strategy Use when Trade-off
domcontentloaded The initial document is enough, or you will wait for a specific element next. Images and client-rendered content may not yet be ready.
load You want the page load event, including load-dependent resources. It does not prove that delayed or application-rendered content is complete.
networkidle2 The page settles to a low level of network activity and that matches the site. Long-lived requests can delay or prevent the condition.
Selector wait A known element indicates the content you need has rendered. You must choose a selector that exists and becomes visible on each target page.

For a selector-based page, navigate first and then wait explicitly before capturing:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.waitForSelector('main article', { visible: true, timeout: timeoutMs });
await page.screenshot({ path: filePath, fullPage: true });

For a fixed delay, use it only when the target site has a known short rendering delay; it can waste time on fast pages and still be too short on slow ones:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await new Promise(resolve => setTimeout(resolve, 1500));
await page.screenshot({ path: filePath, fullPage: true });

4. Configure screenshot output

page.screenshot() supports several useful options. The API reference is authoritative for the installed Puppeteer version; check it when upgrading.

Option Effect Notes
path Writes the image to a file. Use a deterministic, unique path for each row to prevent overwrites.
fullPage Captures the full page instead of only the viewport. Defaults to false. Full-page images can consume more memory and disk space.
type Selects PNG, JPEG, or WebP where supported by the installed browser. PNG is the documented default. Confirm support for your runtime when choosing another type.
quality Sets lossy image quality for supported formats. Does not apply to PNG.
clip Captures a specified rectangular region. Use coordinates and dimensions from the page layout; it is not a substitute for selecting an element.
omitBackground Omits the default background where supported, allowing transparency. Useful for assets that need transparency; verify the target format supports the result.

To capture one element rather than the whole page, use Puppeteer’s documented ElementHandle.screenshot() path:

const element = await page.waitForSelector('.product-card', { visible: true });
if (!element) throw new Error('Product card was not found');
await element.screenshot({ path: filePath, type: 'png' });

To use JPEG and a quality setting, choose a .jpg path and set the type explicitly:

await page.screenshot({ path: filePath.replace(/\.png$/, '.jpg'), type: 'jpeg', quality: 80, fullPage: false });

5. Make output names and reruns predictable

Never derive a path from an unfiltered URL alone: URLs can contain characters that are invalid or awkward in filenames, and multiple rows may point to the same URL. The example prefixes the sanitized host and path with the CSV row number, which keeps duplicate URLs distinct and makes it easy to trace a file back to its source row. If URLs contain sensitive query parameters, omit the query from filenames and keep the original URL only in a protected results log.

For reproducible runs, keep the input CSV unchanged, use a separate output directory per run, and retain the JSON Lines result log. If a process stops unexpectedly, the log shows which rows completed; decide whether to skip existing successful files or recapture them as part of an explicit rerun policy.

6. Process at a measured pace

Sequential processing uses one page at a time and is the simplest starting point for debugging. It limits simultaneous page memory use but may take longer for large lists. Bounded concurrency can increase throughput, but each open page consumes browser resources and target sites may respond differently to parallel requests. The reviewed Puppeteer documentation provides no safe concurrency number or throughput benchmark, so measure representative pages and increase the limit gradually.

Reuse one browser for the whole batch and close each page after its row, as in the example. Use timeouts so a stalled navigation does not hold the job indefinitely. For especially large pages, full-page capture, image decoding, and saving files add memory and disk load; monitor the process and available storage. Keep failures in the log and rerun only failed rows when appropriate rather than silently discarding them.

7. Troubleshooting

Symptom Likely cause Fix
“Could not find Chrome” or browser executable missing The install script was blocked, or puppeteer-core was installed without a browser. With Puppeteer, run npx puppeteer browsers install. With puppeteer-core, install and configure a compatible browser yourself.
Navigation timeout on only some URLs The page is slow, has long-lived requests, or never reaches the selected wait condition. Use an appropriate per-site readiness signal, such as domcontentloaded followed by a selector wait; adjust the configurable timeout based on your workload.
Screenshot is blank or missing app content The screenshot ran before client-side rendering completed, or the expected content did not load. Wait for a visible content selector and record failures when it does not appear. Inspect the URL and page behavior manually.
CSV rows shift or URLs split at commas The file was parsed by splitting on commas instead of using a CSV parser, or the file has malformed quoting. Use a real CSV parser, confirm the header name, and correct malformed CSV quoting.
One bad URL stops the run Validation or navigation errors are not caught inside the per-row loop. Keep validation, navigation, and capture inside a row-level try/catch and append a result for every row.
Images overwrite each other Output names collide for repeated or similar URLs. Include a stable row identifier or unique ID in every filename.
Process runs out of memory or disk fills Too many pages are open, full-page captures are large, or output accumulates. Process sequentially or reduce bounded concurrency, use viewport capture where sufficient, and monitor or clean output storage.

8. Or skip the browser setup

ScreenshotNeo takes a screenshot with one GET request. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

For a CSV batch, keep the same validation and per-row logging pattern and call the endpoint once for each valid URL, saving each response to its own file. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. See ScreenshotNeo and sign up for 1,000 free screenshots a month with no card.

FAQ

How do I take full-page screenshots?

Set fullPage: true in page.screenshot(). Its documented default is false.

Why is Puppeteer missing Chrome after install?

The package install script may have been blocked, or you may have installed puppeteer-core, which does not download Chrome. Install the browser with Puppeteer’s documented browser installer or manage a compatible executable yourself.

How do I keep one failed URL from stopping the batch?

Catch errors inside the loop for each row, record the failed URL and error, close that row’s page in finally, and continue the loop.

Does this batch script guarantee every page is captured exactly as a visitor sees it?

No. Readiness, access controls, site behavior, and rendering vary. Validate the result files for your target sites and choose a readiness signal that matches the content you need.