ScreenshotNeo

BlogHow-to

How to bulk screenshot URLs from a CSV file with Playwright

Read URLs from a CSV, capture each page with Playwright, and save organized screenshots. Includes validation, concurrency, retries, troubleshooting, and an API alternative.

By the ScreenshotNeo team4 October 202612 min read

To bulk screenshot URLs from a CSV file with Playwright, read the CSV in Node.js, validate a named URL column, open each URL in a browser page, and save each screenshot under a predictable filename. The script below runs sequentially by default, records a result for every row, retries transient failures, and always closes the browser.

Playwright’s test documentation notes that its runner operates in Node.js, so a script can read local files and use a CSV library. Its example uses csv-parse/sync with header columns and empty-line skipping. This guide uses that same parser in a standalone script. Playwright: Parameterize tests.

1. Prepare the project and CSV

Use a current Node.js installation. In a new project, install Playwright and the CSV parser, then install a browser binary:

mkdir csv-screenshots
cd csv-screenshots
npm init -y
npm install playwright csv-parse
npx playwright install chromium

Create urls.csv with a header named url. An optional name column can provide readable filename labels.

url,name
https://example.com,example-home
https://www.wikipedia.org/wiki/Playwright,playwright-wikipedia
https://developer.mozilla.org/,mdn

Use a proper CSV writer if values contain commas, quotes, or line breaks. Do not split CSV lines manually: quoted fields can contain commas.

2. Create the bulk capture script

Save this as capture.mjs. It captures the viewport by default. Set FULL_PAGE=true to capture the full scrollable page, or configure the other environment variables shown below.

import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { parse } from 'csv-parse/sync';
import { chromium } from 'playwright';

const inputFile = process.env.CSV_FILE ?? 'urls.csv';
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const urlColumn = process.env.URL_COLUMN ?? 'url';
const nameColumn = process.env.NAME_COLUMN ?? 'name';
const fullPage = process.env.FULL_PAGE === 'true';
const imageType = process.env.IMAGE_TYPE ?? 'png'; // png, jpeg, or webp
const quality = Number(process.env.QUALITY ?? 80); // used for jpeg/webp
const scale = process.env.SCALE ?? 'css'; // css or device
const waitUntil = process.env.WAIT_UNTIL ?? 'load';
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 30000);
const delayMs = Number(process.env.DELAY_MS ?? 0);
const retries = Number(process.env.RETRIES ?? 1);
const concurrency = Number(process.env.CONCURRENCY ?? 1);

if (!['png', 'jpeg', 'webp'].includes(imageType)) {
  throw new Error('IMAGE_TYPE must be png, jpeg, or webp');
}
if (!['css', 'device'].includes(scale)) {
  throw new Error('SCALE must be css or device');
}
if (!['load', 'domcontentloaded', 'networkidle', 'commit'].includes(waitUntil)) {
  throw new Error('WAIT_UNTIL must be load, domcontentloaded, networkidle, or commit');
}
for (const [label, value] of Object.entries({ timeoutMs, retries, concurrency, quality, delayMs })) {
  if (!Number.isFinite(value) || value < 0) throw new Error(`${label} must be a non-negative number`);
}
if (concurrency < 1 || !Number.isInteger(concurrency)) throw new Error('CONCURRENCY must be a positive integer');
if (!Number.isInteger(retries)) throw new Error('RETRIES must be an integer');
if (imageType !== 'png' && (quality < 0 || quality > 100)) throw new Error('QUALITY must be between 0 and 100');

const csvText = await readFile(inputFile, 'utf8');
const rows = parse(csvText, { columns: true, skip_empty_lines: true, bom: true, trim: true });
if (rows.length === 0) throw new Error(`No data rows found in ${inputFile}`);
if (!Object.hasOwn(rows[0], urlColumn)) {
  throw new Error(`CSV is missing required column "${urlColumn}"`);
}

await mkdir(outputDir, { recursive: true });
const extension = imageType === 'jpeg' ? 'jpg' : imageType;
const results = new Array(rows.length);
const browser = await chromium.launch({ headless: true });
let nextIndex = 0;

function safeLabel(value) {
  return String(value ?? '')
    .normalize('NFKD')
    .replace(/[^a-zA-Z0-9_-]+/g, '-')
    .replace(/^-+|-+$/g, '')
    .slice(0, 70);
}

function validateUrl(value) {
  const raw = String(value ?? '').trim();
  if (!raw) throw new Error('URL is blank');
  let parsed;
  try { parsed = new URL(raw); } catch { throw new Error('URL is not a valid absolute URL'); }
  if (!['http:', 'https:'].includes(parsed.protocol)) throw new Error('Only http and https URLs are supported');
  return parsed.href;
}

async function captureRow(index, context) {
  const row = rows[index];
  const rowNumber = index + 2; // header is row 1; data starts on row 2
  let url;
  let outputPath;
  try {
    url = validateUrl(row[urlColumn]);
    const label = safeLabel(row[nameColumn]);
    const base = `${String(rowNumber).padStart(4, '0')}${label ? `-${label}` : ''}`;
    outputPath = path.join(outputDir, `${base}.${extension}`);
    let lastError;

    for (let attempt = 0; attempt <= retries; attempt++) {
      const page = await context.newPage();
      try {
        page.setDefaultNavigationTimeout(timeoutMs);
        await page.goto(url, { waitUntil, timeout: timeoutMs });
        if (delayMs > 0) await page.waitForTimeout(delayMs);
        await page.screenshot({
          path: outputPath,
          type: imageType,
          fullPage,
          scale,
          ...(imageType === 'png' ? {} : { quality }),
          animations: 'disabled'
        });
        results[index] = { row: rowNumber, url, status: 'success', file: outputPath };
        console.log(`OK row ${rowNumber}: ${url} -> ${outputPath}`);
        return;
      } catch (error) {
        lastError = error;
        if (attempt < retries) console.warn(`Retry ${attempt + 1}/${retries} row ${rowNumber}: ${error.message}`);
      } finally {
        await page.close().catch(() => {});
      }
    }
    throw lastError;
  } catch (error) {
    results[index] = { row: rowNumber, url: url ?? String(row[urlColumn] ?? ''), status: 'error', error: error.message };
    console.error(`FAIL row ${rowNumber}: ${url ?? row[urlColumn] ?? '(blank)'} — ${error.message}`);
  }
}

try {
  const context = await browser.newContext({ viewport: { width: 1440, height: 1000 }, deviceScaleFactor: 1 });
  const workers = Array.from({ length: Math.min(concurrency, rows.length) }, async () => {
    while (true) {
      const index = nextIndex++;
      if (index >= rows.length) return;
      await captureRow(index, context);
    }
  });
  await Promise.all(workers);
  await context.close();
} finally {
  await browser.close();
}

await writeFile(path.join(outputDir, 'results.json'), JSON.stringify(results, null, 2) + '\n');
const failed = results.filter(result => result?.status === 'error').length;
console.log(`Finished: ${rows.length - failed} succeeded, ${failed} failed. Details: ${path.join(outputDir, 'results.json')}`);
if (failed > 0) process.exitCode = 1;

Run it with defaults:

node capture.mjs

Use a different CSV or output directory:

CSV_FILE=targets.csv OUTPUT_DIR=shots node capture.mjs

Environment variable examples:

FULL_PAGE=true WAIT_UNTIL=domcontentloaded TIMEOUT_MS=45000 RETRIES=2 node capture.mjs
IMAGE_TYPE=webp QUALITY=75 SCALE=css node capture.mjs
CONCURRENCY=3 DELAY_MS=500 node capture.mjs

3. Understand what the script does

CSV parsing and validation

The parser treats the first row as column names, skips blank lines, trims values, and accepts a UTF-8 byte order mark. The URL header is configurable with URL_COLUMN. Blank values, malformed absolute URLs, and non-HTTP(S) schemes become row errors instead of stopping the batch. Duplicate URLs are captured separately because each CSV row receives a row-based filename. If you prefer to deduplicate, do so explicitly and retain a mapping from skipped rows to the URL that was captured.

Output names and results

Files use the CSV row number, plus a sanitized optional name. The original URL is never used as the filename, since URL characters may be invalid or produce unsafe paths. Each row has a result entry in screenshots/results.json, including its source URL and either the file path or error. A failed row does not silently disappear.

Viewport, full-page, format, and scale

Choice Behavior Use it when
Viewport Captures the currently visible viewport. You need consistent above-the-fold previews.
fullPage: true Captures the full scrollable page. You need a long-page record or page review.
PNG Lossless image output; no quality setting. Visual inspection or pixel-sensitive work matters.
JPEG/WebP Compressed image output; quality is 0–100. Smaller files matter more than lossless pixels.
scale: 'css' One output pixel per CSS pixel. You want more manageable dimensions.
scale: 'device' Uses device pixels and can produce larger high-DPI images. You need device-pixel detail.

Playwright’s screenshot API documents the output path, image type, quality, scale, and full-page option. Page.screenshot API.

Wait conditions

WAIT_UNTIL maps to Playwright navigation lifecycle values. load waits for the load event; domcontentloaded can return earlier; networkidle waits for network activity to quiet and may be unsuitable for pages with persistent connections; commit returns when the response is committed. No single choice guarantees that a specific site’s client-rendered content, fonts, consent state, or lazy images are ready. For a known page, add an explicit page.waitForSelector('.content-ready'), wait for a particular response, or use a bounded delay after navigation. Prefer a real readiness signal to a large fixed delay.

Concurrency and retries

The default concurrency of one is deliberately conservative and easy to diagnose. Playwright supports multiple pages in a browser context, but the right parallelism depends on site behavior, machine memory, page complexity, and network capacity. Increase it gradually; concurrency is not a universal speed setting. Retries help with transient navigation or browser failures, but repeated retries will not fix a persistent 404, access restriction, or broken URL. Keep retries bounded.

4. Optional variants

Set a fixed viewport or emulate a device

Change the context creation to choose the viewport dimensions that suit your comparison. A stable viewport matters because responsive layouts change with width.

const context = await browser.newContext({
  viewport: { width: 1280, height: 800 },
  deviceScaleFactor: 1
});

For a device preset, use Playwright’s device descriptors from playwright and pass one to browser.newContext(). Keep the preset and browser version fixed when comparing images. Viewport, scale, and device scale factor affect the resulting pixels.

Wait for a page-specific selector

For a known target whose main content appears after navigation, wait for a stable selector before taking the screenshot:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.locator('main article').waitFor({ state: 'visible', timeout: 10000 });
await page.screenshot({ path: outputPath, fullPage });

Capture a single element

Use a locator screenshot when the output should be one component rather than the page:

await page.locator('[data-testid="report"]').screenshot({ path: outputPath });

Run URLs as Playwright Test cases

A direct script is the simplest batch utility. If each URL should be a separately reported test, Playwright Test can generate tests from CSV-backed records. The official parameterization guide demonstrates this pattern. Parameterize tests.

5. cURL, Python, and Node.js alternatives

The CSV workflow above uses Playwright in Node.js. For an API-driven batch, these examples send each CSV URL to ScreenshotNeo and save the returned image. Create an API key in ScreenshotNeo and keep it in an environment variable; do not commit secrets to source control. See the ScreenshotNeo API documentation for request options and response details.

cURL, one URL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python, batch a CSV

import csv
import os
from pathlib import Path
import requests

api_key = os.environ["SCREENSHOTNEO_API_KEY"]
out_dir = Path("screenshots")
out_dir.mkdir(parents=True, exist_ok=True)

with open("urls.csv", newline="", encoding="utf-8-sig") as f:
    for row_number, row in enumerate(csv.DictReader(f), start=2):
        url = (row.get("url") or "").strip()
        if not url:
            print(f"FAIL row {row_number}: blank URL")
            continue
        try:
            response = requests.get(
                "https://api.screenshotneo.com/v1/shot",
                params={"access_key": api_key, "url": url},
                timeout=90,
            )
            response.raise_for_status()
            (out_dir / f"{row_number:04d}.webp").write_bytes(response.content)
            print(f"OK row {row_number}: {url}")
        except requests.RequestException as exc:
            print(f"FAIL row {row_number}: {url}: {exc}")

Node.js, batch a CSV

import { readFile, mkdir, writeFile } from 'node:fs/promises';
import { parse } from 'csv-parse/sync';

const apiKey = process.env.SCREENSHOTNEO_API_KEY;
if (!apiKey) throw new Error('Set SCREENSHOTNEO_API_KEY');
const rows = parse(await readFile('urls.csv', 'utf8'), {
  columns: true, skip_empty_lines: true, bom: true, trim: true
});
await mkdir('screenshots', { recursive: true });
for (let i = 0; i < rows.length; i++) {
  const url = String(rows[i].url ?? '').trim();
  const rowNumber = i + 2;
  if (!url) { console.error(`FAIL row ${rowNumber}: blank URL`); continue; }
  const q = new URLSearchParams({ access_key: apiKey, url });
  try {
    const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`, { signal: AbortSignal.timeout(90000) });
    if (!res.ok) throw new Error(`HTTP ${res.status}`);
    await writeFile(`screenshots/${String(rowNumber).padStart(4, '0')}.webp`, Buffer.from(await res.arrayBuffer()));
    console.log(`OK row ${rowNumber}: ${url}`);
  } catch (error) {
    console.error(`FAIL row ${rowNumber}: ${url}: ${error.message}`);
  }
}

The single-request Node.js form is:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

6. Troubleshooting

Symptom Likely cause Fix
“CSV is missing required column” Header spelling differs from url, or the file has no header. Rename the first-column header or set URL_COLUMN to its exact name.
Blank or malformed URL row Empty cell, relative path, typo, or missing scheme. Provide an absolute http:// or https:// URL; inspect the row number in the log.
Navigation timeout Slow server, blocked request, or a page that never reaches the selected lifecycle event. Raise TIMEOUT_MS, try domcontentloaded, or wait for the specific content selector. Check whether the site permits automated access.
Screenshot is blank or incomplete Client rendering or lazy-loaded content was not ready when capture began. Wait for a visible content selector, a known response, or a bounded delay. For lazy content, scroll in steps before capture and verify the page's own loading behavior.
Cookie dialog or popup covers content The page displays its normal visitor consent or promotional UI. For Playwright, handle the site's consent flow where permitted, or hide a known selector with page CSS. Do not assume a navigation event removes overlays.
Access denied or CAPTCHA The site restricts automated visits or requires authentication. Respect the site's access rules. Use an authorized account or request access; do not attempt to bypass the protection.
Images differ between runs Content changes, fonts load at different times, or rendering environment differs. Use a consistent OS, browser version, viewport, scale, and readiness condition; freeze dynamic data where possible.
Large output or slow full-page capture Very tall pages, device scale, or lossless images create larger files. Use CSS scale, viewport capture, JPEG/WebP when acceptable, or capture selected elements.
“Executable doesn't exist” The Playwright browser binary is not installed in the environment. Run npx playwright install chromium in the deployment environment.

7. Performance, reliability, and cost

Local Playwright has no per-screenshot API charge, but the run consumes your machine's CPU, memory, browser storage, and network bandwidth. Full-page, high-resolution, or resource-heavy pages take more work and produce more data. Concurrency can reduce elapsed time when pages are independent, while increasing memory pressure and the chance of throttling or site restrictions. Begin with one page at a time, measure on your own URL set, then raise concurrency gradually.

For reliable batch records, preserve the input CSV, row-numbered filenames, and results JSON. Rerun only failed rows when practical instead of recapturing everything. Keep the browser version, viewport, color scheme, locale, and capture settings consistent for visual comparison. Playwright notes that snapshots can vary with operating system, browser version, settings, hardware, power source, and headless mode. Playwright visual comparisons.

Automated access can fail because of network errors, bot defenses, login walls, or site restrictions. A screenshot run does not grant permission to collect a page; follow the target site's rules and your authorization.

Or skip the browser setup

ScreenshotNeo accepts a URL in one GET request and returns a PNG, JPEG, WebP, or PDF. Its capture flow accepts consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Read the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. Sign up for 1,000 free screenshots a month with no card.

FAQ

Can I screenshot only some CSV rows?

Yes. Filter the parsed rows before assigning work, or add a command-line option that selects a row range. Keep the original CSV row number in each output name so results still map back to the source file.

Can a CSV contain headers other than url?

Yes. Set URL_COLUMN to the header name. The optional filename label column is controlled by NAME_COLUMN.

Does full-page mode automatically load every lazy image?

No readiness policy works for every site. If content loads only as it scrolls into view, scroll through the page and wait for that site's content before calling the screenshot API.

Can I use this on authenticated pages?

Only when you have permission. Configure an authorized browser context with the site's supported login flow or session state, and keep credentials and session files out of source control.