How to Capture Website Screenshots in Bulk from a CSV of URLs for a Report
Use Playwright to capture URLs from a CSV, save consistently named screenshots, record failures, and assemble the results into an auditable report.
To capture website screenshots in bulk from a CSV, use a browser automation script to read each row, visit its URL, save a screenshot with a stable filename, and record success or failure in a manifest. This guide uses Playwright with Node.js. It supports viewport or full-page captures, configurable output format and size, controlled concurrency, and a report-ready CSV manifest.
Playwright provides the browser primitives: navigate to a page, take a screenshot, and optionally capture the full page or a particular element. The CSV loop, naming scheme, retry policy, and manifest below are workflow choices built around those primitives. See the official Playwright screenshot documentation and Page API.
1. Prepare the input CSV
Save a file named urls.csv with a url header and one URL per row. An optional id column gives each image a readable, stable name:
id,url
001,https://example.com/
002,https://www.iana.org/domains/reserved
Use complete URLs, including https://. Keep IDs unique. The script below rejects missing or invalid URLs and creates a safe filename from the ID and hostname.
2. Install Playwright and CSV parsing
From an empty project directory, run:
npm init -y
npm install playwright csv-parse
npx playwright install chromium
Save the script in the next section as capture.mjs. Run it with:
node capture.mjs urls.csv screenshots
The first argument is the input CSV; the second is the output directory. The script creates that directory if needed. Playwright’s browser installation is separate from installing the package.
3. Capture each URL and write a manifest
This runnable script uses a small worker pool so a large input does not open every page at once. It writes PNG files and a manifest.csv containing each source URL, output path, status, and error. Set CONCURRENCY to adjust parallel captures.
import fs from 'node:fs/promises';
import path from 'node:path';
import { parse } from 'csv-parse/sync';
import { chromium } from 'playwright';
const inputPath = process.argv[2] ?? 'urls.csv';
const outputDir = process.argv[3] ?? 'screenshots';
const concurrency = Math.max(1, Number(process.env.CONCURRENCY ?? 3));
const timeoutMs = Math.max(1000, Number(process.env.TIMEOUT_MS ?? 30000));
const fullPage = process.env.FULL_PAGE === '1';
const imageType = (process.env.FORMAT ?? 'png').toLowerCase();
const allowedTypes = new Set(['png', 'jpeg']);
if (!allowedTypes.has(imageType)) throw new Error('FORMAT must be png or jpeg');
const csvText = await fs.readFile(inputPath, 'utf8');
const records = parse(csvText, { columns: true, skip_empty_lines: true, trim: true, bom: true });
if (records.length === 0) throw new Error('CSV has a header row but no URL records');
if (!Object.hasOwn(records[0], 'url')) throw new Error('CSV must have a url column');
await fs.mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = new Array(records.length);
let nextIndex = 0;
function safePart(value) {
return String(value).normalize('NFKD').replace(/[^a-zA-Z0-9._-]+/g, '-').replace(/^-+|-+$/g, '').slice(0, 70) || 'page';
}
function escapeCsv(value) {
const text = String(value ?? '');
return /[",\r\n]/.test(text) ? `"${text.replaceAll('"', '""')}"` : text;
}
async function worker() {
while (true) {
const index = nextIndex++;
if (index >= records.length) return;
const row = records[index];
const rawUrl = row.url?.trim();
const id = row.id?.trim() || String(index + 1).padStart(3, '0');
let url;
let file = '';
try {
if (!rawUrl) throw new Error('Missing URL');
url = new URL(rawUrl);
if (!['http:', 'https:'].includes(url.protocol)) throw new Error('URL must use http or https');
file = `${safePart(id)}-${safePart(url.hostname)}.${imageType === 'jpeg' ? 'jpg' : imageType}`;
const page = await browser.newPage({ viewport: { width: 1440, height: 900 }, deviceScaleFactor: 1 });
try {
page.setDefaultNavigationTimeout(timeoutMs);
const response = await page.goto(url.href, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
// A reached page can return an HTTP error status and still be useful evidence.
await page.screenshot({ path: path.join(outputDir, file), fullPage, type: imageType });
results[index] = { id, url: url.href, file, status: 'ok', http_status: response?.status() ?? '', error: '' };
} finally {
await page.close();
}
} catch (error) {
results[index] = { id, url: rawUrl ?? '', file, status: 'error', http_status: '', error: error instanceof Error ? error.message : String(error) };
}
}
}
try {
await Promise.all(Array.from({ length: Math.min(concurrency, records.length) }, () => worker()));
} finally {
await browser.close();
}
const columns = ['id', 'url', 'file', 'status', 'http_status', 'error'];
const manifest = [columns.join(','), ...results.map(row => columns.map(key => escapeCsv(row[key])).join(','))].join('\n') + '\n';
await fs.writeFile(path.join(outputDir, 'manifest.csv'), manifest, 'utf8');
const failed = results.filter(row => row.status === 'error').length;
console.log(`Finished ${results.length} rows: ${results.length - failed} captured, ${failed} failed. Manifest: ${path.join(outputDir, 'manifest.csv')}`);
if (failed) process.exitCode = 1;
The capture waits for domcontentloaded, which is a practical default for pages that continue loading analytics or other background resources. It records the returned HTTP status but still saves a screenshot for responses such as 404 pages: an error page can itself be useful in a report. A navigation failure, timeout, or invalid URL is recorded as an error and does not stop other rows.
4. Choose capture size, format, and timing
| Need | Setting | Effect |
|---|---|---|
| Only the visible browser viewport | Default | Captures the current viewport, 1440 × 900 CSS pixels at device scale factor 1. |
| The full scrollable page | FULL_PAGE=1 |
Captures beyond the viewport. Very long pages can produce large images and take longer. |
| JPEG instead of PNG | FORMAT=jpeg |
Writes JPEG files, which can be smaller for photographic pages but are lossy. |
| More or fewer simultaneous pages | CONCURRENCY=5 |
Controls how many pages this script captures at once. Start low if memory or target sites are constrained. |
| Longer navigation allowance | TIMEOUT_MS=60000 |
Sets the navigation timeout in milliseconds. |
Examples:
FULL_PAGE=1 node capture.mjs urls.csv screenshots
FORMAT=jpeg CONCURRENCY=2 node capture.mjs urls.csv screenshots
TIMEOUT_MS=60000 CONCURRENCY=2 FULL_PAGE=1 node capture.mjs urls.csv screenshots
To capture one element rather than the page, use Playwright’s locator screenshot API in the page capture section, for example await page.locator('main').screenshot({ path: path.join(outputDir, file), type: imageType }). Replace the page screenshot call; choose a selector that exists on all target pages or record missing-element errors per URL. Playwright documents page, full-page, and element screenshot options in its screenshot guide.
For pages that render content after navigation, wait for a known selector before capturing, such as await page.locator('main').waitFor({ state: 'visible', timeout: timeoutMs }). A fixed delay can be added with await page.waitForTimeout(1500), but selector-based waits are usually more targeted. Avoid waiting for all network activity to stop on sites with persistent connections; it can waste time or never complete.
5. Assemble the report and audit the captures
- Open
screenshots/manifest.csvand filterstatusforokorerror. - Match each row to its image using the
filevalue. Preserve the originalurlin report captions or source notes. - Review a sample of images for consistent viewport, page state, and capture date before distributing the report.
- Retry only failed rows after correcting their cause; keep the manifest with the report so missing captures are visible.
Stable row IDs make filenames predictable and prevent identical hostnames from overwriting each other when the IDs differ. If IDs repeat, the script can overwrite a prior image, so ensure they are unique or add a row number to the naming expression.
6. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. A GET request captures one URL; its bulk capture option accepts up to 100 URLs per call. For a single request, use the documented API pattern below and adapt the URL and filename for each CSV row. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Sign up for free and get 1,000 screenshots a month with no card.
Performance, reliability, and cost considerations
- Concurrency: More workers can reduce elapsed time, but each browser page consumes memory and adds requests to target sites. Increase gradually and respect site access policies.
- Timeouts: A larger timeout helps slow pages but makes a failed row take longer. Keep failures in the manifest and rerun them separately where possible.
- Repeatability: Fix viewport, device scale, format, and wait condition for all rows in a report. Dynamic pages can still vary by time, location, login state, or personalized content.
- Image size: Full-page PNGs can be large. Consider JPEG for photographic content or viewport captures when a report does not require the entire page.
- Local cost: The script has no per-screenshot API charge, but it uses your machine’s CPU, memory, storage, and network. Hosted API pricing and limits depend on the service and plan; check current terms before processing a large CSV.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
CSV must have a url column |
The header is absent, misspelled, or uses different capitalization. | Name the header exactly url; keep a header row in the file. |
Missing URL or invalid URL |
A row has an empty value or lacks a valid scheme. | Fill the URL and use a full https:// or http:// address. |
| Timeout or navigation error | The host is slow, unreachable, or blocks the request. | Check the URL in a browser, raise TIMEOUT_MS, or rerun that row. Do not increase concurrency to address a slow target. |
| Screenshot is blank or incomplete | Content renders after DOM navigation, depends on interaction, or is lazy-loaded. | Wait for a stable content selector; use full-page capture where needed. Some pages require authentication or user actions and will need an explicit setup. |
| Some rows have HTTP status 4xx/5xx but an image exists | The server returned an error document as a normal page response. | Review the image and recorded status. The script preserves it as evidence instead of discarding it. |
| Out of memory or machine becomes slow | Too many pages are open at the same time or full-page images are large. | Lower CONCURRENCY, capture the viewport, or process the CSV in smaller batches. |
| Images overwrite one another | Rows share the same ID and hostname, producing the same filename. | Use unique IDs or include the row index in the filename. |
Cannot find package or browser launch failure |
Dependencies or the Chromium browser binary are missing. | Run npm install playwright csv-parse and npx playwright install chromium in the project. |
FAQ
Should I capture the viewport or the full page?
Use the viewport when the report compares the initial visible state. Use full-page capture when below-the-fold content is part of the evidence; it may create much taller files.
Can I export a PDF instead of image files?
Playwright’s CLI documents PDF export as well as screenshot output. For this CSV workflow, image files plus a manifest make it straightforward to arrange selected captures into a report. Choose PDF when each page’s print layout is the deliverable.
Can I rerun only failed URLs?
Yes. Use the manifest’s error rows to create a smaller input CSV, then run the same script. Keeping the original IDs preserves traceability.
Does a screenshot prove what every visitor saw?
No. It records one browser session and configuration. Personalization, geography, authentication, and time-dependent content can produce different results.


