How to Capture Screenshots of URLs from a CSV File with Playwright
Read URLs from a CSV, capture viewport or full-page screenshots with Playwright, and handle naming, failures, readiness, and batch size.
Use Node.js to parse the CSV with a real CSV parser, open each URL with Playwright, and save a screenshot for each successful navigation. The example below expects a header named url, skips blank lines, writes full-page PNGs into a screenshots directory, and records individual failures so one bad URL does not stop the batch.
This is a standalone Playwright script; you do not need the Playwright Test runner for a simple batch job. Playwright’s documentation notes that its test runner runs in Node.js, so tests can read files and parse them with a CSV library. The same Node.js file and parser approach works for a standalone script. See Playwright’s CSV parameterization guide.
1. Prepare the CSV and project
Save a file named urls.csv in the project directory:
url
https://example.com/
https://playwright.dev/
https://screenshotneo.com/
The header matters: the script reads row.url. Keep a proper CSV parser even if the first version has one column. Splitting lines on commas breaks when fields are quoted or contain commas, and it becomes fragile as the file gains more columns.
Initialize a Node.js project, install Playwright and the CSV parser, and install Chromium:
npm init -y
npm install playwright csv-parse
npx playwright install chromium
Use an ES module project. Add "type": "module" to package.json, or save the script with an .mjs extension. Playwright’s browser installation is separate from installing the package; install the browser you intend to run.
2. Create the screenshot script
Save the following as capture-csv.mjs. It creates predictable row-based filenames, includes a sanitized hostname to make files easier to identify, writes a manifest, and continues after a row fails.
import fs from 'node:fs';
import path from 'node:path';
import { parse } from 'csv-parse/sync';
import { chromium } from 'playwright';
const inputFile = 'urls.csv';
const outputDir = 'screenshots';
const fullPage = true;
const rows = parse(fs.readFileSync(inputFile), {
columns: true,
skip_empty_lines: true,
trim: true,
});
fs.mkdirSync(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
for (const [index, row] of rows.entries()) {
const rowNumber = index + 2; // Header is row 1 in the CSV.
const url = row.url?.trim();
if (!url) {
results.push({ row: rowNumber, url: '', status: 'skipped', reason: 'Missing url' });
continue;
}
let parsedUrl;
try {
parsedUrl = new URL(url);
if (!['http:', 'https:'].includes(parsedUrl.protocol)) {
throw new Error('URL must use http or https');
}
} catch (error) {
results.push({ row: rowNumber, url, status: 'failed', reason: error.message });
console.error(`Invalid URL on CSV row ${rowNumber}: ${url} (${error.message})`);
continue;
}
const host = parsedUrl.hostname.replace(/[^a-z0-9.-]/gi, '_');
const filename = `${String(rowNumber).padStart(4, '0')}-${host}.png`;
const outputPath = path.join(outputDir, filename);
try {
const response = await page.goto(url, {
waitUntil: 'load',
timeout: 30_000,
});
// A non-2xx response can still render a useful error page. Keep it and
// record the HTTP status so the batch output remains inspectable.
await page.screenshot({ path: outputPath, fullPage });
results.push({
row: rowNumber,
url,
status: 'captured',
httpStatus: response?.status() ?? null,
file: outputPath,
});
console.log(`Saved ${outputPath}`);
} catch (error) {
results.push({ row: rowNumber, url, status: 'failed', reason: error.message });
console.error(`Failed CSV row ${rowNumber} (${url}): ${error.message}`);
}
}
} finally {
await browser.close();
fs.writeFileSync('screenshots-manifest.json', JSON.stringify(results, null, 2));
}
const failed = results.filter((result) => result.status === 'failed').length;
console.log(`Finished: ${results.filter((result) => result.status === 'captured').length} captured, ${failed} failed.`);
if (failed > 0) process.exitCode = 1;
Run it with:
node capture-csv.mjs
The row number in each filename prevents duplicate hostnames from overwriting each other. The manifest maps each CSV row and URL to its outcome and output path. This is a script design choice; adapt the filename scheme or manifest format to what consumes the images downstream.
3. Choose the capture size and page readiness
Viewport screenshot or full page
By default, page.screenshot() captures the visible viewport. Set fullPage: true to capture the full scrollable page. The Playwright Page screenshot API documents both the screenshot options and full-page behavior.
Use viewport captures for a consistent above-the-fold snapshot or visual comparison. Use full-page captures when the content below the fold matters. Very long pages can produce large images and take longer to render or save. Some sites load content only when scrolled; if a lazy-loaded section is absent, use a site-specific scroll or readiness step before capture.
Navigation readiness is site-dependent
The script uses waitUntil: 'load' as a practical default. It means the page’s load event has fired; it does not guarantee that every application has finished fetching data, animating, or displaying its final layout. Choose a readiness condition based on the pages you capture.
If a known element indicates that the content is ready, wait for it after navigation:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('main article').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: outputPath, fullPage: true });
Replace main article with a selector that is meaningful for the target site. For a collection of unrelated domains, there may be no single selector that works everywhere. In that case, keep a general navigation condition and allow per-site configuration or a modest delay where you have confirmed it is needed. Avoid assuming that one wait strategy guarantees a complete render for every URL.
4. Adapt the CSV and output behavior
Use a different URL column
If your header is page_url, change row.url to row.page_url. To tolerate either name:
const url = (row.url ?? row.page_url ?? '').trim();
For a CSV with no header, set columns: false and read the first field, such as row[0]. If the file has a header plus extra columns, columns: true makes each row an object keyed by its column names. The parser also supports options for delimiter, quote handling, and other CSV formats; configure those to match the actual input.
Capture a viewport instead
Change const fullPage = true to false, or pass fullPage: false directly. Set the viewport at context or page creation to make viewport captures consistent:
const page = await browser.newPage({ viewport: { width: 1365, height: 768 } });
Save JPEG or WebP
Playwright’s screenshot API supports formats based on the output path and screenshot options. For JPEG, use a .jpg path and set type: 'jpeg'; JPEG supports a quality option from 0 to 100. WebP can be requested with type: 'webp' where supported by the installed browser. Match the filename extension to the selected format.
await page.screenshot({
path: outputPath.replace(/\.png$/, '.jpg'),
type: 'jpeg',
quality: 80,
fullPage,
});
Use one page per URL
The example reuses one page to keep the loop simple. For stricter isolation between sites, create a new browser context or page for each URL, and close it after capture. This can prevent cookies and page state from a previous navigation from influencing a later one, at the cost of more setup and resource use.
for (const row of rows) {
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
const page = await context.newPage();
try {
// Validate the row URL, navigate, and capture here.
} finally {
await context.close();
}
}
Process a large input in chunks
The synchronous CSV parser loads the file and all rows into memory. It is convenient for small and medium input files. For a very large CSV, use a streaming parser and process rows incrementally. Keep concurrency bounded if you add parallel pages: each browser page uses resources, and simultaneous requests can increase load on both your machine and the destination sites.
5. cURL, Python, and Node.js alternatives
Playwright is a Node.js browser automation library, so the complete CSV-to-Playwright workflow above is JavaScript. cURL is useful for an HTTP screenshot API call, not for controlling a local Playwright browser. Python can parse a CSV and call an API, or automate a browser with a Python browser library, but it is not the language used by this title’s Playwright example.
cURL: one screenshot through ScreenshotNeo
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python: one screenshot through ScreenshotNeo
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js: one screenshot through ScreenshotNeo
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
These API examples show a single request. To process a CSV with an API, keep the same parser and row loop, make a request for each valid URL, and write each response to a distinct file. See the ScreenshotNeo API documentation for request options and response details.
6. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A GET request can return a PNG, JPEG, WebP, or PDF. Here is the one-call example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
It accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up free for 1,000 screenshots a month, with no card.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Cannot use import statement outside a module |
Node is treating the file as CommonJS. | Use the .mjs extension or add "type": "module" to package.json. |
| Browser executable is missing | The Playwright package is installed but its browser has not been installed. | Run npx playwright install chromium in the project environment. |
row.url is undefined |
The CSV header differs, has a byte-order mark, or the file has no header. | Inspect the first line and parsed keys; update the column name or parser configuration. For a headerless file, use columns: false. |
| Every row fails with invalid URL | Values may be blank, lack a scheme, or contain whitespace or unexpected characters. | Trim values, require an explicit http:// or https:// URL, and keep invalid rows in the manifest for correction. |
| Navigation times out | The site is slow, unreachable, blocking automation, or waiting indefinitely on a resource. | Check the URL from the same machine, set a deliberate timeout, and select a suitable navigation readiness event. Record the failure and continue the batch. |
| Screenshot looks incomplete | The application renders after the load event, or content is lazy-loaded. | Wait for a relevant selector or application readiness signal; scroll if the site loads content on scroll. A generic delay is not a universal fix. |
| Images are overwritten | Filenames are based only on hostname and multiple input rows share that host. | Include a row index or another unique input identifier in every filename. |
| Batch is slow or the machine runs out of memory | Full-page images, large input, or too many concurrent pages consume resources. | Capture the viewport when sufficient, stream large CSVs, and limit concurrent pages. Start sequentially and measure your own workload before increasing concurrency. |
| Visual output changes between runs | Rendering varies with operating system, browser version, settings, hardware, power source, and headless mode. | For visual comparisons, use the same execution environment and browser setup. Playwright discusses these sources of variation in its visual comparisons guide. |
8. Performance, reliability, and cost
A local Playwright batch has no per-screenshot service fee, but it uses your compute, storage, and network connection. Actual throughput depends on the pages, machine, browser, image dimensions, and readiness waits; there is no universal rate to assume. Sequential navigation is the simplest reliable baseline. If you need more throughput, add bounded concurrency, watch memory use, and respect destination site limits.
For reliability, keep a manifest, preserve the row number and URL for every failure, set a finite navigation timeout, and close the browser in a finally block. Decide whether a non-2xx page should still be captured; the example captures it and records its status. Retry only failures that make sense to retry, with a limit, so a persistently unavailable URL does not stall the entire batch.
For repeatable visual comparisons, pin the runtime and browser installation in your environment and run captures under consistent settings. Playwright warns that screenshots can differ across operating systems, versions, settings, hardware, power sources, and headless mode. See its visual comparison guidance.
9. Frequently asked questions
Can I use a CSV with more than one URL column?
Yes. Choose the column for each row explicitly, or generate more than one capture per row. Keep output names unique across both the row and selected URL.
Does a screenshot capture the page’s current state or a PDF?
page.screenshot() saves an image. If the deliverable is a PDF, use Playwright’s PDF capability in a supported browser workflow or a screenshot API that returns PDFs; an image capture is not itself a PDF document.
Can the script run in CI?
Yes, provided the CI environment installs the required browser and has network access to the target pages. Keep the browser environment consistent if outputs are compared over time, and store generated files or the manifest as build artifacts according to your pipeline.
Why use a CSV parser for a one-column file?
It handles CSV quoting, headers, and blank lines correctly and makes it easier to extend the input later. A line split is only safe for a tightly constrained format that you control.


