How to Capture a Batch of URLs with Playwright and Save One File per Page
Capture a list of URLs with Playwright, save one uniquely named PNG per page, and handle navigation errors, HTTP failures, and readiness.
To capture a batch of URLs with Playwright and save one file per page, navigate to each URL, take a screenshot, and give each output a unique filename. The runnable Node.js example below uses Chromium, visits URLs sequentially, and saves full-page PNGs. It continues after an individual URL fails and records HTTP status separately, since an HTTP 404 or 500 does not by itself make page.goto() throw.
This guide assumes PNG screenshots. If you need PDFs or saved HTML instead, the output step and its options differ.
1. Install Playwright and prepare a URL list
In an existing Node.js project, install Playwright and its Chromium browser:
npm install playwright
npx playwright install chromium
Save the script below as capture-batch.mjs. It uses ES modules, available with the .mjs extension. The URLs can be edited in the script or replaced with values loaded from a file or command-line arguments.
2. Capture one full-page PNG per URL
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';
const urls = [
'https://example.com',
'https://playwright.dev',
'https://example.com/about',
];
const outputDir = 'captures';
const timeoutMs = 30_000;
const failOnHttpError = false;
function safeLabel(input) {
try {
const url = new URL(input);
const label = `${url.hostname}${url.pathname}`
.replace(/[^a-zA-Z0-9]+/g, '-')
.replace(/^-+|-+$/g, '')
.slice(0, 70);
return label || 'page';
} catch {
return 'invalid-url';
}
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const results = [];
try {
// One page is enough for a sequential batch. A fresh context starts without
// cookies or local storage from a previous browser task.
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
const page = await context.newPage();
page.setDefaultNavigationTimeout(timeoutMs);
for (const [index, url] of urls.entries()) {
const number = String(index + 1).padStart(3, '0');
const filename = `${number}-${safeLabel(url)}.png`;
const outputPath = path.join(outputDir, filename);
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
const status = response?.status() ?? null;
if (failOnHttpError && status !== null && status >= 400) {
throw new Error(`HTTP ${status}`);
}
// For client-rendered sites, wait for a meaningful page-specific signal
// here, for example: await page.locator('main').waitFor();
await page.screenshot({ path: outputPath, fullPage: true });
results.push({ url, status, file: outputPath, ok: true });
console.log(`${status ?? 'no main response'} ${url} -> ${outputPath}`);
} catch (error) {
results.push({ url, file: outputPath, ok: false, error: String(error) });
console.error(`Failed ${url}: ${String(error)}`);
}
}
await context.close();
} finally {
await browser.close();
}
const failed = results.filter((result) => !result.ok);
console.log(`\nFinished: ${results.length - failed.length} succeeded, ${failed.length} failed.`);
if (failed.length) process.exitCode = 1;
Run it with:
node capture-batch.mjs
The numeric prefix ensures output paths stay unique even when URLs repeat or different URLs produce the same sanitized label. The label is for readability only; it is not used as a path derived directly from untrusted URL text.
3. Choose navigation and screenshot options
Navigation readiness
page.goto() accepts readiness choices such as commit, domcontentloaded, and load. Use the earliest condition that reliably leaves the content you need available. A client-rendered page may need an additional application-specific wait, such as a locator becoming visible. Playwright describes networkidle as discouraged for testing; it waits for at least 500 ms without network connections and can be a poor fit for pages that keep connections open.
commit: navigation response has started loading; useful when you will explicitly wait for content afterward.domcontentloaded: initial HTML has been parsed; a practical default for many captures.load: page load event has fired, including dependent resources such as images and stylesheets.- Locator or app signal: best when the screenshot depends on a particular rendered component or data state.
Viewport or entire page
fullPage: true captures the full scrollable page. Omit it to capture the current viewport. Very long pages can create large images and use more memory; if downstream limits matter, capture a viewport or divide the page into sections.
Context and page reuse
A browser context can contain multiple pages. For a serial batch, one page reused across navigations is the simplest shape. Use multiple pages when you need multiple tabs at once. Use a separate context per job when cookie and local storage isolation is important; the script’s fresh non-persistent context does not inherit another task’s browser state.
If captures intentionally need an authenticated session, configure the context with the appropriate storage state or authenticate it before the loop. Keep credentials out of source control, and avoid reusing a signed-in context across unrelated targets.
HTTP responses and output formats
A navigation can resolve with an HTTP error response. Inspect response.status() if 4xx or 5xx pages should be treated as failures. The example logs such statuses and captures them by default; set failOnHttpError to true to skip screenshotting responses with status 400 or higher.
For a PDF instead of PNG, Playwright exposes page.pdf() in Chromium. PDF output has its own page size, print background, and margin choices. HTML snapshots require saving the DOM separately and do not preserve the rendered appearance as a screenshot does.
4. Load URLs from a text file
For a larger list, keep one URL per line in urls.txt and replace the hard-coded array with file loading. This runnable variant trims whitespace and skips blank lines and comment lines:
import { readFile } from 'node:fs/promises';
const urls = (await readFile('urls.txt', 'utf8'))
.split(/\r?\n/)
.map((line) => line.trim())
.filter((line) => line && !line.startsWith('#'));
Insert that snippet in place of the const urls = [...] declaration. Validate input from untrusted sources and allow only URL schemes and hosts your workflow is intended to capture.
5. Retry transient failures without hiding them
Retries can help with temporary network errors, but repeated captures can also waste time or overload a target. Keep retries bounded, use a delay, and preserve the final failure in the run summary. For example, wrap the page.goto() call in a small retry loop for thrown errors; do not retry every HTTP status automatically unless that is appropriate for your use case. A 404 is generally not transient.
For reproducible captures, record the URL, status, output filename, and error for every item. If the process is interrupted, the index-based names make it straightforward to identify completed outputs and rerun the remaining URLs.
6. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
Executable doesn't exist or browser launch fails |
The Playwright package is installed but its browser binary is not. | Run npx playwright install chromium in the project environment. |
| Navigation times out | The server is slow, unreachable, or the chosen readiness event never occurs. | Set a realistic navigation timeout, check the URL and network access, and try a readiness event suited to the page. Wait for a specific locator after navigation when needed. |
| Screenshot shows a blank or incomplete app | DOM content loaded before client-side rendering or data fetching finished. | Wait for a page-specific selector or readiness condition before calling screenshot(). |
| A 404/500 image is saved as if successful | goto() does not throw solely because a valid HTTP response has an error status. |
Check response.status() and enable the example’s failOnHttpError behavior if those responses should be skipped. |
| Files overwrite one another | Names came only from a hostname or sanitized path, which can collide for duplicate URLs. | Include a stable index or unique identifier in every output name, as the example does. |
| Later pages appear signed in or personalized | The same page/context carries cookies or local storage between navigations. | Use a new context for each independent job, or deliberately configure the session state required for the capture. |
| The batch stops at the first bad URL | An exception escapes the loop. | Catch errors inside the per-URL loop, log them, and continue; set a nonzero exit code after processing if any failed. |
| Full-page screenshots are huge or memory-heavy | The target has an unusually long page or very large assets. | Use viewport capture, reduce the viewport or split capture into sections, and process a smaller number concurrently. |
7. Performance, reliability, and cost
Sequential capture uses one browser, context, and page, which keeps resource use and session behavior easier to reason about. Its tradeoff is that each navigation waits for the previous one. If throughput matters, add bounded concurrency with a small number of pages or workers and measure resource use on your own workload; no universal speed figure applies. Avoid launching an unbounded browser per URL.
Navigation readiness affects both time and completeness: waiting for load can wait for more resources than domcontentloaded, while an explicit locator can wait for exactly the content that matters. Set timeouts, isolate per-URL failures, and retain a result manifest so transient failures can be retried without losing successful captures.
Playwright is software you run with your own browser environment, so account for the compute, browser installation, storage, and maintenance involved in operating it. This method has no per-screenshot ScreenshotNeo charge, but your infrastructure has its own costs.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request captures a URL as PNG, JPEG, WebP, or PDF. The examples below use the documented API; see the ScreenshotNeo API documentation for parameters and response details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For a batch, make one request per URL and give each returned file a unique name using the same index-and-label approach shown in the Playwright script. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
9. Frequently asked questions
Can I capture the same URL more than once?
Yes. Each iteration gets a distinct index in its filename, so repeated URLs do not overwrite earlier captures.
Does a successful screenshot prove the page returned HTTP 200?
No. A screenshot can be saved for an HTTP error page. Check and record the navigation response status when status matters.
Should I use one page or one browser per URL?
Reuse one page for a simple sequential list. Use separate pages for simultaneous tabs, or separate contexts when jobs need independent browser state.
How do I make captures consistent?
Use the same viewport, browser engine, readiness signal, and session setup for each run. Dynamic content can still vary unless the target provides a stable test state.


