How to Take Batch Screenshots of URLs Using Playwright Workers
Capture many URLs with Playwright by scheduling screenshot jobs with bounded concurrency, saving per-URL results, and handling failures cleanly.
To take batch screenshots with Playwright, load your URLs, run a fixed number of capture jobs concurrently, and save a separate result for each URL. In a custom script, implement that limit with a queue or worker loop: Playwright Test’s workers setting applies to work scheduled by the test runner and does not automatically distribute an arbitrary URL list.
This guide uses Node.js and Playwright. It includes a runnable batch script, explains worker count and browser/context lifetimes, and covers retries, output naming, reproducibility, and common failures.
1. Install Playwright and prepare your URL list
Create a project and install Playwright. The install command downloads the package; install the Chromium browser used by the example as well.
npm init -y
npm install playwright
npx playwright install chromium
Save one URL per line in urls.txt:
https://example.com
https://www.wikipedia.org
https://playwright.dev
Use complete URLs with a scheme such as https://. The script below validates them, drops blank lines, and creates deterministic output names using the input position and a short hash. The position keeps duplicate URLs from overwriting each other.
2. Runnable Node.js batch screenshot script
Save this as batch-screenshots.mjs. It runs at most WORKERS jobs at once, writes screenshots under screenshots/, and produces a JSON report with a status for every valid input URL. Each job gets a fresh browser context, while the browser process is reused for the batch.
import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
const INPUT = process.argv[2] ?? 'urls.txt';
const OUTPUT_DIR = process.argv[3] ?? 'screenshots';
const WORKERS = Math.max(1, Number(process.env.WORKERS ?? 3));
const FULL_PAGE = process.env.FULL_PAGE === '1';
const NAVIGATION_TIMEOUT_MS = Math.max(1000, Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000));
function parseUrls(text) {
return text.split(/\r?\n/)
.map((line) => line.trim())
.filter(Boolean)
.map((raw, index) => {
let url;
try {
url = new URL(raw);
} catch {
throw new Error(`Invalid URL on non-empty input line ${index + 1}: ${raw}`);
}
if (url.protocol !== 'http:' && url.protocol !== 'https:') {
throw new Error(`Unsupported URL scheme on line ${index + 1}: ${url.protocol}`);
}
return { input: raw, normalized: url.href, index };
});
}
function outputPath(job) {
const host = new URL(job.normalized).hostname.replace(/[^a-zA-Z0-9.-]/g, '_');
const hash = createHash('sha256').update(job.normalized).digest('hex').slice(0, 10);
const extension = 'png';
return path.join(OUTPUT_DIR, `${String(job.index + 1).padStart(4, '0')}-${host}-${hash}.${extension}`);
}
async function capture(browser, job) {
const startedAt = Date.now();
const destination = outputPath(job);
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
try {
const page = await context.newPage();
const response = await page.goto(job.normalized, {
waitUntil: 'load',
timeout: NAVIGATION_TIMEOUT_MS,
});
await page.screenshot({ path: destination, fullPage: FULL_PAGE, type: 'png' });
return {
input: job.input,
url: job.normalized,
status: 'success',
httpStatus: response?.status() ?? null,
screenshot: destination,
elapsedMs: Date.now() - startedAt,
};
} catch (error) {
return {
input: job.input,
url: job.normalized,
status: 'error',
error: error instanceof Error ? error.message : String(error),
elapsedMs: Date.now() - startedAt,
};
} finally {
await context.close();
}
}
async function main() {
const jobs = parseUrls(await readFile(INPUT, 'utf8'));
await mkdir(OUTPUT_DIR, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = new Array(jobs.length);
let next = 0;
async function worker() {
while (true) {
const index = next++;
if (index >= jobs.length) return;
results[index] = await capture(browser, jobs[index]);
}
}
try {
await Promise.all(Array.from({ length: Math.min(WORKERS, jobs.length) }, () => worker()));
} finally {
await browser.close();
}
await writeFile(path.join(OUTPUT_DIR, 'results.json'), JSON.stringify(results, null, 2) + '\n');
const failures = results.filter((result) => result.status === 'error').length;
console.log(`Finished ${results.length} URL(s): ${results.length - failures} succeeded, ${failures} failed.`);
console.log(`Screenshots and report: ${OUTPUT_DIR}`);
if (failures) process.exitCode = 1;
}
main().catch((error) => {
console.error(error);
process.exitCode = 1;
});
Run it with the default settings:
node batch-screenshots.mjs urls.txt screenshots
Set options through environment variables:
WORKERS=2 FULL_PAGE=1 NAVIGATION_TIMEOUT_MS=45000 node batch-screenshots.mjs urls.txt screenshots
WORKERS is the maximum number of active URL jobs. FULL_PAGE=1 captures the entire scrollable page; the default captures the visible viewport. NAVIGATION_TIMEOUT_MS bounds navigation wait time. The script waits for the page’s load event, not for every later network request to stop. Pages that continue polling or load content lazily may need a different readiness condition or an explicit wait.
3. How the worker queue works
The next counter assigns each URL to one worker. Each worker takes one job, awaits it, then takes another until the list is exhausted. This bounds simultaneous browser work without creating one promise per URL that all begins navigation at once. Results are stored by input index, so the report follows the original list even when captures complete in a different order.
The example uses Playwright’s Chromium browser, a fresh non-persistent context for each URL, and one page in that context. A browser context is an isolated session; pages are tabs within a context. Fresh contexts help prevent cookies or local storage from one URL’s session from carrying into another job. If your URLs intentionally share login state, you can reuse a context and create pages within it, but then concurrent jobs may share mutable session state. See the official documentation for BrowserContext lifecycle and isolation and pages and contexts.
For very long batches, consider writing each result to a durable report as soon as its job completes. The sample holds all result records in memory and writes the JSON report after the batch; that is appropriate for ordinary URL lists but means an unexpected process stop can leave screenshots without a completed report.
4. Choosing a worker count
There is no universally correct worker count. More workers can increase throughput, but also increase concurrent browser and page resource use and the load sent to target sites. Start with a small cap such as the example’s three, then measure elapsed time, memory use, and failure rate on your own pages. Reduce the cap if the machine runs short of memory, sites begin throttling, or navigation failures rise.
| Workload or constraint | Starting choice | Reason |
|---|---|---|
| Unknown sites and a modest machine | 1–3 workers | Establishes a baseline with limited parallel load; tune from measurements. |
| Heavy pages, large full-page captures | Lower the cap | Each active page can consume significant browser and image memory. |
| Small, stable pages on a capable machine | Increase gradually | More concurrency may shorten the batch, subject to target limits. |
| Visual regression suite expressed as tests | Configure Playwright Test workers | The test runner schedules test work across its worker processes. |
Playwright Test workers are independent OS processes, each of which starts its own browser. The Test runner supports a worker limit in configuration or on the command line. Those controls apply when URL captures are represented as Playwright Test work; for a custom script like the one above, the script’s queue controls concurrency. See Playwright Test parallelism and Test configuration.
When to use Playwright Test workers instead
Use the Test runner when each capture is naturally a test case and you want its test lifecycle and reporting. Put each URL in a test or test data set, then limit runner parallelism with a configuration setting or --workers. Use a custom queue when the input is a general list and you want a simple script to control URL assignment, output naming, and per-URL status. These are architectural choices; the documentation does not provide a universal performance winner for screenshot batches.
5. Capture scope, readiness, and output
Viewport or full page
The default page.screenshot() captures the visible viewport. The sample switches to full-page mode with fullPage: true. Full-page images can be much taller and larger than viewport images, so they take more storage and may need more memory to process. Choose based on whether reviewers need the initial screen or the whole document.
Playwright also lets you capture screenshot bytes instead of writing directly to a path, which is useful if another part of your pipeline uploads or transforms the result. See the Page API and screenshot documentation for full-page and buffer examples. Check the documentation version matching your installed Playwright release before relying on newer options.
Wait for the page you actually need
waitUntil: 'load' waits for the page load event. It does not guarantee that client-side rendering, web fonts, delayed images, or animations have settled. For a known page, wait for a selector that marks the content as ready, or add a short bounded delay when necessary. Avoid waiting indefinitely for “network idle” on pages that poll continuously.
For example, after navigation you could wait for a page-specific marker before taking the screenshot:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.locator('main article').waitFor({ state: 'visible', timeout: 10000 });
await page.screenshot({ path: destination, fullPage: true });
Use selectors that exist on the target pages and set a timeout. A selector that never appears should become a recorded failure or an explicit fallback, not a job that stalls the whole batch.
Stable naming and duplicate URLs
Keep one result per input row. The example includes the row number and a hash of the normalized URL, so repeated URLs get distinct filenames and long query strings do not become path components. If the same batch is rerun, those names are stable. Decide whether a rerun should overwrite old files, write to a run-specific directory, or compare against them before starting.
6. Reliability: retries, isolation, and reproducibility
- Retry selectively. Retry transient navigation timeouts or temporary network errors a limited number of times. Do not repeatedly retry invalid URLs, unsupported schemes, or persistent access-denied responses.
- Keep successes. The JSON report allows a retry job to select only failed URLs instead of repeating completed work.
- Respect target sites. A concurrency cap limits your own parallelism but does not guarantee a site will accept that request rate. Lower concurrency or schedule work over time if sites throttle requests.
- Isolate state intentionally. A separate context per job prevents browser session data from being shared. Reuse contexts only when shared state is desired and safe.
- Avoid shared mutable inputs. Parallel jobs can race if they use the same account, change server-side data, or write to colliding paths. Use unique output names and coordinate shared test data.
- Close resources. Close each context in a
finallyblock and close the browser after all workers finish, as the example does. - Keep rendering stable. Browser screenshots can vary with host OS, browser version, settings, hardware, power source, and headless mode. For visual comparisons, run baseline and current captures in the same environment and record the Playwright and browser versions. See Playwright’s screenshot comparison guidance.
For long-running or scheduled work, also record a run identifier, start time, browser version, and configuration alongside the URL and result. These fields make it easier to tell whether a difference came from the page or from a changed capture environment.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
Executable doesn't exist or browser launch fails |
The Playwright package is installed but its browser binary is missing. | Run npx playwright install chromium. In managed environments, install browser system dependencies as required by Playwright. |
| Many jobs fail with navigation timeouts | Slow sites, an unsuitable timeout, transient network issues, or pages that never reach the chosen event. | Check the failing URLs and error messages, increase the bounded timeout if appropriate, or wait for a page-specific selector after domcontentloaded. |
| Screenshot is blank or content is missing | The page may need client-side rendering, authentication, a readiness selector, or additional time. | Wait for the content marker you need; verify required cookies or credentials in the context; inspect the response and page in a local headed run when diagnosing. |
| Some sites return an error page or deny access | The target may restrict automated traffic, require authentication, or show a bot check. | Check that you are authorized to capture the page and supply permitted session state if needed. Lower concurrency if the site is throttling. A screenshot script cannot guarantee access to every site. |
| Files overwrite or are missing | Output paths collide, the process stops before report writing, or the URL list contains invalid entries. | Use unique stable names, retain the per-input report, and validate input before launching the browser. For durable batches, persist each result as it completes. |
| Machine runs out of memory or slows down | Too many active pages, heavyweight pages, or large full-page images. | Reduce WORKERS, use viewport screenshots where sufficient, and process very large lists in chunks. |
| Captures differ across runs | Rendering environment changes or dynamic page content. | Pin the browser/package environment, use consistent viewport and readiness rules, and control dynamic content where possible. |
8. Performance, reliability, and cost
For a self-hosted Playwright batch, plan for browser startup, navigation, page rendering, screenshot encoding, and file writes. A single browser reused by a bounded set of workers avoids launching a new browser process for each URL in this example, while a new context per job keeps session state isolated. The best balance depends on page weight, available memory, capture scope, and target-site behavior; measure it with your own URL set rather than assuming a specific worker count or speedup.
Playwright itself is the automation library. Your operational costs come from the machine and infrastructure that run the browser, plus storage and any network or proxy services you choose. Full-page images and high concurrency can increase resource use. For reliability, keep a result record for each input, make failed jobs retryable, and preserve successful outputs.
9. Or skip the browser setup
If you do not want to install and operate browsers, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its capture flow accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. Its MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 shots a month without a card; paid plans start at $5 for 3,000 shots. See the ScreenshotNeo API documentation for options and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For a batch, call the endpoint once per URL or use its bulk capture option for up to 100 URLs per call. Sign up for 1,000 free screenshots a month with no card.
10. Frequently asked questions
Does Playwright’s --workers option work in this script?
No. That option controls Playwright Test runner workers. This standalone script uses its own WORKERS queue limit.
Should each URL use a new browser?
Usually the browser process can be reused for a batch. Use separate contexts when jobs need isolated cookies and storage; create separate browser processes only when your operational requirements call for that additional isolation.
Can I capture URLs that require a login?
Yes, when you are authorized and provide the necessary session state. Configure cookies or a login flow in the context and protect any credentials; do not place secrets in a public URL list or report.
Why does full-page output take longer or use more memory?
It has to capture and encode a taller image than a viewport shot. Use viewport mode when it satisfies the review task, and lower concurrency for very large pages.
Will the same URL always produce identical pixels?
No. Dynamic content and differences in browser or host rendering can change the result. Keep the capture environment stable and make page readiness explicit for visual comparisons.


