How to Take Bulk Website Screenshots
Capture a list of URLs reliably with Playwright, then compare local automation with a hosted batch API workflow.
Short answer: put your URLs in a file, open each one with a real browser, wait for the page state you need, and save either a viewport, full-page, or element screenshot. Playwright provides the navigation and screenshot primitives; your batch loop adds retries, concurrency limits, logging, and output naming.
This guide shows a local workflow you control, including complete Node.js and Python scripts, then a hosted alternative for teams that do not want to operate browsers.
1. Choose the capture workflow
| Approach | Best when | Trade-offs |
|---|---|---|
| Local Playwright script | You need browser state, custom JavaScript, private-network access, or post-processing. | You maintain browser binaries, dependencies, retries, and capacity. |
| Hosted batch service | You want URL-list input without running browsers yourself. | Review current limits, privacy terms, output options, failure reporting, and price before committing. |
For either approach, decide these settings before writing code:
- Viewport or full page: a viewport records only what is visible; Playwright’s
fullPageoption captures the entire scrollable page. Playwright documents both modes and screenshot parameters. - Whole page or element: use an element locator when you need one component rather than the document.
- File or bytes: write files for an archive; keep returned bytes in memory for hashing, uploads, or image processing.
- Readiness: use a selector, a known application signal, or a bounded delay. A fixed delay is not universal because pages load content differently.
- Concurrency: start conservatively and increase only when your machine and the target sites can handle it. There is no generally correct number.
2. Prepare a URL list
Create urls.txt with one absolute URL per line. Blank lines and lines beginning with # are ignored by the examples below.
https://example.com/
https://playwright.dev/docs/screenshots
https://stripe.com/
Use HTTPS where possible. If a URL contains spaces or special characters, store the fully encoded URL. Keep a stable input order if the screenshots will be compared across runs.
3. Bulk screenshots with Node.js and Playwright
Install Playwright and its browser once:
npm init -y
npm install playwright
npx playwright install chromium
Save this as bulk-screenshots.mjs. It creates one PNG per URL, captures the full page, limits parallel pages, retries transient failures, and writes a CSV report.
import fs from 'node:fs/promises';
import path from 'node:path';
import { chromium } from 'playwright';
const input = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const concurrency = Number(process.env.CONCURRENCY ?? 3);
const maxAttempts = Number(process.env.ATTEMPTS ?? 2);
const waitFor = process.env.WAIT_FOR ?? '';
const fullPage = (process.env.FULL_PAGE ?? 'true') === 'true';
const raw = await fs.readFile(input, 'utf8');
const urls = raw.split(/\r?\n/)
.map(line => line.trim())
.filter(line => line && !line.startsWith('#'));
await fs.mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const results = [];
let next = 0;
function safeName(url, index) {
const parsed = new URL(url);
const base = `${parsed.hostname}${parsed.pathname}`
.replace(/[^a-z0-9]+/gi, '-')
.replace(/^-|-$/g, '')
.toLowerCase();
return `${String(index + 1).padStart(4, '0')}-${base || 'page'}.png`;
}
async function worker() {
const context = await browser.newContext({
viewport: { width: 1440, height: 900 },
deviceScaleFactor: 1
});
const page = await context.newPage();
while (true) {
const index = next++;
if (index >= urls.length) break;
const url = urls[index];
const filename = safeName(url, index);
const target = path.join(outputDir, filename);
let lastError = '';
const started = Date.now();
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
if (waitFor) await page.locator(waitFor).waitFor({ state: 'visible', timeout: 15000 });
else await page.waitForLoadState('networkidle', { timeout: 15000 }).catch(() => {});
await page.screenshot({ path: target, fullPage });
results.push({ index, url, status: 'ok', file: target, attempts: attempt, ms: Date.now() - started });
lastError = '';
break;
} catch (error) {
lastError = error instanceof Error ? error.message : String(error);
if (attempt < maxAttempts) await new Promise(resolve => setTimeout(resolve, 1000 * attempt));
}
}
if (lastError) results.push({ index, url, status: 'failed', file: '', attempts: maxAttempts, ms: Date.now() - started, error: lastError });
}
await context.close();
}
await Promise.all(Array.from({ length: Math.max(1, concurrency) }, worker));
await browser.close();
results.sort((a, b) => a.index - b.index);
const header = 'index,url,status,file,attempts,ms,error';
const csv = [header, ...results.map(r => [r.index, r.url, r.status, r.file, r.attempts, r.ms, r.error ?? '']
.map(value => `"${String(value).replaceAll('"', '""')}"`).join(','))].join('\n');
await fs.writeFile(path.join(outputDir, 'report.csv'), csv + '\n');
console.log(`Completed ${results.filter(r => r.status === 'ok').length}/${urls.length}; report: ${outputDir}/report.csv`);
Run it with:
node bulk-screenshots.mjs urls.txt screenshots
CONCURRENCY=5 FULL_PAGE=false node bulk-screenshots.mjs urls.txt screenshots
WAIT_FOR='main' node bulk-screenshots.mjs urls.txt screenshots
Capture one element instead of the page
Replace the screenshot line with a locator screenshot after waiting for the component:
const card = page.locator('[data-testid="pricing-card"]').first();
await card.waitFor({ state: 'visible', timeout: 15000 });
await card.screenshot({ path: target });
Use bytes instead of writing immediately
const bytes = await page.screenshot({ fullPage: true, type: 'png' });
// Send bytes to storage, hash them, or pass them to an image-processing library.
4. Bulk screenshots with Python and Playwright
Install the package and Chromium:
python -m pip install playwright
python -m playwright install chromium
Save as bulk_screenshots.py:
import csv
import os
import re
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright
INPUT = Path(os.getenv("INPUT", "urls.txt"))
OUTPUT = Path(os.getenv("OUTPUT", "screenshots"))
CONCURRENCY = int(os.getenv("CONCURRENCY", "3"))
ATTEMPTS = int(os.getenv("ATTEMPTS", "2"))
FULL_PAGE = os.getenv("FULL_PAGE", "true").lower() == "true"
WAIT_FOR = os.getenv("WAIT_FOR", "")
urls = [line.strip() for line in INPUT.read_text().splitlines()
if line.strip() and not line.lstrip().startswith("#")]
OUTPUT.mkdir(parents=True, exist_ok=True)
def filename(url, index):
parsed = urlparse(url)
base = re.sub(r"[^a-zA-Z0-9]+", "-", parsed.netloc + parsed.path).strip("-").lower()
return OUTPUT / f"{index + 1:04d}-{base or 'page'}.png"
def capture(index_url):
index, url = index_url
target = filename(url, index)
started = time.time()
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page(viewport={"width": 1440, "height": 900}, device_scale_factor=1)
error = ""
for attempt in range(1, ATTEMPTS + 1):
try:
page.goto(url, wait_until="domcontentloaded", timeout=45000)
if WAIT_FOR:
page.locator(WAIT_FOR).wait_for(state="visible", timeout=15000)
else:
try:
page.wait_for_load_state("networkidle", timeout=15000)
except Exception:
pass
page.screenshot(path=str(target), full_page=FULL_PAGE)
browser.close()
return {"index": index, "url": url, "status": "ok", "file": str(target),
"attempts": attempt, "ms": int((time.time() - started) * 1000), "error": ""}
except Exception as exc:
error = str(exc)
if attempt < ATTEMPTS:
time.sleep(attempt)
browser.close()
return {"index": index, "url": url, "status": "failed", "file": "",
"attempts": ATTEMPTS, "ms": int((time.time() - started) * 1000), "error": error}
with ThreadPoolExecutor(max_workers=max(1, CONCURRENCY)) as pool:
results = list(pool.map(capture, enumerate(urls)))
with (OUTPUT / "report.csv").open("w", newline="") as handle:
writer = csv.DictWriter(handle, fieldnames=["index", "url", "status", "file", "attempts", "ms", "error"])
writer.writeheader()
writer.writerows(results)
print(f"Completed {sum(r['status'] == 'ok' for r in results)}/{len(urls)}")
Run it with python bulk_screenshots.py. This simple version launches one browser per worker task. For large batches, keep a browser alive and create a context or page per job so startup overhead does not dominate; also measure memory before increasing concurrency.
5. Add readiness controls
Wait for a selector
A selector is usually more reliable than sleeping when the page exposes a clear ready state:
await page.locator('.report-chart').waitFor({ state: 'visible', timeout: 20000 });
Wait for a bounded delay
Use a delay only when the site has a known animation or delayed widget and no useful selector. Keep it bounded and record it in your configuration.
await page.waitForTimeout(1500);
Handle lazy-loaded images
Some pages load images only as they approach the viewport. A full-page screenshot does not guarantee every lazy asset has loaded. Scroll in increments, wait for image completion where possible, then capture:
await page.evaluate(async () => {
await new Promise(resolve => {
const step = 600;
const timer = setInterval(() => {
window.scrollBy(0, step);
if (window.innerHeight + window.scrollY >= document.body.scrollHeight) {
clearInterval(timer);
resolve();
}
}, 100);
});
});
await page.waitForTimeout(500);
await page.screenshot({ path: target, fullPage: true });
Adapt this to the page’s own loading behavior. A single fixed wait cannot cover every site.
6. Output options and browser state
Playwright’s screenshot API accepts parameters for image format, clip area, quality, and more. See the official API examples.
- PNG: lossless and suitable for visual diffs.
- JPEG: smaller files when quality loss is acceptable; set a quality value.
- WebP: compact output when your downstream tools support it.
- Viewport: set width, height, device scale factor, color scheme, and locale in the browser context.
- Authenticated pages: load a saved storage state or set cookies and headers before navigation. Treat those files as secrets.
- Animations: inject CSS to pause transitions if deterministic pixels matter.
- Privacy: local execution keeps page contents and screenshots in your environment, subject to your own logs, storage, and third-party resources.
7. Reliability checklist
- Record the URL, timestamp, viewport, browser version, wait strategy, status, and error.
- Use bounded navigation and selector timeouts so one page cannot block the batch forever.
- Retry transient navigation failures with backoff, but do not retry invalid URLs indefinitely.
- Write to a temporary filename and rename after a successful capture to avoid partial files.
- Keep concurrency low enough that your machine does not run out of memory and target sites are not overloaded.
- Preserve failed URLs in a report for a second, targeted run.
- For visual comparisons, keep browser version, fonts, viewport, timezone, locale, and device scale factor stable.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation timeout | The page or a dependency never finished loading. | Use a bounded timeout, capture after domcontentloaded, wait for a specific selector, and record the failure. |
| Blank or incomplete image | Client-side rendering or lazy content was not ready. | Wait for the application’s ready selector, scroll to trigger lazy loading, and verify image completion. |
| Cookie banner covers content | The page requires consent before showing its normal layout. | Use a locator to click the appropriate consent control, or supply the site’s consent cookie in the browser context. |
| CAPTCHA or bot check | The target is challenging automated browsing. | Do not attempt to bypass access controls. Mark the URL failed or use an authorized capture path. |
| Only part of the page is captured | Viewport capture was used or the document height changed during capture. | Set fullPage: true, stabilize dynamic content, and compare the resulting dimensions. |
| Out of memory | Too many concurrent pages, very large documents, or unclosed contexts. | Lower concurrency, reuse a browser, close pages and contexts, and split the input into batches. |
| Missing fonts or different wrapping | Fonts are unavailable or the environment differs. | Install required fonts, wait for document.fonts.ready, and pin the capture environment. |
| Rate limiting | The target site limits request volume. | Lower concurrency, add delays, identify your traffic where appropriate, and follow the site’s terms. |
9. Performance, cost, and scale
Batch duration is driven by navigation time, readiness waits, screenshot size, browser startup, and concurrency. Measure your own pages rather than assuming a universal throughput number. Reusing one browser and creating isolated contexts can reduce startup work, while too many simultaneous pages increase memory pressure.
Local automation has no per-shot vendor charge, but you pay for compute, storage, browser maintenance, and engineering time. Hosted services trade that operational work for a current service price and limits. Compare privacy requirements, URL volume, failure handling, output formats, and retention before choosing.
10. Or skip the browser setup
ScreenshotNeo accepts one GET request per URL and also supports bulk capture of up to 100 URLs per call. It removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server so Claude, Cursor, and other MCP clients can use take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for the full option set. A single request looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
It supports full-page and element captures, dark mode, device presets or custom viewports, retina scale, PDF output, custom CSS and JavaScript, clicks, selector and network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify a migration.
The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.
11. FAQ
Should I capture the viewport or the full page?
Use the viewport for a realistic above-the-fold image. Use full page for archives, audits, and visual comparisons of the entire document.
How many URLs should I process at once?
There is no universal ideal. Start with a small concurrency value, observe memory and target-site responses, then adjust.
Can screenshots be returned without files?
Yes. Playwright returns screenshot bytes when no path is supplied, so you can upload or transform them directly.
Why do two runs have different pixels?
Dynamic content, animations, ads, fonts, timezones, viewport settings, and browser versions can all change output. Stabilize those inputs and record them with each run.
When is a hosted service preferable?
Choose one when you want URL-list or bulk handling without maintaining browser workers, and when its privacy, limits, output, and failure reporting fit your workload.


