How to Create Thumbnail Images for Hundreds of URLs with Playwright
Build a reliable Playwright batch that captures, resizes, and tracks thumbnails for hundreds of URLs, with bounded concurrency and retryable results.
Use Playwright to visit each URL, wait for the page content you need, capture a viewport or selected element, and save the image under a stable filename. For hundreds of URLs, reuse a browser, limit concurrent pages, record each result in a manifest, and retry failures deliberately. Playwright does not choose your thumbnail dimensions: capture first, then crop or resize to your destination’s specification.
This guide uses Node.js and Playwright. It covers a runnable batch script, readiness choices, output sizing, failure handling, concurrency, and an API alternative.
1. Install Playwright and prepare a URL list
Create a project and install Playwright. Install its Chromium browser once on the machine that will run the job:
npm init -y
npm install playwright
npx playwright install chromium
Save one URL per line in urls.txt. Blank lines and lines beginning with # are ignored by the example:
https://example.com/
https://www.wikipedia.org/
# Add the rest of your URLs here
Run the script below with Node.js. It creates thumbnails/ and manifest.jsonl. Each manifest line records the input URL, final URL when available, HTTP status, output path, and any error. The index-based filenames are stable for an unchanged input order and avoid unsafe URL characters and filename collisions.
2. Runnable Node.js batch script
const fs = require('node:fs/promises');
const path = require('node:path');
const { chromium } = require('playwright');
const INPUT = process.env.URLS_FILE || 'urls.txt';
const OUTPUT_DIR = process.env.OUTPUT_DIR || 'thumbnails';
const MANIFEST = process.env.MANIFEST || 'manifest.jsonl';
const CONCURRENCY = Math.max(1, Number(process.env.CONCURRENCY || 4));
const NAV_TIMEOUT_MS = Math.max(1, Number(process.env.NAV_TIMEOUT_MS || 30000));
const READY_SELECTOR = process.env.READY_SELECTOR || '';
const FULL_PAGE = process.env.FULL_PAGE === '1';
const CAPTURE_HTTP_ERRORS = process.env.CAPTURE_HTTP_ERRORS === '1';
const MAX_ATTEMPTS = Math.max(1, Number(process.env.MAX_ATTEMPTS || 2));
function safeExtension(format) {
if (format === 'jpeg') return 'jpg';
if (format === 'png' || format === 'webp') return format;
throw new Error('IMAGE_FORMAT must be png, jpeg, or webp');
}
async function readUrls(file) {
const text = await fs.readFile(file, 'utf8');
return text.split(/\r?\n/)
.map(line => line.trim())
.filter(line => line && !line.startsWith('#'));
}
async function main() {
const urls = await readUrls(INPUT);
await fs.mkdir(OUTPUT_DIR, { recursive: true });
await fs.writeFile(MANIFEST, '', 'utf8');
const format = process.env.IMAGE_FORMAT || 'png';
const extension = safeExtension(format);
const browser = await chromium.launch({ headless: true });
let next = 0;
async function worker() {
while (true) {
const index = next++;
if (index >= urls.length) return;
const inputUrl = urls[index];
const outputPath = path.join(OUTPUT_DIR, `${String(index + 1).padStart(4, '0')}.${extension}`);
let record = {
index: index + 1,
inputUrl,
finalUrl: null,
status: null,
outputPath,
error: null,
attempts: 0
};
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
record.attempts = attempt;
let page;
try {
page = await browser.newPage({ viewport: { width: 1280, height: 800 } });
page.setDefaultNavigationTimeout(NAV_TIMEOUT_MS);
const response = await page.goto(inputUrl, { waitUntil: 'domcontentloaded' });
record.finalUrl = page.url();
record.status = response ? response.status() : null;
if (!CAPTURE_HTTP_ERRORS && response && response.status() >= 400) {
throw new Error(`HTTP ${response.status()}`);
}
if (READY_SELECTOR) {
await page.locator(READY_SELECTOR).waitFor({ state: 'visible', timeout: NAV_TIMEOUT_MS });
}
const imageType = format === 'jpeg' ? 'jpeg' : format;
await page.screenshot({
path: outputPath,
type: imageType,
fullPage: FULL_PAGE,
animations: 'disabled',
timeout: NAV_TIMEOUT_MS
});
record.error = null;
break;
} catch (error) {
record.error = error instanceof Error ? error.message : String(error);
if (attempt < MAX_ATTEMPTS) {
await new Promise(resolve => setTimeout(resolve, 500 * attempt));
}
} finally {
if (page) await page.close().catch(() => {});
}
}
await fs.appendFile(MANIFEST, `${JSON.stringify(record)}\n`, 'utf8');
console.log(`${record.error ? 'FAILED' : 'OK'} ${inputUrl}${record.error ? `: ${record.error}` : ''}`);
}
}
try {
await Promise.all(Array.from({ length: Math.min(CONCURRENCY, urls.length) }, () => worker()));
} finally {
await browser.close();
}
console.log(`Processed ${urls.length} URLs. Results: ${MANIFEST}`);
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
Save this as thumbnails.cjs. Start with four workers and adjust based on the machine and sites in your list; that value is a conservative starting point for this example, not a universal throughput recommendation. To run:
node thumbnails.cjs
Useful environment settings:
| Setting | Default | Effect |
|---|---|---|
URLS_FILE |
urls.txt |
Input file path. |
OUTPUT_DIR |
thumbnails |
Directory for saved files. |
MANIFEST |
manifest.jsonl |
One JSON result per input URL. |
CONCURRENCY |
4 |
Maximum active pages. Raise gradually while watching memory and site responses. |
NAV_TIMEOUT_MS |
30000 |
Navigation, readiness selector, and screenshot timeout in milliseconds. |
READY_SELECTOR |
empty | Optional CSS selector that must become visible before capture. |
FULL_PAGE |
0 |
Set to 1 to capture the full scrollable page. |
IMAGE_FORMAT |
png |
Choose png, jpeg, or webp. |
CAPTURE_HTTP_ERRORS |
0 |
Set to 1 to save HTTP 4xx/5xx pages as diagnostic images. |
MAX_ATTEMPTS |
2 |
Attempts per URL, including the first try. Retries wait briefly between attempts. |
3. Pick the right readiness signal
page.goto() supports commit, domcontentloaded, load, and networkidle for its waitUntil option. The script uses domcontentloaded so it does not wait for every image, analytics request, or long-lived connection before proceeding. If the thumbnail needs client-rendered content, wait for an observable site-specific condition, such as a card, heading, or product image becoming visible.
For example, if the content to capture appears in an element with #main-content, run:
READY_SELECTOR='#main-content' node thumbnails.cjs
Choose a selector that means the content is actually useful in the image. A generic page shell may appear before its data has loaded. For a site without a suitable selector, add a site-specific check or a bounded delay after navigation. Playwright documents networkidle as discouraged as a general testing readiness signal; network silence does not prove the right content is present. See the Playwright Page API.
4. Choose capture scope and image output
Viewport thumbnail
The default screenshot is the visible viewport, 1280 by 800 CSS pixels in this script. This is a reasonable starting capture for a page preview. Fix the viewport across the batch if you want comparable compositions. Pages may still render differently due to responsive breakpoints, fonts, animation, or browser state.
Full-page capture
Set FULL_PAGE=1 to capture the full scrollable document:
FULL_PAGE=1 node thumbnails.cjs
Full-page images can be very tall and use more memory and storage. They are useful when the entire document matters, but are often a poor direct fit for a compact card thumbnail. Consider resizing and cropping them to the destination aspect ratio.
Capture a single element
For a selected element, use a locator screenshot instead of the page screenshot:
await page.locator('main article').screenshot({ path: outputPath, type: 'png' });
Wait for that locator to be visible first. Element screenshots represent the element’s bounds, so check that the selector identifies one unambiguous region and that it is not clipped by overflow or overlays. Playwright documents page, full-page, buffer, and element screenshot options in its screenshots guide.
Format, quality, and fixed dimensions
PNG is lossless and often larger. JPEG is suitable for photographic content and accepts a quality option from 0 to 100. WebP can be useful where your publishing pipeline supports it. For example, change the screenshot options to include quality: 80 when using JPEG or WebP:
await page.screenshot({ path: outputPath, type: 'jpeg', quality: 80 });
Playwright captures the browser rendering; it does not impose your destination’s thumbnail size. If the target requires, for example, a fixed 320 by 180 pixel file, process the captured output with an image library or your image pipeline. Decide whether to crop, letterbox, or stretch; stretching distorts the page. Capture buffers with const bytes = await page.screenshot() when the next step needs image bytes rather than a file. Playwright’s screenshot documentation describes saving to a path and returning bytes.
5. Python alternative with Playwright
If your batch pipeline is in Python, install the Python package and Chromium, then use the same basic pattern: navigate, check the response, wait for useful content, and capture. This runnable script reads urls.txt, writes numbered PNG files, and records JSON Lines results.
pip install playwright
playwright install chromium
import asyncio
import json
import os
from pathlib import Path
from playwright.async_api import async_playwright
INPUT = Path(os.getenv('URLS_FILE', 'urls.txt'))
OUTPUT_DIR = Path(os.getenv('OUTPUT_DIR', 'thumbnails'))
MANIFEST = Path(os.getenv('MANIFEST', 'manifest.jsonl'))
CONCURRENCY = max(1, int(os.getenv('CONCURRENCY', '4')))
TIMEOUT_MS = max(1, int(os.getenv('NAV_TIMEOUT_MS', '30000')))
READY_SELECTOR = os.getenv('READY_SELECTOR', '')
FULL_PAGE = os.getenv('FULL_PAGE', '0') == '1'
CAPTURE_HTTP_ERRORS = os.getenv('CAPTURE_HTTP_ERRORS', '0') == '1'
MAX_ATTEMPTS = max(1, int(os.getenv('MAX_ATTEMPTS', '2')))
async def main():
urls = [line.strip() for line in INPUT.read_text(encoding='utf-8').splitlines()
if line.strip() and not line.strip().startswith('#')]
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
MANIFEST.write_text('', encoding='utf-8')
semaphore = asyncio.Semaphore(CONCURRENCY)
async with async_playwright() as p:
browser = await p.chromium.launch()
async def capture(index, input_url):
output_path = OUTPUT_DIR / f'{index + 1:04d}.png'
result = {'index': index + 1, 'inputUrl': input_url, 'finalUrl': None,
'status': None, 'outputPath': str(output_path),
'error': None, 'attempts': 0}
async with semaphore:
for attempt in range(1, MAX_ATTEMPTS + 1):
result['attempts'] = attempt
page = await browser.new_page(viewport={'width': 1280, 'height': 800})
page.set_default_navigation_timeout(TIMEOUT_MS)
try:
response = await page.goto(input_url, wait_until='domcontentloaded')
result['finalUrl'] = page.url
result['status'] = response.status if response else None
if (not CAPTURE_HTTP_ERRORS and response
and response.status >= 400):
raise RuntimeError(f'HTTP {response.status}')
if READY_SELECTOR:
await page.locator(READY_SELECTOR).wait_for(
state='visible', timeout=TIMEOUT_MS)
await page.screenshot(path=str(output_path), full_page=FULL_PAGE,
animations='disabled', timeout=TIMEOUT_MS)
result['error'] = None
break
except Exception as exc:
result['error'] = str(exc)
if attempt < MAX_ATTEMPTS:
await asyncio.sleep(0.5 * attempt)
finally:
await page.close()
with MANIFEST.open('a', encoding='utf-8') as manifest:
manifest.write(json.dumps(result, ensure_ascii=False) + '\n')
print(('FAILED' if result['error'] else 'OK'), input_url,
result['error'] or '')
try:
await asyncio.gather(*(capture(i, url) for i, url in enumerate(urls)))
finally:
await browser.close()
asyncio.run(main())
6. cURL and Node.js with ScreenshotNeo
If you want screenshots without installing and managing a local browser, ScreenshotNeo is a website screenshot API and MCP server. Its API accepts one URL per request and returns an image or PDF. For hundreds of URLs, call it once per URL with your own bounded request concurrency and retain the same input-to-output manifest idea. See the ScreenshotNeo API documentation for request options.
cURL: one URL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python: one URL
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js: one URL
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
The API accepts many screenshot options, including viewport and device presets, full-page capture, element selection, output format, custom CSS and JavaScript, wait conditions, request blocking, cookies and headers, caching, and asynchronous jobs. It also supports bulk capture of up to 100 URLs per call. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Cookie and consent banners, newsletter popups, and chat widgets can be removed before capture, with each cleanup step configurable. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. Every feature is available on every plan; 1,000 shots a month are free with no card, and paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card required.
7. Reliability, performance, and cost considerations
- Bound concurrency. Each active page consumes browser and machine resources, and sites may throttle bursts. Increase
CONCURRENCYgradually while observing memory, failures, and response behavior. There is no universal worker count or runtime for hundreds of arbitrary URLs. - Reuse the browser. The example launches Chromium once and creates a fresh page for each attempt. This avoids browser startup for every URL while isolating page state. If sites need shared login state, use an explicitly configured browser context and protect its cookies.
- Keep partial results. The manifest is appended as workers finish, so successful captures remain recorded if a later URL fails. For large production runs, consider writing to a temporary manifest and atomically finalizing it, or include a run identifier.
- Retry selectively. The example retries every caught error up to a small limit. A production pipeline should distinguish transient timeouts from permanent invalid URLs, denied access, and HTTP errors. Use backoff and a retry cap to avoid hammering a struggling site.
- Plan storage and image size. Full-page images can be much larger than viewport images. JPEG or WebP may reduce storage where supported; choose quality according to visual needs. Preserve originals if a later crop or reprocessing step matters.
- Respect access boundaries. Capture only pages you are permitted to access. Authenticated sites may require session setup; private pages should not be exposed in public output directories or logs.
8. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation times out | The site is slow, has long-running requests, or is unreachable. | Use an appropriate timeout, choose domcontentloaded or commit when suitable, and add a site-specific readiness check. Retry only when the failure may be transient. |
| A 404 or 500 page was captured | HTTP error responses do not necessarily make page.goto() throw. |
Inspect the response status. The script marks status 400 or higher as failed by default; set CAPTURE_HTTP_ERRORS=1 when you want diagnostic captures. See the Page API. |
| Thumbnail is blank or missing content | Capture began before client-rendered content appeared, or the selected selector is too generic. | Set READY_SELECTOR to a visible content element or implement a site-specific readiness check. Verify the selector against the page. |
| Images or fonts are incomplete | They load after DOM content is ready, or they are lazy-loaded below the viewport. | Wait for a relevant image or selector, or use a deliberate bounded wait. Full-page capture may expose lazy content, but behavior depends on the page; inspect representative output. |
| Browser launch fails | The Playwright package or browser binary is missing for the runtime environment. | Run npx playwright install chromium. In Python, run playwright install chromium. |
| Some output files are missing | Those entries exhausted retries or failed before screenshot save. | Check the manifest’s error and attempts values, correct the cause, then rerun only failed inputs. |
| Workers slow down or the process runs out of memory | Too many concurrent pages, oversized full-page captures, or constrained host resources. | Lower CONCURRENCY, use viewport captures, and process in smaller batches. |
| Output names overwrite each other | A different input list or order reused the same output directory. | Use a separate directory per run, or derive stable IDs from a canonicalized URL and handle hash collisions. |
9. FAQ
Does Playwright automatically make a file at my required thumbnail dimensions?
No. It captures the browser viewport, full page, or a locator. Resize or crop the resulting image to the dimensions your destination requires.
Should I capture HTTP error pages?
That depends on the purpose. For a thumbnail set, usually record them as failures and retry or review them. For diagnostics, saving error pages can be useful.
Can I resume a batch after interruption?
The example records completed results, but it does not skip existing manifest entries automatically. Add a resume pass that loads completed input IDs and schedules only missing or failed records.
Is there a universal best concurrency setting?
No. It depends on the host, page weight, and the sites being visited. Start low, observe resource use and failures, and tune for your workload.
Sources
- Playwright Page API: navigation, response status, navigation errors, and readiness options.
- Playwright screenshots guide: viewport, full-page, and element screenshots.
- Playwright Python screenshots guide: Python screenshot usage.
- ScreenshotNeo documentation: API usage and options.


