How to capture website thumbnails for a list of URLs with Puppeteer
Capture a thumbnail for every URL with Puppeteer. Learn safe filenames, page readiness, output options, failure handling, and when to use an API instead.
Use Puppeteer to launch one browser, visit each URL, and save a screenshot to a unique file path. For a first-screen thumbnail, use the default viewport capture; set fullPage: true only when you want the whole document. The loop, filename mapping, and per-URL error handling are your application logic; Puppeteer provides the page navigation and screenshot operations.
The example below uses Node.js, Puppeteer, and a newline-delimited urls.txt. It creates a safe, unique filename for every input URL, records a manifest, continues after individual failures, and closes the browser even when the batch encounters an error. See the Puppeteer Screenshots guide and Page.screenshot() API for the documented screenshot behavior and options.
1. Set up a URL list
Create a project and install Puppeteer:
mkdir thumbnail-batch
cd thumbnail-batch
npm init -y
npm install puppeteer
Puppeteer downloads a compatible browser as part of its standard installation. If your environment instead uses puppeteer-core, provide the path to an installed browser with executablePath; the code below uses the standard puppeteer package.
Create urls.txt, with one absolute HTTP or HTTPS URL per line:
https://example.com
https://www.iana.org/domains/reserved
https://developer.mozilla.org/
Only feed the script URLs you are authorized to visit. If the list comes from users or another untrusted source, apply an allowlist or other destination policy appropriate to your environment before navigating; blindly visiting supplied URLs can expose the machine running the browser to unintended internal or external destinations.
2. Capture one thumbnail per URL
Save this as capture.mjs and run it with node capture.mjs. The script uses a 1280 by 800 viewport, waits for the load event, and applies a 30-second navigation timeout. Those values are explicit starting choices, not guarantees that every site will finish rendering within that time.
import puppeteer from 'puppeteer';
import { createHash } from 'node:crypto';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
const inputPath = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'thumbnails';
const viewport = { width: 1280, height: 800, deviceScaleFactor: 1 };
const navigationTimeoutMs = 30_000;
function parseUrl(line, lineNumber) {
let url;
try {
url = new URL(line);
} catch {
throw new Error(`Line ${lineNumber}: not a valid absolute URL: ${line}`);
}
if (url.protocol !== 'http:' && url.protocol !== 'https:') {
throw new Error(`Line ${lineNumber}: only http and https URLs are accepted: ${line}`);
}
return url;
}
function outputName(url, index) {
const host = url.hostname.toLowerCase().replace(/[^a-z0-9.-]+/g, '-').slice(0, 70) || 'site';
const pathPart = url.pathname.split('/').filter(Boolean).slice(0, 2).join('-')
.toLowerCase().replace(/[^a-z0-9-]+/g, '-').replace(/-+/g, '-').replace(/^-|-$/g, '').slice(0, 35);
const digest = createHash('sha256').update(url.href).digest('hex').slice(0, 10);
return `${String(index).padStart(4, '0')}-${host}${pathPart ? `-${pathPart}` : ''}-${digest}.png`;
}
const lines = (await readFile(inputPath, 'utf8'))
.split(/\r?\n/)
.map((value) => value.trim())
.filter((value) => value.length > 0 && !value.startsWith('#'));
const jobs = lines.map((line, index) => ({ url: parseUrl(line, index + 1), index: index + 1 }));
await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
const results = [];
try {
const page = await browser.newPage();
await page.setViewport(viewport);
page.setDefaultNavigationTimeout(navigationTimeoutMs);
for (const job of jobs) {
const filename = outputName(job.url, job.index);
const outputPath = path.join(outputDir, filename);
try {
const response = await page.goto(job.url.href, { waitUntil: 'load' });
// A non-2xx response can still have a useful page to capture. Record its status.
await page.screenshot({ path: outputPath, type: 'png' });
results.push({ input: job.url.href, file: filename, status: 'captured', httpStatus: response?.status() ?? null });
console.log(`Captured ${job.url.href} -> ${outputPath}`);
} catch (error) {
results.push({ input: job.url.href, file: filename, status: 'failed', error: String(error) });
console.error(`Failed ${job.url.href}: ${error.message}`);
}
}
} finally {
await browser.close();
await writeFile(path.join(outputDir, 'manifest.json'), JSON.stringify(results, null, 2) + '\n');
}
const failures = results.filter((result) => result.status === 'failed').length;
console.log(`Finished: ${results.length - failures} captured, ${failures} failed. Manifest: ${path.join(outputDir, 'manifest.json')}`);
if (failures > 0) process.exitCode = 1;
The output name combines the input position, a readable host/path fragment, and a short SHA-256 digest of the full URL. The digest helps distinguish URLs that share a host and path but differ in query string or fragment. The index also keeps repeated identical URLs as separate output rows. The manifest preserves the original URL to saved-file mapping and records failures and HTTP status.
Run the script:
node capture.mjs
# Or choose an input and output directory:
node capture.mjs urls.txt output-shots
3. Choose what counts as ready
The screenshot guide demonstrates navigation with waitUntil: 'networkidle2'. This is one option, not a universal readiness guarantee. A site with long-polling or persistent network requests may not become idle; meanwhile, client-rendered content or lazy images may still be absent when network activity pauses.
| Signal | Use it when | Tradeoff |
|---|---|---|
domcontentloaded |
The initial document structure is sufficient. | Images and later scripts may not have completed. |
load |
You want the page load event, including ordinary dependent resources. | It does not prove that a client app or delayed content is ready. |
networkidle0 or networkidle2 |
A quiet network is a useful proxy for your target pages. | Persistent traffic can delay or prevent the condition; quiet does not mean visually complete. |
| Selector wait | A known page element marks the content you need. | Requires a selector that exists and an appropriate timeout. |
| Bounded delay | A known site needs a short extra render interval. | Can waste time on fast pages and still be too short on slow ones. |
For a page-specific readiness selector, use this pattern after navigation and before the screenshot:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.waitForSelector('main article', { timeout: 10_000 });
await page.screenshot({ path: outputPath });
Replace main article with an element that actually signals that your target content has rendered. If no stable selector exists, a bounded delay is possible with await new Promise(resolve => setTimeout(resolve, 1500)); treat it as a site-specific fallback, not proof of completion.
4. Configure thumbnail size and format
For an at-a-glance thumbnail, the default capture is the viewport. Set the viewport before navigation so responsive layouts render at the intended size. deviceScaleFactor controls pixel density and can increase output dimensions and file size.
await page.setViewport({ width: 1200, height: 750, deviceScaleFactor: 1 });
await page.goto(url, { waitUntil: 'load' });
await page.screenshot({ path: outputPath, type: 'webp', quality: 80 });
| Option | What it changes | Notes |
|---|---|---|
path |
Writes the image to a file. | Without a path, Page.screenshot() returns image bytes instead. |
type |
Image format: PNG, JPEG, or WebP where supported by the installed Puppeteer/browser combination. | PNG is lossless; JPEG and WebP are lossy formats. |
quality |
Lossy image quality. | Use with supported lossy types; it does not improve PNG. |
fullPage |
Captures the full document instead of just the viewport. | Defaults to false. Full-page output is usually taller and larger than a thumbnail. |
clip |
Captures a specified rectangular region. | Useful for a known region; coordinates and dimensions must fit the page capture. |
omitBackground |
Omits the default white background where transparency is supported. | Useful for transparent output, typically with PNG. |
Example full-page archive image:
await page.screenshot({ path: outputPath, fullPage: true, type: 'png' });
For a single component rather than a page thumbnail, select its element and call ElementHandle.screenshot(). Puppeteer’s screenshot guide says this method attempts to scroll the element into view when it is hidden:
const card = await page.$('.product-card');
if (!card) throw new Error('Product card not found');
await card.screenshot({ path: outputPath, type: 'png' });
5. Process more URLs without losing reliability
The example reuses one page sequentially. This is a conservative starting point for a batch because it limits simultaneous browser work and makes per-URL failures easy to associate with inputs. It may take longer than parallel work, but the Puppeteer documentation cited here does not establish a universally safe concurrency setting.
- Reuse one browser: launching once avoids repeatedly starting a browser for every URL.
- Use one page sequentially for a small batch: the simplest resource and failure model.
- Use a bounded worker pool for larger batches: give each active worker its own page, cap the number of concurrent navigations, and measure memory and stability in your environment.
- Keep a manifest: persist input URL, file, status, response code, and error so the batch can be audited or retried.
- Retry selectively: retry transient navigation timeouts or temporary network failures a small, bounded number of times; do not retry malformed URLs or permanent access denials indefinitely.
- Close resources: close pages when workers finish and close the browser in a
finallyblock.
A worker pool should never create one browser per URL. Browser processes and pages consume memory and other system resources; increase concurrency gradually and watch for navigation failures and memory pressure. There is no evidence-backed throughput number or safe parallelism value that applies to every site and machine.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser launch fails | Browser download is missing, runtime libraries are unavailable, or a sandbox restriction blocks launch. | Confirm Puppeteer installed its browser, use the deployment environment’s documented launch requirements, or point puppeteer-core at an installed browser. |
| Navigation timeout | The host is slow, unreachable, or never reaches the selected readiness condition. | Check the URL and connectivity; select a readiness signal that fits the page, set a bounded timeout, and record the failure for retry or review. |
| Screenshot is blank or incomplete | The page captured before app rendering, fonts, images, or lazy content appeared. | Wait for a relevant selector; use an appropriate load signal or a bounded delay. For full-page captures, note that below-the-fold lazy content may need scrolling or additional page-specific handling. |
| Some URLs overwrite the same file | Filename generation used only a host or raw path. | Include a stable unique component such as an index and a digest of the complete URL, as in the example. |
| Invalid path or file write error | Raw URL characters were used in a filename, or the output directory is not writable. | Sanitize names, create the output directory, and check filesystem permissions and available storage. |
| HTTP error page appears as an image | Navigation completed with an HTTP error status; that does not necessarily throw as a navigation exception. | Inspect the recorded response status and decide whether to keep, flag, or exclude that screenshot. |
| Batch stops after one bad site | Failure handling surrounds the entire loop instead of each URL. | Catch errors per job, append a failed manifest record, then continue to the next URL. |
| Very tall or unexpectedly large files | fullPage was enabled or the viewport/device scale is larger than intended. |
Use viewport capture for thumbnails; reduce viewport dimensions or device scale, or choose an appropriate lossy format and quality. |
7. Cost, performance, and operational notes
Puppeteer is a browser automation library, so a self-hosted batch consumes the compute, memory, storage, and network capacity of the machine running it. The runtime depends on the pages, readiness condition, browser environment, and concurrency; the available documentation does not establish a benchmark for this URL-list workflow. Keep timeouts bounded, avoid unbounded parallelism, and monitor output storage.
For repeat runs, a manifest lets you retry only failures. If your source URLs are stable and screenshots do not need to reflect every visit, you can also maintain your own cache keyed by URL and capture settings. Decide explicitly whether query strings and fragments should distinguish outputs: the example hashes the full URL, so they do.
For current Puppeteer screenshot option behavior, consult the official screenshot guide and ScreenshotOptions API. API details can vary with the package version installed in your project.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Send a GET request with a URL to receive an image or PDF; see the ScreenshotNeo API documentation for configuration and response details. This cURL example saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.
FAQ
Does Puppeteer have a built-in batch screenshot method?
The screenshot guide documents page-level operations. A URL list is handled by your own loop or worker pool.
How do I keep the filenames linked to the input URLs?
Write a manifest alongside the images with each original URL, output filename, status, and any error. The example creates manifest.json.
Should I use screenshots of the viewport or the full page?
Use the viewport for a compact preview of the initial view. Use fullPage: true when the goal is to preserve the whole document.
Can I capture an element instead of the whole page?
Yes. Use ElementHandle.screenshot() on the selected element when the desired image is a component preview.


