ScreenshotNeo

BlogHow-to

How to Create Website Thumbnails from a List of URLs

Turn a URL list into consistently sized website thumbnails with Playwright. Get runnable batch code, failure handling, output options, and a hosted API alternative.

By the ScreenshotNeo team4 October 202612 min read

To create website thumbnails from a list of URLs, use a browser automation script to open each URL, capture a screenshot, and save it under a unique filename. The example below uses Node.js and Playwright: it reads a text file, captures a consistent viewport preview for each valid URL, records success or failure in a JSON manifest, and continues when an individual site fails.

A viewport capture is usually the right starting point for thumbnail cards. Use full-page capture when you need the entire scrollable page; those images can be much taller and less uniform. Playwright documents navigation and screenshot capture in its Screenshots guide and Page API.

1. Set up a local batch thumbnail script

Install Node.js, create a project, and install Playwright and its Chromium browser:

mkdir url-thumbnails
cd url-thumbnails
npm init -y
npm install playwright
npx playwright install chromium

Create urls.txt with one address per line. Blank lines and lines beginning with # are ignored.

https://example.com
https://stripe.com
# Add one URL per line

Save this as thumbnails.mjs. It is a complete runnable script. The default settings make fixed-size viewport previews, use a CSS-pixel screenshot scale, save WebP files, run two pages at once, and retry transient failures once.

import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { createHash } from 'node:crypto';

const INPUT = process.argv[2] ?? 'urls.txt';
const OUTPUT_DIR = process.argv[3] ?? 'thumbnails';
const WIDTH = 1280;
const HEIGHT = 800;
const CONCURRENCY = 2;
const RETRIES = 1;
const NAVIGATION_TIMEOUT_MS = 30_000;
const SCREENSHOT_TIMEOUT_MS = 15_000;
const FORMAT = 'webp'; // png, jpeg, or webp
const QUALITY = 82; // used for jpeg/webp; range 0–100
const FULL_PAGE = false;
const WAIT_UNTIL = 'domcontentloaded'; // load, domcontentloaded, or commit

function parseUrl(line) {
  const input = line.trim();
  if (!input || input.startsWith('#')) return null;

  // Only infer HTTPS for an unambiguous bare hostname. Other inputs need a scheme.
  const candidate = /^[\w.-]+\.[a-z]{2,}(?::\d+)?(?:\/|$)/i.test(input)
    ? `https://${input}`
    : input;
  const url = new URL(candidate);
  if (!['http:', 'https:'].includes(url.protocol)) {
    throw new Error(`Unsupported URL scheme: ${url.protocol}`);
  }
  return url.toString();
}

function filenameFor(url) {
  const parsed = new URL(url);
  const readable = `${parsed.hostname}${parsed.pathname}`
    .replace(/[^a-z0-9._-]+/gi, '-')
    .replace(/^-+|-+$/g, '')
    .slice(0, 80) || 'page';
  const hash = createHash('sha256').update(url).digest('hex').slice(0, 10);
  return `${readable}-${hash}.${FORMAT}`;
}

async function captureOne(browser, url) {
  const filename = filenameFor(url);
  const outputPath = path.join(OUTPUT_DIR, filename);
  let lastError;

  for (let attempt = 0; attempt <= RETRIES; attempt++) {
    const context = await browser.newContext({
      viewport: { width: WIDTH, height: HEIGHT },
      deviceScaleFactor: 1,
    });
    const page = await context.newPage();
    try {
      const response = await page.goto(url, {
        waitUntil: WAIT_UNTIL,
        timeout: NAVIGATION_TIMEOUT_MS,
      });
      // A 4xx/5xx response still renders a page; preserve it and record its status.
      await page.screenshot({
        path: outputPath,
        type: FORMAT,
        quality: FORMAT === 'png' ? undefined : QUALITY,
        fullPage: FULL_PAGE,
        scale: 'css',
        animations: 'disabled',
        timeout: SCREENSHOT_TIMEOUT_MS,
      });
      await context.close();
      return {
        url,
        status: 'ok',
        httpStatus: response?.status() ?? null,
        file: outputPath,
        attempts: attempt + 1,
      };
    } catch (error) {
      lastError = error;
      await context.close().catch(() => {});
      if (attempt < RETRIES) {
        await new Promise(resolve => setTimeout(resolve, 500 * (attempt + 1)));
      }
    }
  }

  return {
    url,
    status: 'error',
    file: null,
    attempts: RETRIES + 1,
    error: String(lastError?.message ?? lastError),
  };
}

async function main() {
  const lines = (await readFile(INPUT, 'utf8')).split(/\r?\n/);
  const entries = [];
  for (const [index, line] of lines.entries()) {
    try {
      const url = parseUrl(line);
      if (url) entries.push({ line: index + 1, url });
    } catch (error) {
      entries.push({ line: index + 1, invalid: true, input: line, error: error.message });
    }
  }
  await mkdir(OUTPUT_DIR, { recursive: true });
  const browser = await chromium.launch({ headless: true });
  const results = new Array(entries.length);
  let next = 0;

  async function worker() {
    while (true) {
      const index = next++;
      if (index >= entries.length) return;
      const entry = entries[index];
      results[index] = entry.invalid
        ? { line: entry.line, input: entry.input, status: 'invalid', error: entry.error }
        : { line: entry.line, ...(await captureOne(browser, entry.url)) };
    }
  }

  try {
    await Promise.all(Array.from(
      { length: Math.min(CONCURRENCY, entries.length) },
      () => worker(),
    ));
  } finally {
    await browser.close();
  }

  const manifestPath = path.join(OUTPUT_DIR, 'manifest.json');
  await writeFile(manifestPath, JSON.stringify(results, null, 2) + '\n');
  const succeeded = results.filter(item => item.status === 'ok').length;
  console.log(`Saved ${succeeded}/${results.length} captures. Manifest: ${manifestPath}`);
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it with node thumbnails.mjs, or pass different paths as arguments: node thumbnails.mjs input.txt output-folder. Each image filename combines a readable hostname/path with a short hash of the full URL, which reduces collisions when URLs share a path or have query strings. The manifest preserves the original normalized URL, output path, response status, retry count, and any failure.

2. Choose what the thumbnail should show

Viewport or full page

For a grid of preview cards, keep FULL_PAGE = false and choose a shared viewport such as 1280 by 800 CSS pixels. This captures the initial visible area and keeps aspect ratios predictable. For an archive or review, set FULL_PAGE = true; Playwright describes this as capturing the full scrollable page. Full-page output dimensions vary by site and may be very tall. Lazy-loaded images below the fold may need an explicit scroll-and-wait step before capture.

Format, quality, and scale

  • PNG: lossless and often larger. Do not set quality for PNG.
  • JPEG: useful for photographic pages and broadly compatible; set quality from 0 to 100.
  • WebP: often a practical choice for web thumbnail storage; set quality from 0 to 100.
  • CSS scale: one output pixel per CSS pixel, which keeps high-DPI captures smaller and dimensions consistent with the viewport.
  • Device scale: includes device pixel ratio, so the resulting image can be larger and sharper on high-density displays.

These screenshot controls are documented in the Playwright Page API. If a downstream card requires exact dimensions, resize or crop the captured file to a fixed canvas in an image-processing step. Viewport dimensions alone do not guarantee that the source content has the same composition on every site.

Readiness and dynamic pages

The sample waits for domcontentloaded, which avoids waiting for every image and third-party resource to finish. Change WAIT_UNTIL to load when the page needs its load event before capture, or commit when you want to start as soon as the response begins. Playwright documents networkidle but discourages using it as a general readiness signal; pages with analytics, streaming, or recurring requests may never become idle. Prefer waiting for a known selector when the content matters:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('main h1').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: outputPath, type: 'webp', quality: 82, scale: 'css' });

For a site that fills content only after scrolling, scroll in steps and allow the page to render before taking a full-page image. A short fixed delay can help diagnose a rendering race, but a selector or another site-specific readiness condition is more reliable than an arbitrary delay.

Capture one element

To make a thumbnail from a card, hero, or other specific component, wait for its selector and capture the locator instead of the whole page:

const card = page.locator('main');
await card.waitFor({ state: 'visible', timeout: 10_000 });
await card.screenshot({ path: outputPath, type: 'webp', quality: 82 });

The element screenshot is clipped to that element’s size and position. See the ElementHandle API for the documented element capture behavior.

3. Process larger lists safely

The loop, manifest, retries, filename scheme, and concurrency limit in the example are batch-processing choices built around Playwright’s page navigation and screenshot operations. Playwright does not provide a turnkey URL-list batch command or a universal recommended concurrency level.

  1. Start with low concurrency. The example uses two pages at once. Increase gradually while watching memory, CPU, target-site rate limits, and failure rate. Large pages and full-page captures consume more resources.
  2. Reuse the browser. The script launches Chromium once and creates an isolated context for each attempt. Closing each context releases page state while avoiding a fresh browser process for every URL.
  3. Keep failures per URL. The script catches navigation and screenshot errors individually, retries once, records the final error, and continues the queue. Tune retry count and backoff to the target sites; retries cannot fix a persistent block or invalid address.
  4. Persist a manifest. Use the JSON mapping to resume only failed or missing captures, audit outputs, and associate each image with its input. For a very large or interruptible job, write results incrementally rather than only at the end.
  5. Normalize carefully. The example adds HTTPS only to clear bare hostnames. It rejects non-HTTP schemes. Preserve meaningful query strings because they can identify distinct pages; the hash in the filename distinguishes them.
  6. Resize after capture when required. A screenshot captures browser pixels; it does not automatically create a uniform card crop. Apply a consistent crop/fit policy downstream and decide whether to letterbox, center-crop, or preserve the full image.

If URLs come from untrusted users, treat the capture worker as a network boundary. Allow only intended public HTTP/HTTPS destinations and prevent access to loopback, private, link-local, and internal network addresses. Recheck resolved addresses and redirects; validating only the original hostname is insufficient for a robust server-side capture service. Run workers with limited network permissions and avoid forwarding secrets or privileged cookies to arbitrary destinations.

4. Python alternative with Playwright

If the rest of your pipeline is Python, install Playwright and its Chromium browser:

python -m pip install playwright
python -m playwright install chromium

Save as thumbnails.py. This version handles each URL independently and produces a JSON manifest. It uses a sequential loop for simplicity; add a bounded worker pool only after measuring resource use.

import asyncio
import hashlib
import json
import re
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright

INPUT = Path('urls.txt')
OUT = Path('thumbnails')
WIDTH, HEIGHT = 1280, 800


def normalize(raw):
    value = raw.strip()
    if not value or value.startswith('#'):
        return None
    if re.match(r'^[\\w.-]+\\.[a-z]{2,}(?::\\d+)?(?:/|$)', value, re.I):
        value = 'https://' + value
    parsed = urlparse(value)
    if parsed.scheme not in ('http', 'https') or not parsed.netloc:
        raise ValueError('Expected an http:// or https:// URL')
    return value


def output_name(url):
    parsed = urlparse(url)
    readable = re.sub(r'[^a-zA-Z0-9._-]+', '-', parsed.netloc + parsed.path).strip('-')[:80]
    suffix = hashlib.sha256(url.encode()).hexdigest()[:10]
    return f'{readable or "page"}-{suffix}.webp'


async def main():
    OUT.mkdir(parents=True, exist_ok=True)
    entries = []
    for line_no, line in enumerate(INPUT.read_text(encoding='utf-8').splitlines(), 1):
        try:
            url = normalize(line)
            if url:
                entries.append({'line': line_no, 'url': url})
        except Exception as exc:
            entries.append({'line': line_no, 'status': 'invalid', 'input': line, 'error': str(exc)})

    results = []
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            for entry in entries:
                if 'url' not in entry:
                    results.append(entry)
                    continue
                context = await browser.new_context(viewport={'width': WIDTH, 'height': HEIGHT}, device_scale_factor=1)
                page = await context.new_page()
                try:
                    response = await page.goto(entry['url'], wait_until='domcontentloaded', timeout=30_000)
                    name = output_name(entry['url'])
                    await page.screenshot(path=str(OUT / name), type='webp', quality=82, scale='css', timeout=15_000)
                    results.append({**entry, 'status': 'ok', 'httpStatus': response.status if response else None, 'file': str(OUT / name)})
                except Exception as exc:
                    results.append({**entry, 'status': 'error', 'error': str(exc)})
                finally:
                    await context.close()
        finally:
            await browser.close()

    (OUT / 'manifest.json').write_text(json.dumps(results, indent=2) + '\n', encoding='utf-8')
    print(f"Saved {sum(x.get('status') == 'ok' for x in results)}/{len(results)} captures")


asyncio.run(main())

Run python thumbnails.py. As with the Node.js version, the screenshot API exposes image options while the input processing, naming, and failure policy are application code.

5. Troubleshooting

Symptom Likely cause Fix
Browser executable missing Playwright package installed but its browser was not downloaded. Run npx playwright install chromium or python -m playwright install chromium.
Navigation timeout Slow site, stalled resource, blocked request, or an overly strict wait condition. Use a reasonable per-page timeout; try domcontentloaded or commit, and wait for the specific content selector separately. Record failures rather than hanging the whole batch.
Screenshot is blank or missing content Capture occurred before client-rendered content appeared, or a consent wall/login is blocking the page. Wait for a visible content selector. Check the resulting page and response status. The local script does not automatically remove consent banners or bypass access controls.
Some pages show an error page The server returned an HTTP error, redirected, denied automation, or displayed a challenge. Inspect the manifest status and page manually. A completed navigation can still return a 4xx/5xx response; the sample saves the rendered result and records its HTTP status.
WebP/JPEG quality option errors Quality was set for PNG, or the selected format/extension does not match. Set the screenshot type and filename extension consistently. Omit quality for PNG.
Output files overwrite each other Names were based only on a host or sequential counter reused across runs. Include the path and a stable hash of the full normalized URL, as the sample does. Keep the manifest as the authoritative mapping.
Process runs out of memory Too many simultaneous pages, very large full-page captures, or contexts not being closed. Lower concurrency, use viewport captures where possible, close each context, and resize images after capture.
Lazy images are absent Images load only when their section approaches the viewport. For full-page capture, scroll down in increments and wait for expected images/content to load before screenshotting.
URL rejected by parser Missing scheme on a non-obvious input, malformed hostname, or unsupported protocol. Enter a full https:// or http:// URL. The sample only infers HTTPS for a recognizable bare hostname.

6. Performance, reliability, and cost

Local Playwright has no per-screenshot API charge, but it uses your compute, storage, bandwidth, and engineering time. Browser installation and maintenance, queueing, retries, monitoring, image storage, and safe network isolation become part of the job. Hosted capture can reduce the browser operations you maintain, but compare current price, concurrency, rendering controls, privacy and retention terms, and failure semantics against your workload before choosing a provider.

Capture time depends on the sites and their assets, so there is no universal reliable batch size or throughput figure. Keep concurrency bounded, set navigation and screenshot timeouts, save status for every URL, and retry only failures likely to be transient. For repeat runs, skip captures whose inputs have not changed if freshness requirements allow it. Screenshot file size depends on image format, quality, scale, viewport, and page content; full-page screenshots can be substantially larger than viewport thumbnails.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One request can turn a URL into an image, so you do not need to install and operate a browser for this capture. The endpoint accepts the parameter names other screenshot APIs use, which can make switching easier. See the ScreenshotNeo API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For a list, call the endpoint once per URL and save each returned image with your own unique filename and manifest. ScreenshotNeo’s bulk capture supports up to 100 URLs per call. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Can I generate thumbnails without saving them to disk?

Yes. Playwright’s screenshot call can return image bytes when you omit the path. You can pass those bytes to an image processor or upload them to storage; keep a manifest that records the destination and result.

Should I capture the full page for a thumbnail?

Usually not for a compact preview card. A fixed viewport gives more predictable dimensions. Choose full-page capture when the purpose is to inspect or preserve the whole page.

Can the batch continue if one URL fails?

Yes. Catch errors around each URL, store its failure in the manifest, and continue with the rest. The examples follow that pattern.

How do I make thumbnails visually consistent across sites?

Use the same viewport, color scheme, scale, format, and crop policy. Sites still control their own responsive layouts, so resize or crop the final output when exact canvas dimensions are required.