ScreenshotNeo

BlogHow-to

How to Bulk Screenshot a List of URLs and Generate a Contact Sheet

Capture a URL list with Playwright, assemble labeled screenshots into a contact sheet with Python, and handle failures, sizing, and hosted capture options.

By the ScreenshotNeo team4 October 202611 min read

To bulk screenshot a list of URLs and generate a contact sheet, use two steps: capture one image per URL, then arrange those images into a labeled grid. Playwright can capture pages locally; Pillow can build the sheet. The example below uses consistent viewport screenshots so tiles are easy to compare. Use full-page screenshots when seeing the entire scrollable page matters more than uniform tile dimensions.

For a hosted route, ScreenshotNeo can capture URLs through one API call per URL and supports bulk capture of up to 100 URLs per call. It returns individual captures; the contact-sheet layout remains a separate step.

1. Choose a capture workflow

Route Best when Trade-offs
Playwright plus Pillow You want control over browser behavior, dimensions, file naming, and the sheet layout. You install and run a browser, and handle site-specific failures yourself.
ScreenshotNeo API You want hosted captures without managing browser installation, or need batch requests. You still need to assemble returned images into a sheet. Review the current API documentation for request limits and output details.
ScreenshotRun batch capture You want a hosted option documented for thumbnail capture from a URL list. The reviewed documentation describes batch screenshots, but does not establish built-in contact-sheet composition. Its docs say each URL counts as one screenshot toward monthly quota; check current plan terms before relying on that.

ScreenshotNeo is the first hosted option to try: it removes known consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and its paid plans start at $5 for 3,000 shots. Its documentation describes the API and available parameters.

2. Prepare the URL list

Keep input order and source URLs traceable to each image. Save one URL per line in urls.txt. Blank lines and lines beginning with # are ignored by the example. Use complete URLs with a scheme such as https://. Do not put credentials or private tokens in URLs if the output directory or logs could be shared.

https://example.com/
https://www.python.org/
https://playwright.dev/

Choose the capture mode before you start:

  • Viewport: fixed-size images show the initial view and make a regular grid. This is the mode used below.
  • Full page: captures the full scrollable page, but pages can have very different heights. Long pages may become too small to inspect when scaled into equal-sized sheet cells.
  • Element: captures a specific component, such as a product card, when the comparison is about that component rather than the whole page.

3. Capture screenshots with Playwright

Install Playwright and its Chromium browser. This script reads the URL file, uses a fresh page for every URL, writes an indexed PNG per successful capture, and writes a CSV manifest that maps each input to its output or error. It continues after individual navigation failures.

python -m pip install playwright
python -m playwright install chromium
import asyncio
import csv
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright

INPUT = Path("urls.txt")
OUT = Path("captures")
VIEWPORT = {"width": 1280, "height": 800}
TIMEOUT_MS = 30_000


def load_urls(path):
    urls = []
    for line_number, raw in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
        value = raw.strip()
        if not value or value.startswith("#"):
            continue
        parsed = urlparse(value)
        if parsed.scheme not in {"http", "https"} or not parsed.netloc:
            raise ValueError(f"Line {line_number}: expected an http(s) URL, got {value!r}")
        urls.append(value)
    return urls


async def main():
    urls = load_urls(INPUT)
    OUT.mkdir(parents=True, exist_ok=True)
    manifest_rows = []

    async with async_playwright() as p:
        browser = await p.chromium.launch()
        for index, url in enumerate(urls, start=1):
            filename = f"{index:04d}.png"
            page = await browser.new_page(viewport=VIEWPORT, device_scale_factor=1)
            error = ""
            title = ""
            try:
                response = await page.goto(url, wait_until="domcontentloaded", timeout=TIMEOUT_MS)
                # Give client-rendered pages a short chance to paint after DOM readiness.
                await page.wait_for_timeout(750)
                title = await page.title()
                await page.screenshot(path=str(OUT / filename), full_page=False)
                status = str(response.status) if response else "no-response"
                if response and response.status >= 400:
                    error = f"HTTP {response.status}; screenshot saved for review"
            except Exception as exc:
                status = "failed"
                error = str(exc)
            finally:
                await page.close()
            manifest_rows.append({"index": index, "url": url, "title": title,
                                  "file": filename if (OUT / filename).exists() else "",
                                  "status": status, "error": error})
            print(f"{index}/{len(urls)} {status} {url}")
        await browser.close()

    with (OUT / "manifest.csv").open("w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=["index", "url", "title", "file", "status", "error"])
        writer.writeheader()
        writer.writerows(manifest_rows)
    print(f"Saved captures and manifest in {OUT}/")


if __name__ == "__main__":
    asyncio.run(main())

Run it from the directory containing urls.txt:

python capture.py

The script waits for DOM content rather than full network idle. Some sites keep analytics or streaming requests open indefinitely, so waiting for network idle can make a bulk job hang. The short post-navigation pause is only a practical default; increase it or wait for a site-specific selector if important content renders later.

4. Assemble a labeled contact sheet with Pillow

Install Pillow, then run this script. It reads the manifest, fits each screenshot into a consistent tile without stretching it, and labels each tile with its input index and URL. Failed captures appear as labeled placeholders, so the grid still reflects the complete input list.

python -m pip install pillow
import csv
import math
from pathlib import Path
from PIL import Image, ImageDraw, ImageFont, ImageOps

CAPTURES = Path("captures")
MANIFEST = CAPTURES / "manifest.csv"
OUTPUT = Path("contact-sheet.jpg")
COLS = 3
TILE_W, TILE_H = 400, 300
LABEL_H = 62
GAP = 16
MARGIN = 16
BG = (238, 240, 243)

try:
    FONT = ImageFont.truetype("DejaVuSans.ttf", 14)
except OSError:
    FONT = ImageFont.load_default()

with MANIFEST.open(encoding="utf-8", newline="") as f:
    rows = list(csv.DictReader(f))

if not rows:
    raise SystemExit("Manifest is empty; there are no URLs to put on the sheet.")

rows_count = math.ceil(len(rows) / COLS)
cell_w = TILE_W
cell_h = TILE_H + LABEL_H
sheet_w = MARGIN * 2 + COLS * cell_w + (COLS - 1) * GAP
sheet_h = MARGIN * 2 + rows_count * cell_h + (rows_count - 1) * GAP
sheet = Image.new("RGB", (sheet_w, sheet_h), BG)
draw = ImageDraw.Draw(sheet)

for i, row in enumerate(rows):
    col, r = i % COLS, i // COLS
    x = MARGIN + col * (cell_w + GAP)
    y = MARGIN + r * (cell_h + GAP)
    image_path = CAPTURES / row["file"] if row.get("file") else None
    if image_path and image_path.exists():
        with Image.open(image_path) as source:
            tile = ImageOps.contain(source.convert("RGB"), (TILE_W, TILE_H))
        tile_x = x + (TILE_W - tile.width) // 2
        tile_y = y + (TILE_H - tile.height) // 2
        sheet.paste(tile, (tile_x, tile_y))
        draw.rectangle((x, y, x + TILE_W - 1, y + TILE_H - 1), outline=(190, 195, 202))
    else:
        draw.rectangle((x, y, x + TILE_W - 1, y + TILE_H - 1), fill=(220, 223, 228))
        draw.text((x + 12, y + 12), "Capture failed", fill=(90, 30, 30), font=FONT)

    label = f"{row['index']}. {row['url']}"
    draw.text((x, y + TILE_H + 7), label[:70], fill=(25, 28, 34), font=FONT)
    if row.get("error"):
        draw.text((x, y + TILE_H + 31), row["error"][:65], fill=(130, 45, 45), font=FONT)

sheet.save(OUTPUT, quality=88, optimize=True)
print(f"Wrote {OUTPUT} ({sheet_w} x {sheet_h})")

Run:

python make_sheet.py

Adjust COLS, TILE_W, and TILE_H to change the layout. The sheet uses JPEG to keep a large grid more compact; save as PNG if lossless text and edges matter more than file size. Keep the original PNGs for detailed review.

5. Handle full-page captures and element captures

To capture the full scrollable page instead of the viewport, change the screenshot line in the Playwright script to:

await page.screenshot(path=str(OUT / filename), full_page=True)

Playwright documents full-page screenshots and screenshot buffers for post-processing. It also supports capturing a single element by locating it and calling screenshot on that locator. For example:

card = page.locator("main article").first
await card.screenshot(path=str(OUT / filename))

Use a selector appropriate to the target site and handle cases where the element is absent or hidden. Full-page captures can be much taller than the fixed contact-sheet tile. Pillow’s ImageOps.contain preserves the complete image but scales it down; for a scannable overview, viewport shots or a consistent crop usually work better. Retain full-size originals when page detail matters.

6. Hosted capture alternatives

ScreenshotNeo

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API can return PNG, JPEG, WebP, or PDF output, and supports bulk capture of up to 100 URLs per call. The examples below show a single URL; for a contact sheet, capture each URL, retain the URL-to-image mapping, then use the assembly step above. See the API documentation for current parameters and bulk request format.

ScreenshotRun

ScreenshotRun’s batch documentation describes thumbnail captures from a list of websites and says each URL in a batch counts as one screenshot toward the monthly quota. The reviewed documentation does not establish that it composes a contact sheet. Check its live documentation for current limits, pricing, result format, and quota terms before selecting it.

Or skip the browser setup

A single ScreenshotNeo call captures one URL; repeat it for each address in your list, or use its bulk capture option for up to 100 URLs per call. This cURL example saves a WebP image:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot; those steps can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers identify the page verdict and whether the request was billed.
  • An MCP server lets AI agents, including Claude and Cursor, take screenshots with tools for screenshots, page information, and PDF capture.
  • 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000; every feature is available on every plan.

Create a free ScreenshotNeo account to get 1,000 screenshots per month without a card. See the documentation for API setup.

Troubleshooting

Symptom Likely cause Fix
Invalid URL error The input line lacks http:// or https://, or has no hostname. Correct the line in urls.txt. The script reports the line number during validation.
Navigation timeout The page is slow, unreachable, or keeps connections open. Check the URL manually, increase TIMEOUT_MS, or use a more targeted readiness condition. Do not use network-idle waiting indiscriminately on pages with ongoing requests.
Screenshot is blank or missing content Content may render after DOM readiness, require scrolling, depend on consent, or be blocked by bot protection. Increase the post-navigation wait or wait for a known selector. For lazy-loaded content, scroll incrementally before capture and allow images to load. Some bot checks cannot be solved with generic automation; review the page and use an authorized capture route.
Browser executable not found Playwright’s Chromium browser was not installed for the active environment. Run python -m playwright install chromium in the same environment as the script.
Contact sheet reports empty manifest No valid URLs were provided or the capture step did not write its manifest. Check urls.txt, run the capture script, and confirm captures/manifest.csv exists.
Some cells are placeholders Those captures failed before an image was saved. Read the manifest’s error field, fix the affected URL or wait condition, and rerun those entries.
Sheet is too large or tiles are unreadable There are many URLs, oversized captures, or full-page screenshots with long pages. Split the input into smaller groups, reduce tile dimensions, use viewport captures for overview, or produce separate sheets by category.
Hosted API response is an error instead of an image Credentials, URL encoding, request parameters, or a service-side condition may be wrong. Check the HTTP status and response headers/body, verify the key and encoded URL, and compare parameters with the provider’s current documentation.

Performance, reliability, and cost

  • Control concurrency: the local example captures sequentially, which is straightforward and limits simultaneous browser work. For larger lists, bounded concurrency can improve throughput, but too many pages consume memory and can trigger rate limits or bot defenses. Choose a modest limit and measure on your own URLs.
  • Reuse the browser: the example launches Chromium once and opens a page per URL. For reproducibility, use a consistent browser version, viewport, device scale factor, and locale if relevant.
  • Keep a manifest: indexed filenames, URLs, titles, status, and errors make reruns and review possible. Store input order even if you later sort tiles by title or domain.
  • Retry selectively: retry transient timeouts or network errors with a small capped retry count. Avoid blindly retrying authorization failures, invalid URLs, or bot challenges. Preserve the first error in logs.
  • Manage output size: PNG preserves image detail but many large captures take disk space. A JPEG contact sheet is smaller; keep source captures if you need to inspect fine text. Full-page images can be especially tall.
  • Account for hosted quotas: each captured URL may count separately under a provider’s rules. ScreenshotRun documents per-URL quota accounting for its batch feature; verify live plan terms. ScreenshotNeo lists 1,000 free shots per month, then paid tiers from $5 for 3,000 shots.
  • Protect sensitive URLs: URLs may expose internal identifiers or signed parameters. Avoid putting secrets in shared manifests, contact sheets, command history, or logs. Confirm that a hosted provider is appropriate for the URLs before sending them.

FAQ

Can I make one sheet from hundreds of URLs?

Yes. Capture and assemble in batches if memory, sheet dimensions, or reviewability become a problem. A set of smaller sheets is often easier to scan and share.

Should the sheet show the URL or the page title?

Use the URL when source identity matters; use the title when readers recognize page names more easily. A manifest preserves the full mapping either way. The sample labels the URL and index.

Can I compare pages at the same size?

Use the same viewport and device scale factor for every capture. If comparing a specific component, capture the same type of element on each page and expect selectors to vary by site.

Does a batch screenshot API automatically create the contact sheet?

Do not assume so. The reviewed ScreenshotRun batch documentation establishes list-based capture, not sheet composition. ScreenshotNeo provides captures and batch capture; use an image-processing step to arrange the results.

Sources