ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Every URL in a CSV and Create a Contact Sheet

Use Python and Playwright to capture each URL in a CSV, track failures, and build a labeled contact sheet you can scan quickly.

By the ScreenshotNeo team4 October 202611 min read

To capture screenshots of every URL in a CSV and create a contact sheet, use Python’s CSV reader to validate and preserve the input rows, Playwright to capture each page, and Pillow to resize and arrange the resulting images. The script below saves one screenshot per usable URL, records successes and failures in a manifest, and creates a labeled contact sheet in CSV order.

This approach works well for a modest batch you can run from your own machine. It handles blank rows and per-page navigation failures without stopping the batch. Websites can redirect, load slowly, block automation, or render differently depending on consent, authentication, and network conditions, so inspect the manifest and sample captures before relying on the sheet.

1. Install the dependencies

Use Python 3.9 or later. Playwright controls Chromium; Pillow creates thumbnails and the contact sheet.

python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install playwright pillow
python -m playwright install chromium

Save the script below as csv_contact_sheet.py. Put your CSV beside it, or pass the CSV path as the first command-line argument. By default, the script expects a header named url.

2. Capture each URL and build the sheet

#!/usr/bin/env python3
"""Capture URLs from a CSV and assemble a labeled contact sheet."""

import csv
import json
import re
import sys
from pathlib import Path
from urllib.parse import urlparse

from PIL import Image, ImageDraw, ImageFont
from playwright.sync_api import TimeoutError as PlaywrightTimeoutError
from playwright.sync_api import sync_playwright

INPUT = Path(sys.argv[1]) if len(sys.argv) > 1 else Path("urls.csv")
URL_COLUMN = "url"
OUTPUT = Path("captures")
SCREENSHOT_FORMAT = "png"  # png, jpeg, or webp
FULL_PAGE = False           # True includes the page's full scrollable height
VIEWPORT = {"width": 1440, "height": 1000}
NAVIGATION_TIMEOUT_MS = 30_000
WAIT_UNTIL = "domcontentloaded"  # load, domcontentloaded, or networkidle
THUMBNAIL_SIZE = (360, 220)
LABEL_HEIGHT = 48
COLUMNS = 3
GAP = 16


def safe_slug(url: str) -> str:
    """Create a readable filename component without trusting URL characters."""
    host = urlparse(url).netloc.lower() or "page"
    host = re.sub(r"[^a-z0-9.-]+", "-", host).strip("-.") or "page"
    return host[:70]


def read_rows(path: Path):
    # utf-8-sig also accepts UTF-8 files with a byte-order mark from spreadsheet apps.
    with path.open("r", encoding="utf-8-sig", newline="") as source:
        reader = csv.DictReader(source)
        if not reader.fieldnames or URL_COLUMN not in reader.fieldnames:
            raise ValueError(
                f"CSV must have a header named {URL_COLUMN!r}; found {reader.fieldnames!r}"
            )
        return list(reader)


def fit_thumbnail(image: Image.Image, box):
    image = image.convert("RGB")
    image.thumbnail(box, Image.Resampling.LANCZOS)
    canvas = Image.new("RGB", box, "#f1f3f5")
    x = (box[0] - image.width) // 2
    y = (box[1] - image.height) // 2
    canvas.paste(image, (x, y))
    return canvas


def make_contact_sheet(entries, destination: Path):
    successful = [entry for entry in entries if entry["status"] == "ok"]
    if not successful:
        print("No screenshots succeeded; contact sheet not created.", file=sys.stderr)
        return

    cell_width = THUMBNAIL_SIZE[0]
    cell_height = THUMBNAIL_SIZE[1] + LABEL_HEIGHT
    rows = (len(successful) + COLUMNS - 1) // COLUMNS
    sheet = Image.new(
        "RGB",
        (GAP + COLUMNS * (cell_width + GAP), GAP + rows * (cell_height + GAP)),
        "white",
    )
    draw = ImageDraw.Draw(sheet)
    try:
        font = ImageFont.truetype("DejaVuSans.ttf", 13)
    except OSError:
        font = ImageFont.load_default()

    for position, entry in enumerate(successful):
        col = position % COLUMNS
        row = position // COLUMNS
        x = GAP + col * (cell_width + GAP)
        y = GAP + row * (cell_height + GAP)
        with Image.open(OUTPUT / entry["file"]) as captured:
            thumb = fit_thumbnail(captured, THUMBNAIL_SIZE)
        sheet.paste(thumb, (x, y))
        label = f"CSV row {entry['row']}: {entry['url']}"
        # Keep the label readable and bounded to two lines.
        max_chars = 48
        label = label if len(label) <= max_chars else label[: max_chars - 1] + "…"
        draw.text((x, y + THUMBNAIL_SIZE[1] + 6), label, fill="#202124", font=font)

    sheet.save(destination, format="PNG", optimize=True)


def main():
    OUTPUT.mkdir(parents=True, exist_ok=True)
    rows = read_rows(INPUT)
    manifest = []

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch()
        context = browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
        page = context.new_page()
        page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

        for csv_row_number, row in enumerate(rows, start=2):
            original_url = (row.get(URL_COLUMN) or "").strip()
            entry = {"row": csv_row_number, "url": original_url, "file": None, "status": None, "error": None}
            if not original_url:
                entry["status"] = "skipped"
                entry["error"] = "blank URL"
                manifest.append(entry)
                print(f"Row {csv_row_number}: skipped blank URL")
                continue
            parsed = urlparse(original_url)
            if parsed.scheme not in ("http", "https") or not parsed.netloc:
                entry["status"] = "skipped"
                entry["error"] = "URL must be absolute and use http or https"
                manifest.append(entry)
                print(f"Row {csv_row_number}: skipped invalid URL {original_url!r}")
                continue

            filename = f"{csv_row_number:04d}-{safe_slug(original_url)}.{SCREENSHOT_FORMAT}"
            entry["file"] = filename
            try:
                response = page.goto(original_url, wait_until=WAIT_UNTIL)
                # A navigation response with an HTTP error status can still render a useful page.
                page.screenshot(path=str(OUTPUT / filename), full_page=FULL_PAGE, type=SCREENSHOT_FORMAT)
                entry["status"] = "ok"
                entry["http_status"] = response.status if response else None
                print(f"Row {csv_row_number}: captured {original_url}")
            except (PlaywrightTimeoutError, Exception) as error:
                entry["status"] = "failed"
                entry["error"] = f"{type(error).__name__}: {error}"
                print(f"Row {csv_row_number}: failed {original_url}: {entry['error']}", file=sys.stderr)
            manifest.append(entry)

        context.close()
        browser.close()

    manifest_path = OUTPUT / "manifest.json"
    manifest_path.write_text(json.dumps(manifest, indent=2, ensure_ascii=False), encoding="utf-8")
    make_contact_sheet(manifest, OUTPUT / "contact-sheet.png")
    succeeded = sum(item["status"] == "ok" for item in manifest)
    skipped = sum(item["status"] == "skipped" for item in manifest)
    failed = sum(item["status"] == "failed" for item in manifest)
    print(f"Done: {succeeded} captured, {skipped} skipped, {failed} failed.")
    print(f"Manifest: {manifest_path}; output directory: {OUTPUT.resolve()}")


if __name__ == "__main__":
    try:
        main()
    except (OSError, ValueError, csv.Error) as error:
        raise SystemExit(f"Input or setup error: {error}")

Example urls.csv:

url
https://example.com/
https://www.python.org/
https://playwright.dev/python/

Run it with python csv_contact_sheet.py urls.csv. The captures directory contains individual screenshots, manifest.json, and contact-sheet.png. CSV row numbers in the labels count the header as row 1, making them correspond to spreadsheet row numbers.

3. Understand the capture and layout choices

Viewport or full-page screenshots

The script defaults to viewport captures at 1440 by 1000 CSS pixels. This yields similarly sized previews that are easier to compare in a grid. Set FULL_PAGE = True to capture the full scrollable page; Playwright describes this as capturing the page as if it fit on a very tall screen. Full-page images can have different heights and shrink to unreadable thumbnails, so use them when below-the-fold content matters and consider a separate sheet or gallery.

Playwright’s Python screenshot API supports PNG, JPEG, and WebP output, along with viewport and full-page capture. The script uses PNG by default for crisp text. Choose JPEG or WebP when smaller individual files matter more than lossless text edges; the contact sheet remains PNG. See the Playwright screenshots guide and Page screenshot API.

Load state, timeout, and dynamic pages

domcontentloaded waits for the document to be parsed, not for every image or third-party request to finish. It usually avoids waiting on analytics connections that never become idle. Use load if the page needs ordinary load events. Use networkidle only when the target pages settle their network activity; applications with polling or long-lived requests may never reach it. For a site-specific application, wait for a known selector after navigation, for example page.locator("main article").wait_for(), or use a short explicit delay for a known animation. Avoid adding long fixed waits to every URL without a reason.

Increase NAVIGATION_TIMEOUT_MS for slow sites. A timeout is recorded as a row failure and the loop continues. Navigation may redirect; the saved capture is of the final rendered page, while http_status records the initial navigation response when available.

CSV format and traceability

csv.DictReader maps each data row to a dictionary keyed by the header. Opening with newline="" is recommended by Python’s CSV documentation, and utf-8-sig accepts common UTF-8 files with a byte-order mark. If a spreadsheet exported semicolon-separated data, pass delimiter=";" to csv.DictReader. Change URL_COLUMN to match your header exactly. Keep the manifest with the images: it records original URLs, CSV row numbers, filenames, statuses, errors, and response status. See the Python CSV documentation.

Contact-sheet sizing and labels

THUMBNAIL_SIZE controls the preview box, COLUMNS sets the grid width, and GAP sets spacing. Images preserve aspect ratio and are centered in a fixed box. The contact sheet includes successful captures only; failures and skipped rows remain visible in the manifest and terminal output. If the list is large, generate multiple sheets in batches or make an HTML gallery that links to full-size screenshots. A single enormous bitmap becomes slow to open and difficult to scan.

4. Run the same workflow with cURL, Python, or Node.js

The full batch and image-layout logic is most convenient in Python. If another program already manages your CSV, ScreenshotNeo can return an image for each URL through one GET request. The following examples demonstrate one URL; a caller can repeat the request for each CSV row and assemble the returned image files locally.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/ \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

For API options and response headers, see the ScreenshotNeo API documentation.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns a screenshot for a URL; run that call once for each CSV entry, then use the contact-sheet function above to arrange the saved images. Cookie banners are accepted like a visitor and removed before the shot, along with known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. AI agents can capture through its MCP server. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/ \
  -o shot.webp

Repeat with each CSV URL, naming each output by its row number, then build the contact sheet from those files. The API also supports PNG, JPEG, and WebP output. Read the API documentation, then sign up free for 1,000 screenshots a month with no card.

5. Troubleshooting

Symptom Likely cause Fix
The script says the URL column is missing The header is different, contains whitespace, or uses a different delimiter. Set URL_COLUMN to the exact header. For semicolon-separated CSV, create the reader with delimiter=";". Check the first row in a text editor.
Rows are skipped as invalid The cell contains a relative path, lacks a scheme, or has a typo. Provide an absolute address beginning with https:// or http://, and inspect the row number printed in the log.
Chromium executable is missing The Python package installed, but its browser binary did not. Run python -m playwright install chromium in the active environment.
Navigation times out The site is slow, hangs on requests, or the selected wait state is too strict. Increase the timeout, use domcontentloaded, or wait for a specific page element. Treat persistent failures as inaccessible or automation-blocked pages and retain the manifest entry.
Screenshot is blank or incomplete The app renders after navigation, requires authentication, or loads content on scroll. Wait for a stable selector, configure authentication for that site, or scroll through the page before capture to trigger lazy loading. Verify the page manually in a browser when access is uncertain.
Labels are too long or hard to read Long URLs consume the available label space or the sheet contains too many tiles. Shorten the label to host and path, increase LABEL_HEIGHT and draw wrapped lines, enlarge thumbnails, or split the output across sheets.
The contact sheet is missing No capture succeeded, or Pillow could not reopen an output image. Check terminal errors and manifest.json. Confirm screenshot files exist and that disk space is available.

6. Performance, reliability, and cost considerations

This script uses one browser, one context, and one page, processing URLs sequentially. That keeps the workflow simple and limits simultaneous load on target sites, but total runtime grows with the number of URLs and each page’s navigation and rendering time. Avoid opening too many pages concurrently: it increases memory use and may trigger rate limits or bot defenses. If you add concurrency, use a small bounded worker pool, keep per-row error handling, and respect the websites’ access policies.

Browser captures depend on local CPU, memory, network, browser binaries, and the behavior of each destination site. Keep a manifest and rerun only failed rows when possible. Stable filenames include the CSV row number, so duplicate URLs still produce distinct outputs. Consider a fresh output directory for each run if you do not want old screenshots from a prior CSV left alongside new files.

Playwright itself is an open-source browser automation library; this local workflow does not charge a per-screenshot API fee, but it uses your machine’s resources and network. ScreenshotNeo’s API has a free allowance of 1,000 shots per month and paid tiers from $5 for 3,000; only clean shots are billed under the stated product rules. For large repeat batches, compare the plan allowance with your expected usable captures and account for your own contact-sheet processing. Do not treat a failed local capture as equivalent to the API’s billing verdict.

7. Frequently asked questions

Does the script preserve the original order?

Yes. It processes CSV rows in order and lays successful captures into the sheet in that same order. Row numbers in labels preserve the source location when invalid or failed entries leave gaps.

Can I capture pages that require a login?

Only if you configure an authenticated browser context for the site, such as loading saved storage state or setting the required cookies. Protect that state file and do not commit credentials to source control.

Can I create a PDF instead of a contact-sheet image?

The script writes a PNG contact sheet. Pillow can save the composed image as a PDF, but a multi-page report or HTML gallery is usually easier to navigate for a large set of captures.

Will a screenshot include content loaded only after scrolling?

Not necessarily. Some pages lazy-load images or sections as they enter the viewport. Scroll through the page or use a full-page capture strategy appropriate to the site, then check representative output files.