ScreenshotNeo

BlogHow-to

How to Capture Website Screenshots in Bulk for a Research Report

Capture a URL list consistently with Playwright, save attributable images and a manifest, and handle failures before using screenshots as report evidence.

By the ScreenshotNeo team4 October 20269 min read

To capture website screenshots in bulk for a research report, automate browser navigation over a fixed URL list, save one uniquely named image per URL, and record each result in a manifest. Playwright is a flexible choice when you need code-level control; its CLI or shot-scraper can suit simpler command-line runs. Decide in advance whether the report needs the visible viewport, a full page, or a specific element, then keep capture settings and the environment consistent across the batch.

This guide uses Python with Playwright for a runnable batch script, then covers capture scope, formats, auditability, failures, and alternatives. Playwright supports navigation and screenshots to files or buffers, including full-page captures. Playwright screenshot documentation.

1. Define what the report needs to show

Choose the capture scope before running the batch. Different scopes answer different research questions; mixing them can make comparisons misleading.

Scope Use it when Trade-off
Viewport You are comparing the initial presentation or above-the-fold content. Content below the visible area is omitted.
Full page The report needs the page’s scrollable content. The resulting image can be very tall and harder to inspect at normal size.
Element You need evidence of one chart, panel, or component. You must identify a stable selector for the target element.

Choose PNG, JPEG, or WebP based on the report’s needs and destination. Higher device-pixel scale can improve legibility but increases image dimensions. Inspect representative output at the size it will appear in the report.

2. Prepare a URL list and output folder

Put one URL per line in urls.txt. Use complete URLs, including https:// where applicable. Create an empty output directory such as captures/, and keep the input list and script with the results so the batch can be repeated.

https://example.com/
https://www.python.org/
https://playwright.dev/

Install Playwright and its Chromium browser. These commands use Python’s module runner, so the install and script use the same Python environment.

python -m pip install playwright
python -m playwright install chromium

3. Run a batch with Python and Playwright

Save this as capture_batch.py. It captures full pages as PNG, assigns a unique indexed filename to every input line, and writes a CSV manifest containing the requested URL, final URL when navigation succeeds, UTC capture time, filename, and outcome. Failed navigation is recorded and the script continues to the next URL. The script deliberately does not claim that a successful navigation proves the screenshot is authoritative evidence.

import asyncio
import csv
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.async_api import async_playwright

INPUT_FILE = Path("urls.txt")
OUTPUT_DIR = Path("captures")
MANIFEST = OUTPUT_DIR / "manifest.csv"
NAVIGATION_TIMEOUT_MS = 30_000
VIEWPORT = {"width": 1440, "height": 1000}
FULL_PAGE = True


def safe_host(url: str) -> str:
    host = urlparse(url).hostname or "unknown-host"
    return "".join(c if c.isalnum() or c in ".-_" else "_" for c in host)


async def main() -> None:
    urls = [line.strip() for line in INPUT_FILE.read_text(encoding="utf-8").splitlines()
            if line.strip() and not line.lstrip().startswith("#")]
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

    fields = ["index", "requested_url", "final_url", "captured_at_utc", "filename", "outcome", "error"]
    rows = []

    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        context = await browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
        page = await context.new_page()
        page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

        for index, url in enumerate(urls, start=1):
            stamp = datetime.now(timezone.utc).isoformat()
            filename = f"{index:04d}-{safe_host(url)}.png"
            target = OUTPUT_DIR / filename
            final_url = ""
            try:
                response = await page.goto(url, wait_until="domcontentloaded")
                final_url = page.url
                # A missing response can occur for non-HTTP(S) navigation. HTTP error
                # responses still render; record their status so they can be reviewed.
                status_note = ""
                if response is not None and response.status >= 400:
                    status_note = f"HTTP {response.status}"
                await page.screenshot(path=str(target), full_page=FULL_PAGE, type="png")
                rows.append({"index": index, "requested_url": url, "final_url": final_url,
                             "captured_at_utc": stamp, "filename": filename,
                             "outcome": status_note or "captured", "error": ""})
            except Exception as exc:
                rows.append({"index": index, "requested_url": url, "final_url": page.url,
                             "captured_at_utc": stamp, "filename": "",
                             "outcome": "failed", "error": f"{type(exc).__name__}: {exc}"})

        await browser.close()

    with MANIFEST.open("w", newline="", encoding="utf-8") as handle:
        writer = csv.DictWriter(handle, fieldnames=fields)
        writer.writeheader()
        writer.writerows(rows)

    captured = sum(row["outcome"] in ("captured",) or row["outcome"].startswith("HTTP ") for row in rows)
    failed = len(rows) - captured
    print(f"Finished: {captured} captured, {failed} failed; manifest: {MANIFEST}")


if __name__ == "__main__":
    asyncio.run(main())

Run it from the directory containing urls.txt and capture_batch.py:

python capture_batch.py

The script uses one browser context and processes URLs sequentially to keep the capture configuration stable and limit concurrent load on target sites. For viewport-only images, set FULL_PAGE = False. For very large URL lists, consider chunking the input and preserving a manifest for each run.

4. Adapt capture scope and output

Viewport screenshots

Set FULL_PAGE = False in the script. The viewport size is configured by VIEWPORT; use the same dimensions for every URL in a comparison set.

Element screenshots

After navigation and before saving, locate the element and screenshot its bounding box:

locator = page.locator("main .research-chart").first
await locator.wait_for(state="visible", timeout=10_000)
await locator.screenshot(path=str(target), type="png")

Replace the selector with one appropriate to the site. If a selector matches nothing, is hidden, or matches multiple candidates unexpectedly, the capture may fail or target the wrong component. Validate it on a sample page first.

Format and resolution

Playwright’s screenshot API supports image type and device-pixel output. PNG is a useful lossless default for text and charts; JPEG and WebP can be appropriate when smaller image files matter and their rendering is acceptable for your report. Set type="jpeg" or type="webp" and use the matching file extension. Browser screenshot output uses CSS-pixel dimensions unless device scale affects the captured pixels; keep device_scale_factor fixed when comparable output matters. See the screenshot options.

The example waits for domcontentloaded, a practical starting point that avoids waiting for every resource. Sites that render important content later may need a different readiness condition, a targeted selector wait, or a short delay. Waiting for all network activity can be unreliable on pages that maintain long-lived connections. Choose a condition that reflects the content your report needs, and record any non-default wait procedure.

5. Make the batch comparable and auditable

  • Keep the browser engine and version, operating environment, viewport, device scale, and procedure stable across the batch.
  • Keep each filename unique. The indexed host-based names above avoid collisions between repeated hostnames in one run; retain the manifest to map files to exact requested URLs.
  • Record requested URL, final URL, capture time, filename, and outcome. Add browser/version, viewport, scale, and any special wait or interaction settings if the report may need to be reproduced.
  • Run a representative sample first. Inspect redirects, blank pages, unexpected consent dialogs, failed loads, and image readability before scaling up.
  • Review the output directory and manifest together. A row marked captured means the script saved a file; it does not establish that the page content is correct or complete.
  • In the report, identify the screenshot and capture date and cite the original page as well. A screenshot records appearance at a moment; it does not by itself establish authorship, publication date, accuracy, or continued availability.

Rendering can vary with operating system, browser version, settings, hardware, power source, and headless mode. Playwright recommends using the same environment as the visual baseline when comparing screenshots. See its visual comparison guidance.

6. Other do-it-yourself options

Option Good fit Considerations
Playwright script Custom interaction, programmatic error handling, and manifest metadata. Requires maintaining a script and browser environment.
Playwright CLI One-off or shell-driven captures using documented options for viewport, element, full page, filename, format, and device scale. More complex per-URL logic and audit metadata may be easier in a script. See the CLI documentation.
shot-scraper A command-line utility for automated website screenshots built on Playwright. Check its documentation for the exact invocation and options needed by your workflow: shot-scraper documentation.

The available documentation does not establish comparative speed, reliability, or pricing benchmarks for these approaches. Choose based on interaction needs, capture scope, output configuration, and how much metadata and error handling the report requires.

7. Troubleshooting

Symptom Likely cause Fix
Browser executable missing Playwright’s Python package is installed, but its browser was not installed in that environment. Run python -m playwright install chromium with the same Python interpreter used to run the script.
Navigation timeout The site is slow, unreachable, or waiting for the selected navigation condition never completes. Check the URL and network access, then increase NAVIGATION_TIMEOUT_MS cautiously. Use a readiness condition suited to the page instead of waiting indefinitely for all activity.
Screenshot is blank or incomplete The page may render content after DOM readiness, require scrolling, or depend on an interaction. Wait for a meaningful selector or a bounded delay; inspect the page manually and add the necessary interaction. Recheck a sample before rerunning the batch.
Element capture fails The selector does not match, the element is hidden, or it never becomes visible. Confirm the selector in the page, wait for visibility, and decide how to handle pages where the component is absent.
HTTP error page was captured The server returned an HTTP status of 400 or greater, but the browser still displayed a page. Review the manifest’s HTTP status outcome. Decide whether to retain the error page as evidence or retry after checking the URL and access conditions.
Files overwrite or are hard to map Names are not unique or do not identify the input reliably. Use indexed names and keep the URL-to-file manifest. If running multiple batches in the same folder, use a separate folder per run.
Images differ between runs Browser or host environment, content, viewport, scale, or timing changed. Stabilize the environment and settings, record capture time, and compare only captures made under a consistent procedure.

8. Performance, reliability, and cost

For modest batches, sequential capture is simple to audit and avoids creating many browser pages at once. It can take longer than parallel capture, but the cited documentation supplies no speed benchmark. If you add concurrency, use a bounded number of pages and check that target sites and the capture machine can handle it; record the concurrency setting so a later run is interpretable.

Reliability comes from making each URL an independent row in the manifest, continuing after individual failures, and reviewing the output. A navigation or screenshot exception is not the only possible bad result: redirects, HTTP error responses, and visually blank pages also need inspection. Preserve the original URL list and run settings alongside the images.

The reviewed workflow uses software tools. The research sources do not establish a required physical product or provide tool pricing. Account for the machine and time needed to run the batch, and check current terms separately if selecting a hosted service.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request can return an image or PDF. Use this cURL example for a single URL; for a batch, call the endpoint once per URL and save each response under a unique filename while retaining the same manifest fields described above. The full configuration and parameters are in the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie and consent banners are accepted like a visitor and removed before capture, along with supported newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, with no card required.

10. Frequently asked questions

Should every page use the same capture scope?

Use the same scope when the report compares like with like. If different pages need different scopes, record the choice per capture and explain it in the report.

Does a screenshot prove a page’s claims?

No. It preserves a visual record at capture time. Link the source page and state when the image was captured; evaluate the underlying claims separately.

Can I capture a very large URL list in one run?

You can process a list in chunks to simplify recovery and review. Keep the original list and a manifest for each run so every output remains attributable.

Which option should I start with?

Start with a Playwright script when you need tailored interactions, error handling, or an audit manifest. Try the CLI or shot-scraper when their documented command options fit the capture without custom per-page logic.