ScreenshotNeo

BlogHow-to

How to Bulk Screenshot URLs and Log HTTP Errors

Build a repeatable URL screenshot batch that records redirects, HTTP status, capture results, and errors in a durable CSV log.

By the ScreenshotNeo team4 October 202612 min read

Use a browser automation script that records the main document’s HTTP response separately from whether the screenshot succeeded. For every input URL, keep the original URL, final URL after redirects, response status (or null when no response arrived), capture outcome, screenshot filename, and any error details. An HTTP 404 or 500 page can still render and be captured; a timeout or DNS failure may prevent both a response and a screenshot.

The example below uses Python and Playwright. It writes one CSV row per URL as the batch runs, saves screenshots with unique filenames, continues after per-URL failures, and optionally captures the full scrollable page. Playwright documents both screenshot files and the full_page option in its screenshot guide.

1. Choose what counts as an error

Keep these outcomes distinct in your log:

Outcome What it means Typical handling
HTTP response received The main document returned a status such as 200, 404, or 503. Record the status and capture the rendered page when possible. Decide separately whether the status meets your reporting policy.
Navigation or network failure No usable main-document response arrived, for example because of DNS failure, connection refusal, or timeout. Record a null status, a failure category, and the exception. Continue with the next URL.
Screenshot failure A page may have loaded, but the browser could not save its screenshot. Keep the received HTTP status and log the screenshot exception independently.
Browser or script failure The browser process or batch itself stopped unexpectedly. Incremental CSV output preserves rows already completed. Restart or split the remaining work.

Do not use screenshot existence as a proxy for HTTP success: servers often return error pages with normal HTML that a browser can display. Likewise, a browser timeout is not an HTTP 5xx response. A hosted screenshot service describes this distinction between page verdict and billing status in its response behavior; the implementation below applies the general logging distinction directly. See ScreenshotCenter’s batch screenshot overview.

2. Install Playwright and prepare the URL list

Use Python 3.9 or newer, create a project directory, and install the package and its Chromium browser:

python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
python -m playwright install chromium

On Windows PowerShell, activate with .venv\Scripts\Activate.ps1. Save one URL per line in urls.txt:

https://example.com/
https://example.org/
https://example.net/missing-page

Install the browser in the same environment that will run the job. In containers or CI, include browser dependencies and persist the output directory and CSV as artifacts.

3. Run a batch that logs each result incrementally

Save this as bulk_screenshot.py. The script captures the main navigation response, follows redirects through normal browser navigation, records the final page URL, and appends each completed result to CSV immediately. HTTP error statuses do not stop the batch. Navigation and screenshot failures get separate fields.

import argparse
import csv
import re
import sys
from pathlib import Path
from urllib.parse import urlsplit

from playwright.sync_api import (
    Error as PlaywrightError,
    TimeoutError as PlaywrightTimeoutError,
    sync_playwright,
)

FIELDS = [
    "input_url",
    "final_url",
    "http_status",
    "capture_ok",
    "screenshot_path",
    "navigation_error_type",
    "navigation_error",
    "screenshot_error",
]


def read_urls(path: Path) -> list[str]:
    urls = []
    for line_number, raw in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
        value = raw.strip()
        if not value or value.startswith("#"):
            continue
        if not urlsplit(value).scheme or not urlsplit(value).netloc:
            print(f"Skipping invalid URL on line {line_number}: {value}", file=sys.stderr)
            continue
        urls.append(value)
    return urls


def safe_stem(index: int, url: str) -> str:
    parts = urlsplit(url)
    host = parts.netloc or "url"
    path = parts.path.strip("/") or "root"
    stem = re.sub(r"[^A-Za-z0-9._-]+", "_", f"{host}_{path}")
    return f"{index:05d}_{stem[:100]}"


def main() -> int:
    parser = argparse.ArgumentParser(description="Screenshot URLs and log main-document HTTP results.")
    parser.add_argument("--input", type=Path, default=Path("urls.txt"))
    parser.add_argument("--out", type=Path, default=Path("screenshots"))
    parser.add_argument("--log", type=Path, default=Path("results.csv"))
    parser.add_argument("--timeout-ms", type=int, default=30000)
    parser.add_argument("--full-page", action="store_true", help="Capture the full scrollable page.")
    parser.add_argument("--wait-until", choices=["load", "domcontentloaded", "networkidle", "commit"], default="load")
    args = parser.parse_args()

    if args.timeout_ms <= 0:
        parser.error("--timeout-ms must be a positive number")
    urls = read_urls(args.input)
    args.out.mkdir(parents=True, exist_ok=True)
    args.log.parent.mkdir(parents=True, exist_ok=True)

    # Start a fresh log for each run. Use a unique log path if prior run history matters.
    with args.log.open("w", newline="", encoding="utf-8") as log_file:
        writer = csv.DictWriter(log_file, fieldnames=FIELDS)
        writer.writeheader()
        log_file.flush()

        with sync_playwright() as p:
            browser = p.chromium.launch()
            try:
                for index, url in enumerate(urls, 1):
                    page = browser.new_page()
                    status = None
                    final_url = ""
                    screenshot_path = ""
                    navigation_error_type = ""
                    navigation_error = ""
                    screenshot_error = ""
                    capture_ok = False

                    try:
                        response = page.goto(
                            url,
                            wait_until=args.wait_until,
                            timeout=args.timeout_ms,
                        )
                        final_url = page.url
                        if response is not None:
                            status = response.status
                    except PlaywrightTimeoutError as exc:
                        navigation_error_type = "timeout"
                        navigation_error = str(exc)
                        final_url = page.url
                    except PlaywrightError as exc:
                        navigation_error_type = "navigation_error"
                        navigation_error = str(exc)
                        final_url = page.url

                    # Attempt a screenshot when a document rendered, including HTTP 4xx/5xx pages.
                    # After a navigation failure, a browser error page may still be screenshot-able;
                    # keeping that image is useful, but the null status and error remain explicit.
                    try:
                        filename = safe_stem(index, url) + ".png"
                        target = args.out / filename
                        page.screenshot(path=str(target), full_page=args.full_page, animations="disabled")
                        screenshot_path = str(target)
                        capture_ok = True
                    except PlaywrightError as exc:
                        screenshot_error = str(exc)

                    writer.writerow({
                        "input_url": url,
                        "final_url": final_url,
                        "http_status": "" if status is None else status,
                        "capture_ok": str(capture_ok).lower(),
                        "screenshot_path": screenshot_path,
                        "navigation_error_type": navigation_error_type,
                        "navigation_error": navigation_error,
                        "screenshot_error": screenshot_error,
                    })
                    log_file.flush()
                    page.close()
                    print(f"{index}/{len(urls)} status={status} captured={capture_ok} url={url}")
                browser.close()
            except BaseException:
                browser.close()
                raise

    return 0


if __name__ == "__main__":
    raise SystemExit(main())

Run a viewport capture (the browser’s current viewport) or add --full-page for the full scrollable document:

python bulk_screenshot.py --input urls.txt --out screenshots --log results.csv
python bulk_screenshot.py --input urls.txt --out screenshots --log results.csv --full-page --timeout-ms 45000

Each CSV record has the input URL, final URL, main-document status when available, screenshot success, path, and error fields. An empty http_status means no navigation response was observed; it does not mean status zero. The URL reader skips blank lines, comments, and malformed URLs. The sequence number in every filename prevents two URLs with similar or identical paths from overwriting each other.

4. Tune readiness, page size, and status policy

Choose a navigation readiness condition

  • load waits for the load event and is a reasonable default for ordinary pages.
  • domcontentloaded can return sooner when the screenshot does not need every load-event resource.
  • networkidle waits for network activity to settle, but analytics, polling, or long-lived connections may prevent it from becoming idle.
  • commit returns when the response is committed and can be useful for pages that keep loading, but the screenshot may show incomplete content.

There is no universal wait that guarantees a single-page application is visually ready. If a site has a known selector that signals the content you need, wait for it before the screenshot. For a fixed delay, add page.wait_for_timeout(1000) after navigation, but prefer a meaningful selector where possible. Apply the same readiness rule across a batch so results are comparable.

Decide how to treat HTTP statuses

The script logs all response statuses and continues, which is useful for audits. If a downstream report considers only 2xx successful, derive that classification from http_status rather than suppressing screenshots of error pages. Redirects usually lead to a final response; keeping both the original and final URL makes redirects visible. If you need every intermediate redirect status, capture request/response events and write a separate per-request log: the main navigation response alone is not a redirect-chain history.

Viewport or full page

Viewport screenshots are smaller and make page layouts easier to compare. Full-page screenshots include the scrollable document, which is useful for visual review but can produce very tall images and consume more memory. Lazy-loaded content may not appear unless the page scrolls it into view first. If complete lazy content matters, add a site-appropriate scrolling/readiness step and inspect representative output.

5. Keep durable, useful logs

  • Keep one row per input. This lets you identify skipped, failed, and completed URLs without inferring from image files.
  • Flush after each row. The example does this so a process interruption does not discard the entire batch’s log.
  • Separate runs. The example starts a fresh CSV. Choose a new log filename per run or change it to append with a run identifier if you need history.
  • Preserve raw errors. Store exception text for diagnosis, but avoid putting secrets from authenticated URLs or headers into logs.
  • Keep artifacts together. Archive the CSV and screenshots together, and use a stable run date or job ID in the containing directory.
  • Make retries explicit. Retry transient network failures selectively, and record attempt count or a separate attempt log. Do not silently replace the first failure with a later success.

CSV is easy to inspect and import. JSON Lines is a good alternative when each record needs nested data such as redirect chains or multiple attempts; write one JSON object per line and flush after each record.

6. cURL, Python, and Node.js options for hosted captures

A plain HTTP request to a web page downloads its response body; it does not render the page as a browser screenshot. If you want hosted rendering instead of managing a browser, ScreenshotNeo accepts a URL and returns a screenshot or PDF. The examples below save the returned image bytes; check the response headers to distinguish the page verdict and billing outcome, and keep your own URL-to-file log in the batch runner. See the ScreenshotNeo API documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp \
  -D shot.headers

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
print("page verdict:", r.headers.get("X-Page-Verdict"))
print("billed:", r.headers.get("X-Billed"))

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
console.log('page verdict:', res.headers.get('x-page-verdict'));
console.log('billed:', res.headers.get('x-billed'));

For a bulk job, wrap the API request in the same per-URL loop and persist a CSV row after each response. Record the input URL, screenshot path, response outcome headers, request-level HTTP status, and exception. A screenshot API’s request status and its verdict about the target page are different fields; neither should be confused with the target site’s HTTP status unless the service explicitly provides that status. ScreenshotNeo also offers bulk capture for up to 100 URLs per call, async jobs with signed webhooks, and a usage API; consult its documentation for the current request and response fields.

7. Troubleshooting

Symptom Likely cause Fix
CSV status is empty Navigation timed out, DNS failed, or no main-document response was returned. Read navigation_error_type and navigation_error; check the URL and network access, then retry transient failures deliberately.
Status is 404 or 500 but capture is true The server returned an error page that the browser rendered. This is a valid and useful result. Keep the status and screenshot as separate facts.
Screenshot is blank or incomplete The page was captured before its client-side content was ready, or the site requires authentication. Wait for a content selector, configure the needed session, or choose a suitable readiness condition. Inspect a sample before processing the full list.
Images or lower-page content are missing Images are lazy-loaded or only loaded near the viewport. Use full-page capture and, if necessary, scroll the page in a site-appropriate way before taking the screenshot.
Navigation hangs on some URLs A site never reaches the selected event, or the timeout is too short for that page. Try domcontentloaded or a selector-based readiness check, and set a bounded timeout. Avoid using network idle blindly on pages with continual traffic.
Later URLs overwrite earlier images Filenames were based only on a path or hostname that repeated. Include a sequence number or unique input identifier, as the script does.
Batch stops after one problematic page An exception escaped the per-URL handling or the browser process failed. Keep per-URL exceptions caught, flush each row, and split very large runs into smaller restartable batches.
Browser executable or library error Playwright’s browser was not installed in the active environment, or runtime dependencies are missing. Run python -m playwright install chromium in the same environment and install the OS dependencies required by your deployment image.

8. Performance, reliability, and cost

The example uses one browser page at a time. This is slower than parallel capture but limits memory use, simplifies logs, and reduces load on target sites. If you add concurrency, start small, isolate each URL’s output and result row, and cap retries. Heavy pages and full-page images can consume substantial memory; avoid opening an unbounded number of pages in one process.

For reliability, set finite navigation timeouts, persist progress incrementally, and make each run restartable. Keep browser and dependency versions controlled in scheduled jobs. Test a small sample containing redirects, a known 404, a slow page, and a page with dynamic content before scaling up. The dossier provides no independent benchmark or universal throughput figure, so estimate duration from your own representative pages rather than assuming a fixed rate.

Self-managed Playwright has no per-screenshot API charge, but you own compute, browser maintenance, storage, and operational work. Full-page images and retained history increase storage needs. Hosted capture exchanges browser operations for service pricing and service-specific limits; verify current price, retention, supported authentication, and batch limits directly before committing. ScreenshotNeo’s listed plans are 1,000 screenshots/month free with no card, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Only clean shots are billed, while bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Every feature is on every plan.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns an image or PDF. For a single capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners are accepted like a visitor and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing outcome. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. It also supports bulk capture for up to 100 URLs per call. See the API docs and start with 1,000 free screenshots a month, no card.

FAQ

Does a 404 mean the screenshot failed?

No. A 404 is an HTTP response status. If the page rendered and the screenshot was saved, the capture succeeded while the target returned an error status.

Should I save screenshots of 4xx and 5xx responses?

Usually yes when the goal is auditing or debugging. The error page is often the evidence you need; classify it from the status field separately.

How can I tell whether a failure happened before the server responded?

A blank status together with a navigation error indicates that the script did not observe a main-document response. The error text helps distinguish timeout, DNS, and other browser failures.

Should I use CSV or JSON?

CSV works well for a flat one-row-per-URL report. Use JSON Lines if you need nested redirect chains, retries, or multiple capture attempts per URL.

Sources