ScreenshotNeo

BlogHow-to

How to Schedule Screenshots of Multiple URLs from a CSV File

Build a scheduled job that reads URLs from a CSV, captures each page, and reports failures. Includes Playwright, shot-scraper, and GitHub Actions.

By the ScreenshotNeo team4 October 202611 min read

To schedule screenshots of multiple URLs from a CSV file, use a script to read and validate one URL per row, capture each page with a browser tool such as Playwright, and run the script on a schedule with GitHub Actions. Keep a per-URL manifest so failures are visible, choose stable output names, and pin the browser environment if you need to compare screenshots over time.

This guide uses Python, Playwright, and GitHub Actions. It also shows a batch-oriented shot-scraper option. CSV parsing is your script’s job: neither Playwright nor shot-scraper should be assumed to ingest an arbitrary CSV directly. For scheduling, the workflow invokes the same capture script on every run. Playwright screenshot documentation, shot-scraper documentation, and GitHub Actions schedule documentation describe the underlying tools.

1. Define the CSV and output policy

Use a header named url and one absolute HTTP or HTTPS URL in each non-empty row:

url
https://example.com/
https://example.org/pricing
https://example.net/docs

Decide these details before scheduling:

  • Validation: report blank or malformed rows instead of attempting navigation.
  • Names: include a row index and hostname so paths on one domain do not overwrite each other.
  • History: either overwrite the latest capture or retain a dated archive. The example below retains dated files.
  • Scope: viewport captures are bounded; full-page captures include scrollable content and can be much taller.
  • Failures: record each URL’s result and make the job fail if any row fails, while still attempting the rest.

Do not put private URLs, credentials, or authenticated screenshots in a public repository. Treat screenshot output as potentially sensitive. Store credentials in the scheduler’s secret store and restrict access to the artifacts or repository that receives captures.

2. Create a runnable Playwright capture script

Install Python dependencies locally or in CI:

python -m pip install playwright
python -m playwright install --with-deps chromium

Save the following as capture_csv.py. It reads the CSV with Python’s standard library, checks absolute HTTP(S) URLs, attempts every valid row, writes PNG files and a JSON manifest, and returns a failing exit code if any row failed.

import csv
import json
import re
import sys
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import sync_playwright

CSV_PATH = Path("urls.csv")
OUTPUT_DIR = Path("screenshots")
FULL_PAGE = True
VIEWPORT = {"width": 1440, "height": 900}
NAVIGATION_TIMEOUT_MS = 45_000


def valid_http_url(value: str) -> bool:
    parsed = urlparse(value)
    return parsed.scheme in {"http", "https"} and bool(parsed.netloc)


def safe_part(value: str) -> str:
    cleaned = re.sub(r"[^A-Za-z0-9_-]+", "-", value).strip("-_")
    return cleaned[:80] or "page"


def main() -> int:
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    run_id = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
    manifest = []
    failures = 0

    with CSV_PATH.open("r", newline="", encoding="utf-8-sig") as source:
        reader = csv.DictReader(source)
        if not reader.fieldnames or "url" not in reader.fieldnames:
            raise ValueError("CSV must contain a header named 'url'")
        rows = list(reader)

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch()
        for row_number, row in enumerate(rows, start=2):
            raw_url = (row.get("url") or "").strip()
            item = {
                "row": row_number,
                "url": raw_url,
                "captured_at": datetime.now(timezone.utc).isoformat(),
                "status": "failed",
            }
            if not raw_url:
                item["error"] = "Blank URL"
                manifest.append(item)
                failures += 1
                continue
            if not valid_http_url(raw_url):
                item["error"] = "Expected an absolute HTTP(S) URL"
                manifest.append(item)
                failures += 1
                continue

            host = safe_part(urlparse(raw_url).hostname or "page")
            output_path = OUTPUT_DIR / f"{run_id}-{row_number:04d}-{host}.png"
            page = browser.new_page(viewport=VIEWPORT, device_scale_factor=1)
            page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
            try:
                response = page.goto(raw_url, wait_until="networkidle")
                # A response can be absent for some navigation types. Record it when present.
                item["http_status"] = response.status if response else None
                page.screenshot(path=str(output_path), full_page=FULL_PAGE)
                item.update({"status": "ok", "output": str(output_path)})
            except Exception as error:
                item["error"] = str(error)
                failures += 1
            finally:
                page.close()
            manifest.append(item)
        browser.close()

    manifest_path = OUTPUT_DIR / f"{run_id}-manifest.json"
    manifest_path.write_text(json.dumps(manifest, indent=2), encoding="utf-8")
    print(f"Wrote {len(manifest)} row results to {manifest_path}; failures: {failures}")
    return 1 if failures else 0


if __name__ == "__main__":
    sys.exit(main())

Run it with python capture_csv.py from the directory containing urls.csv. The script waits for networkidle, which is useful for pages that continue loading resources after the initial document response. Some sites keep connections open or poll continuously, so network idle may never occur. For those pages, use wait_until="domcontentloaded" or wait_until="load", then add a site-specific readiness condition or a bounded delay before the screenshot. A fixed delay is simple but may waste time or still be too short.

Useful Playwright choices

  • full_page=True captures the full scrollable document; set it to False for a viewport-only image.
  • Set viewport to the CSS pixel dimensions you need. Keep it fixed between runs for comparison.
  • Set device_scale_factor to 2 for a higher-density image, at the cost of larger files and more rendering work.
  • Use a separate browser context or page per URL when pages need isolated cookies or storage. This sample creates a fresh page per row.
  • For authentication, configure a context with the required state or credentials. Do not write secrets or storage state into public artifacts.
  • For dynamically rendered pages, wait for a meaningful selector with Playwright’s locator APIs before capturing rather than assuming the navigation event means the page is ready.

3. Schedule it with GitHub Actions

Commit capture_csv.py and urls.csv, then create .github/workflows/screenshots.yml:

name: Scheduled screenshots

on:
  workflow_dispatch:
  schedule:
    - cron: "17 8 * * 1-5"

permissions:
  contents: write

jobs:
  capture:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: python -m pip install playwright
      - run: python -m playwright install --with-deps chromium
      - run: python capture_csv.py
      - name: Commit screenshots and manifest
        if: success()
        run: |
          git config user.name "github-actions[bot]"
          git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
          git add screenshots/
          git diff --cached --quiet || git commit -m "Update scheduled screenshots"
          git push

The cron expression uses POSIX cron fields and runs at 08:17 UTC Monday through Friday. GitHub’s scheduled workflows use UTC by default, run against the latest commit on the default branch, and have a minimum supported interval of five minutes. Scheduled runs can be delayed during high load, so do not use this mechanism when an exact-to-the-minute execution time is essential. Keep the workflow on the default branch.

The example commits screenshots only when the capture script exits successfully. That avoids presenting a partial batch as a complete scheduled update; the failed run still leaves its logs available in Actions. If you want to retain partial captures for diagnosis, change the commit step to run with if: always() and explicitly account for the privacy and retention of those files. For sensitive or large output, use an access-controlled artifact or storage destination instead of committing images.

4. Batch alternative: shot-scraper

shot-scraper packages common screenshot tasks behind a CLI and supports multiple targets through configuration. It is a good fit when a declarative batch file is enough; use Playwright directly when individual pages need custom authentication, interactions, readiness logic, or error handling. The docs’ multi-shot configuration is the relevant mechanism. Since the input here is CSV, convert validated rows into that documented configuration or write a wrapper that invokes shot-scraper once per URL. Choose explicit output paths when you need predictable names; its default naming behavior can add numbered suffixes when a file already exists.

A basic command-line capture of one URL looks like this:

python -m pip install shot-scraper
shot-scraper install
shot-scraper https://example.com/ -o screenshots/example.png

For a CSV batch, retain the same validation and manifest responsibilities as in the Playwright example. A scheduler can call the wrapper or configured batch command on each run. Consult the shot-scraper documentation for the current configuration shape and CLI options, and pin the installed version in your environment once you settle on it.

5. Or skip the browser setup

ScreenshotNeo provides a screenshot API and MCP server. Your CSV job can make one GET request per URL, save the response, and record the response headers alongside the output. Its parameter names also work with those used by other screenshot APIs, which can make an existing integration easier to switch. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/ -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Use the API from the same validated CSV loop and derive a distinct filename for each row. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response includes X-Page-Verdict and X-Billed headers to identify the result. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, or another MCP client.

ScreenshotNeo includes full-page capture with lazy images loaded, CSS selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, click and hide selectors, wait conditions, request blocking, custom headers, cookies, user agent and authorization, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI spec. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, with yearly billing giving two months free. Every feature is on every plan. Visit ScreenshotNeo for product details, then sign up for 1,000 free screenshots a month with no card.

6. Reliability, performance, and cost

Make comparisons meaningful

Screenshot output can vary with operating system, browser version, browser settings, hardware, power conditions, and headless mode. Pin or otherwise keep the runner, Python version, browser version, viewport, and device scale factor stable when comparing images over time. Treat a pixel difference as evidence to inspect, not automatic proof that the website changed. Record run time and tool versions in the manifest if the captures are used for audits or visual review.

Plan batch runtime

Sequential navigation is easy to reason about and avoids launching too much browser work at once, but total runtime grows with the number of pages and each page’s load time. Full-page captures, high device scale factors, and very long pages increase rendering time and image size. Start with sequential captures; add bounded concurrency only after measuring your workload and ensuring the runner has enough memory and CPU. Put a timeout on navigation and avoid unbounded waits.

Track partial results and retries

Use the manifest to distinguish invalid input, navigation failure, and successful output. A transient timeout may merit a limited retry, but retry only the affected URL and record each attempt so the batch does not conceal instability. A page that blocks automation, requires login, changes its layout, or never reaches the chosen readiness condition needs site-specific handling. Do not silently mark a partial run successful.

Understand cost

With self-hosted or CI Playwright, account for runner time, browser installation, image storage, and maintenance of the browser environment; actual cost depends on your runner and retention choices. Full-page and high-resolution images consume more storage. GitHub Actions execution and repository storage are governed by your account and repository plan; check the current GitHub plan details before relying on a particular allowance. A managed screenshot API trades browser maintenance for per-plan usage. ScreenshotNeo’s free tier is 1,000 shots monthly with no card; paid tiers are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Only clean shots are billed, according to the product details above.

7. Troubleshooting

Symptom Likely cause Fix
Script says the CSV has no url column Header is missing, misspelled, or has unexpected whitespace. Use the exact header url. Inspect the first row and save as UTF-8; the example handles a UTF-8 byte-order mark.
Rows fail URL validation Values are relative paths, missing a scheme, or blank. Provide absolute https:// or http:// URLs. Correct or remove blank rows.
Navigation times out The host is slow, blocks automation, or never becomes network idle. Use a realistic bounded timeout; try domcontentloaded or load and wait for the page-specific content you need.
Screenshot is blank or incomplete The capture ran before client-side content appeared, or content loads only after scrolling or interaction. Wait for a meaningful selector, trigger required interaction, or scroll as the page requires. Check the manifest and output before treating the run as valid.
Images differ between scheduled runs Browser, operating system, viewport, scale factor, fonts, or rendering conditions changed; the site itself may also have changed. Stabilize the runtime and capture settings. Review differences instead of assuming every pixel change is a site regression.
Same-site captures overwrite one another Output names use only the hostname. Include row number or a URL-derived path slug, plus a run identifier if archiving.
Scheduled workflow does not run at the expected time Schedule is evaluated in UTC, workflow is not on the default branch, or GitHub delayed a run under load. Check the cron fields and default branch. Allow for scheduling delays; use a different scheduler if exact timing is mandatory.
Commit step reports nothing to commit Files are unchanged or no successful output was written. The sample checks staged changes before committing. Inspect the manifest and Actions logs to tell these cases apart.
Repository grows quickly Each run archives full-page images indefinitely. Use a retention policy, overwrite latest captures, or store images in an access-controlled artifact or object store.

8. Frequently asked questions

Can I include extra CSV columns?

Yes. The script currently uses only url; additional columns can carry labels or per-page settings. Validate any setting that influences navigation or output before using it.

Should I capture the full page?

Choose full-page capture when below-the-fold content matters. Use viewport capture for a consistent, bounded view or when full-page images become unwieldy. Test long pages that lazy-load content.

Can the job capture pages behind a login?

Yes, if you provide browser state or credentials to the capture process and the target permits that access. Keep secrets out of CSV files, logs, committed screenshots, and public manifests.

Can I schedule more often than once per day?

GitHub Actions supports schedule expressions at a minimum interval of once every five minutes, though actual starts may be delayed. For high-frequency capture, consider runtime, output volume, and the target site’s access limits.