ScreenshotNeo

BlogHow-to

How to Schedule Daily Screenshots for a List of Website URLs

Schedule a daily screenshot run for many URLs, save the results, and handle rate limits, retries, and missing captures.

By the ScreenshotNeo team4 October 202613 min read

To schedule daily screenshots for a list of website URLs, set up two separate parts: a daily trigger and a workflow that processes each URL, stores each result, and records failures. For a no-code setup, Urlbox documents a Zapier schedule workflow and an example that saves screenshots to Google Drive. For a developer workflow, run a scheduled job, queue URLs in batches, call a screenshot API, and persist the images and run status. ScreenshotOne documents a bulk endpoint for multiple URLs; its bulk requests share the regular one-minute request bucket, so pace the work and retry failures.

For a managed URL-list workflow, Urlbox CaptureDeck describes pasting URLs or importing a CSV and scheduling recurring captures. Confirm current destinations, retention, volume limits, and plan terms before depending on them. If you need to control scheduling, retries, storage, and review yourself, use the API workflow below.

1. Choose the workflow that fits your list

Approach How URLs enter Scheduling and processing Best fit
ScreenshotNeo API Your application or job supplies URLs Use your scheduler and queue; capture individual URLs or use bulk capture for up to 100 URLs per call Developer-owned automation with control over capture options and output handling
Urlbox with Zapier Workflow inputs such as a prepared URL list Urlbox documents a Zapier scheduling route for hourly, daily, weekly, or specific-time captures A no-code recurring workflow; its guide illustrates saving results to Google Drive
Urlbox CaptureDeck Paste a URL list or import a CSV Product page describes recurring scheduled captures A managed URL-list workflow with reports and downloadable archives
ScreenshotOne bulk API Send multiple URLs to its bulk endpoint Pair the endpoint with your own daily trigger, queue, and storage A code-driven multi-URL workflow using ScreenshotOne

ScreenshotNeo is the first API to try if you want clean screenshots: cookie and consent banners, newsletter popups, and chat widgets are removed before capture, and only clean shots are billed. Its plans include the same features, with a free tier and paid plans starting at $5. See ScreenshotNeo and its API documentation.

Sources: Urlbox’s scheduled screenshot guide, CaptureDeck, ScreenshotOne’s bulk API documentation.

2. Define the daily run before writing code

Decide what counts as a successful run and where its output belongs. A trigger firing successfully does not mean every site rendered successfully. Track each URL separately so one timeout or bot check does not hide the status of the rest.

  1. Maintain a source of URLs. Use a CSV, a database table, or a version-controlled list. Keep a stable identifier for each URL, such as a short site name, to make output files easy to find.
  2. Set a schedule and timezone. Pick a time when the destination and downstream jobs are ready. Store the intended timezone explicitly; schedulers may otherwise interpret schedules in their own default timezone.
  3. Choose capture settings. Set the output format, viewport, and full-page behavior based on what you need to compare. Keep settings consistent between runs unless you intentionally version them.
  4. Choose a destination. Save images to object storage, a shared drive, or another system your users can access. Decide naming, access controls, and retention. Urlbox’s guide gives Google Drive as an example destination; check current product support and terms before relying on a specific destination.
  5. Record run state. Keep a run ID, URL, start time, result status, output location, and error details. This lets you find missing captures and retry only failed URLs.

3. Build a reliable API workflow

A typical developer setup is a daily scheduler, a durable queue, workers that call the capture API, storage for successful outputs, and a small run-status table. The scheduler should enqueue work rather than try to capture a large list in one long-running task. A durable queue helps work survive process restarts and supports multiple workers.

Suggested processing sequence

  1. Load active URLs and assign the daily run an ID.
  2. Enqueue one job per URL, or enqueue bounded batches if your API supports bulk requests.
  3. Have workers capture URLs with a bounded concurrency. Pace requests according to the service’s current limits.
  4. Store successful image bytes or returned links, using a run ID and URL identifier in the object key.
  5. Record success or failure per URL. Retry transient failures with backoff; do not retry a permanent input error forever.
  6. At the end of the run, report successes, failures, and URLs that were skipped or disabled.

ScreenshotOne recommends batching URLs, retrying failed requests, and checking its usage endpoint’s concurrency remaining and reset values. Its bulk requests share the same one-minute request bucket as regular screenshot requests, so bulk should not be treated as unlimited throughput. See its guide to screenshots of multiple URLs and bulk endpoint documentation.

Example: schedule and queue with your own infrastructure

The following Python example shows the core of a daily worker for a CSV list using ScreenshotNeo’s single-URL API. It downloads each successful image to a local run directory and logs failures per URL. Schedule this script with your operating system or cloud scheduler, and replace local storage with your durable destination for production. Install the dependency with python -m pip install requests. Set SCREENSHOTNEO_API_KEY in the job environment and create urls.csv with a url column.

import csv
import os
import time
from pathlib import Path
from urllib.parse import urlparse

import requests

API_KEY = os.environ["SCREENSHOTNEO_API_KEY"]
INPUT_CSV = Path("urls.csv")
OUTPUT_DIR = Path("captures")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)


def safe_name(url: str) -> str:
    parsed = urlparse(url)
    host = parsed.netloc or "page"
    path = parsed.path.strip("/").replace("/", "_") or "root"
    return f"{host}_{path}".replace(":", "_")


with INPUT_CSV.open(newline="", encoding="utf-8") as source:
    urls = [row["url"].strip() for row in csv.DictReader(source) if row.get("url", "").strip()]

for index, url in enumerate(urls, start=1):
    if not url.startswith(("http://", "https://")):
        print(f"SKIP invalid URL: {url}")
        continue

    for attempt in range(3):
        try:
            response = requests.get(
                "https://api.screenshotneo.com/v1/shot",
                params={"access_key": API_KEY, "url": url},
                timeout=90,
            )
            response.raise_for_status()
            output = OUTPUT_DIR / f"{index:04d}_{safe_name(url)}.png"
            output.write_bytes(response.content)
            print(f"OK {url} -> {output}")
            break
        except requests.RequestException as exc:
            if attempt == 2:
                print(f"FAILED {url}: {exc}")
            else:
                time.sleep(2 ** attempt)

This minimal worker stores the response body after an HTTP success. For production, also inspect ScreenshotNeo’s response headers: X-Page-Verdict and X-Billed indicate the page verdict and billing status. Keep these values with the run record so you can distinguish a clean capture from a bot check, blank page, timeout, failed load, or cache hit. ScreenshotNeo bills only clean shots; those other outcomes and cache hits cost nothing.

Scheduling examples

The exact scheduler command depends on the host. For a Linux machine using cron, this example runs daily at 02:15 UTC and writes a log. Confirm that Python and the script path are correct for the cron environment:

15 2 * * * cd /srv/site-captures && /usr/bin/python3 daily_capture.py >> /var/log/site-captures.log 2>&1

For a cloud scheduler, use its daily cron expression or recurring schedule to invoke a short job that enqueues URLs. Keep the scheduled entry point quick and let queue workers do the capture work. Avoid overlapping runs unless your storage keys and queue logic are designed for them.

cURL, Python, and Node.js single-capture calls

These are complete request examples for one URL. Put the API key in a secret store or environment variable in deployed jobs; do not commit it to source control. The ScreenshotNeo docs describe the API and available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

In Node.js, the final line uses Bun’s file writer. With Node.js, save the response using the filesystem module:

import { writeFile } from 'node:fs/promises';

const q = new URLSearchParams({ access_key: process.env.SCREENSHOTNEO_API_KEY, url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

4. Handle multiple URLs, rate limits, and retries

For a long list, do not start an unbounded request for every URL at once. Use bounded concurrency and check the provider’s current usage limits. ScreenshotOne specifically advises batching and retrying, and says its bulk requests use the same one-minute request bucket as regular screenshot calls. Its docs also expose usage values for concurrency remaining and reset; use those values to pace work instead of assuming the bulk endpoint bypasses limits.

  • Retry transient failures: timeouts and temporary server errors may succeed on another attempt. Use exponential backoff and a maximum attempt count.
  • Do not retry invalid input unchanged: malformed URLs, unsupported settings, and authorization errors need correction first.
  • Make retries safe: use stable output keys or a run ID so a retry does not create confusing duplicate files.
  • Prevent duplicate daily runs: use a lock or unique run record if a scheduler may retry its trigger.
  • Limit concurrency: tune worker count to your account limits, average page load time, and downstream storage capacity.
  • Review incomplete runs: compare the number of requested URLs with the terminal success, failure, and skipped counts.

5. Choose capture settings for comparable results

Daily screenshots are only useful for comparison when important rendering inputs stay consistent. The website itself may change, but viewport, device scale, color mode, and wait behavior should be held steady unless those are the variables you intend to track.

Setting When it matters Operational note
Viewport and device preset Layout shifts at mobile or desktop breakpoints Use the same dimensions for every run; ScreenshotNeo offers 12 device presets and custom viewports
Full-page capture You need content below the fold ScreenshotNeo loads lazy images for full-page capture; long pages can take more time and produce larger files
Element selector You need a stable component rather than the whole page Capture one CSS-selected element to reduce irrelevant page content
Wait condition Pages render content asynchronously Wait for a selector, a delay, or network idle; choose a condition that reflects the content you need
Interactions and exclusions A dialog must be dismissed or an area omitted ScreenshotNeo supports clicking an element and hiding selectors; each consent-cleanup step can also be disabled
Resource handling Pages load ads or trackers that slow captures Options include blocking ads, trackers, requests, or resource types; blocking a needed resource can change the page
Output format and size Storage cost or downstream image handling matters PNG, JPEG, and WebP are available; resizing can reduce stored bytes
Authenticated or localized pages Content varies by account or region Custom headers, cookies, user agent, Authorization, timezone, and geolocation are supported; protect captures and credentials

ScreenshotNeo also supports dark mode, retina scale, transparent backgrounds, custom CSS and JavaScript, custom headers and cookies, caching with a chosen TTL, signed links for public <img> tags, PDF output with paper size, margins, orientation, and page ranges, async jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI spec. These options can help adapt the workflow, but document the chosen settings with each run so later comparisons remain interpretable.

6. Storage, retention, performance, and cost

Storage and file organization

Use a predictable key such as daily/YYYY-MM-DD/run-id/site-id.webp. Avoid putting credentials or sensitive query parameters in filenames. Store a small manifest alongside the images with the requested URL, capture time, settings, verdict, and output location. Set access controls to match the page contents, especially for authenticated captures. Decide how long to retain images and manifests before the first scheduled run.

Performance and reliability

  • Break large lists into batches and use bounded workers so one scheduled run cannot overwhelm API limits or storage.
  • Use a durable queue if runs must survive restarts or be processed by multiple workers.
  • Set request timeouts long enough for slow pages, but keep retries finite so one URL cannot stall the whole run.
  • Use an explicit wait condition for dynamic pages. A fixed delay that is too short captures incomplete content; one that is too long wastes time.
  • Record missing and failed captures, and alert on repeated failures or a large change in success count.
  • Use caching only when repeated captures are allowed to reuse a prior result. For a daily change-monitoring run, ensure the cache TTL does not cause yesterday’s image to be returned as today’s capture.

Cost

Estimate volume as active URLs multiplied by scheduled runs, then account for retries and any additional viewport or format variants. Verify third-party API quotas and pricing from current vendor account terms; the cited research does not establish current prices or a neutral performance comparison. Storage and retention costs depend on image size, run frequency, and how long you keep files.

ScreenshotNeo’s listed plans are Free: 1,000 shots per month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed; bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Check the current ScreenshotNeo site for plan details.

7. No-code: schedule captures without maintaining workers

  1. Prepare the URLs as a list or CSV. CaptureDeck describes both pasted URLs and CSV import; Urlbox’s docs also mention lists from Google Sheets and Airtable.
  2. For the Zapier route, create the Urlbox and Zapier accounts required by the documented workflow. Add a schedule trigger for the desired daily time and connect it to the screenshot action.
  3. Choose a destination. Urlbox’s guide illustrates Google Drive. Verify that your account currently supports the destination and the file naming and retention you require.
  4. Run a small representative set first. Check that each page loads as expected, the output has the correct viewport, and the workflow reports failures clearly.
  5. Confirm the daily volume against the current plan and integration limits, then enable the recurring schedule.

References: Urlbox’s guide to scheduling screenshots, CaptureDeck’s URL-list workflow, and Urlbox documentation. Integrations, destinations, retention, and pricing can change, so verify details in the current product account.

8. Troubleshooting scheduled runs

Symptom Likely cause What to do
No run started Scheduler timezone, disabled trigger, or deployment path is wrong Check scheduler history, timezone, script path, environment variables, and the job’s service account
Some URLs are missing A worker exited early, the queue lost jobs, or the run did not wait for completion Compare enqueued and terminal job counts; use a durable queue and record per-URL state
Rate limit or concurrency errors Requests exceeded the current request bucket or concurrency allowance Reduce worker concurrency, pace batches, and check the provider usage endpoint and reset values where available
Blank or incomplete page Content is delayed, requires interaction, or did not load successfully Choose an appropriate wait condition, selector, or click action; inspect the page verdict and retry transient failures
CAPTCHA or bot check in output The destination is challenging automated visits Record the outcome rather than treating it as a valid page capture. ScreenshotNeo reports page verdict and billing status in response headers and does not bill bot checks or CAPTCHAs
Image from an earlier day A cache served an earlier capture or the output key was overwritten incorrectly Review cache TTL and use date- and run-specific keys when each daily result must be retained
Images are unexpectedly large Full-page, high retina scale, or an unnecessarily large format increased bytes Capture only the needed area, adjust scale, choose a suitable format, or resize output
Access denied or wrong page The page requires authentication, cookies, headers, or a particular region Supply supported auth and locale inputs securely, and verify that the account is authorized to capture the page
Retries create duplicate or confusing files Output names are based only on timestamps or random values Use deterministic URL identifiers plus run IDs and make the run trigger idempotent

Or skip the browser setup

ScreenshotNeo turns a URL into an image or PDF with one GET request. Its cookie and consent handling accepts the banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are never billed, and responses include X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The API also supports bulk capture for up to 100 URLs per call, which you can invoke from your daily scheduler.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. One thousand screenshots a month are free with no card, and paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo and schedule your first URL list.

9. Frequently asked questions

Can I schedule screenshots of URLs from a Google Sheet?

Urlbox’s documentation mentions URL lists from Google Sheets and Airtable. For a custom workflow, read the sheet in the scheduled job and enqueue each active URL.

Does a bulk screenshot endpoint run every day automatically?

A bulk endpoint processes multiple URLs per request; it does not inherently define the daily schedule. The cited ScreenshotOne docs describe the bulk endpoint, so pair it with a separate scheduler.

How do I know whether a scheduled screenshot is fresh?

Store the capture time and run ID with each output, use date-specific storage keys, and check that your cache settings cannot return a previous day’s image.

A scheduled image alone does not establish provenance or legal sufficiency. If the capture is evidence, define the required chain of custody and retention separately with the appropriate experts.