ScreenshotNeo

BlogHow-to

How to Archive Website Screenshots Automatically with an API

Build a repeatable screenshot archive with scheduled captures, useful metadata, retries, and durable storage you control.

By the ScreenshotNeo team4 October 202612 min read

To archive website screenshots automatically, schedule a capture for each URL, request an image with explicit viewport and format settings, save the returned file to durable storage, and record metadata that lets you find and interpret it later. A managed archive service can schedule captures and retain history for you; a capture API workflow can export files into storage you control.

This guide covers both patterns, a runnable Python pipeline that saves captures to local disk, production design choices, failure handling, retention, and the questions to settle before relying on an archive.

1. Define the archive job

Start with the purpose of the archive. A record used to review visual changes has different requirements from one used to preserve a broad set of public pages over time. Decide these details before choosing an API:

  • URLs: Which pages matter, and how will you maintain the URL list?
  • Schedule: How often does each page change, and how soon would you need to notice a change? Choose a frequency that fits the purpose; capturing more often is not automatically more useful.
  • Viewport: Which screen size or device view should the archive represent?
  • Page extent: Is the initial viewport enough, or do you need a full-page capture? Full-page captures may need extra time for long pages and lazy-loaded content.
  • Output: Choose an image format supported by the provider and suitable for your downstream workflow.
  • Retention and access: How long must files remain available, and which services or people need access?
  • Failure policy: Decide how to alert on repeated failure and how to represent a missed capture in your records.

Record the original URL even if you also create a normalized URL or stable identifier for filenames. Query parameters, redirects, and URL changes can affect what page a capture represents.

2. Choose a capture pattern

Managed monitor and archive

With this pattern, create a monitor for a URL and let the service schedule captures. The API can then be used to trigger a capture, list historical snapshots, and download a file. Snapshot Archive documents HTTPS REST endpoints for creating monitors, triggering captures, retrieving snapshots, and downloading files; its API reference describes JSON request bodies, bearer-token authentication, and access for Starter plans and above. Its allowed capture intervals vary by plan, so verify current limits before designing a schedule. Snapshot Archive API documentation and scheduling documentation.

Capture API plus your own storage

In this pattern, a scheduler or existing job runner calls a screenshot endpoint, receives a file or job result, and exports the file to storage you select. Add Screenshots documents single and bulk screenshot workflows, including uploads to repositories such as S3 and Azure Blob Storage. ScreenshotAPI.net advertises scheduled captures and private cloud bucket storage. Treat these as vendor-described capabilities and verify current behavior, supported formats, and limits in the provider’s documentation. Add Screenshots API documentation and ScreenshotAPI.net.

Decision Managed archive API plus your storage
Scheduling Service schedules according to its monitor and plan options. Your scheduler or workflow decides when each URL is due.
Historical retrieval Use the service’s snapshot listing and download endpoints, where available. Retrieve from the storage system and metadata index you maintain.
Retention Check provider retention windows and export options. Set retention rules in the chosen storage and any backups.
Operational ownership Less scheduling and archive infrastructure to maintain; still handle credentials, monitoring, and export needs. More control over file organization and retention, with more responsibility for scheduling, retries, storage, and monitoring.
Costs to verify Plan tier, capture interval, API access, retention, and overages. Capture requests, storage, transfer, job-runner usage, and operational effort.

There is no neutral price, reliability, or performance comparison established by the available vendor documentation. Estimate your own expected capture volume and check each provider’s current limits, retention terms, and pricing before committing.

3. Build a recurring capture pipeline

A robust archive is a workflow around a screenshot request, not just a request that happens to run repeatedly. The following steps are general implementation guidance; providers differ in whether captures are immediate or asynchronous.

  1. Schedule: A cron job, queue, workflow, or managed monitor identifies URLs that are due.
  2. Capture: Send the URL with explicit viewport, page extent, and output settings.
  3. Complete: For a synchronous endpoint, check the response status and file type. For an asynchronous API, poll or handle its documented completion mechanism before treating the capture as ready.
  4. Persist: Write the returned bytes to durable storage using a deterministic key or path.
  5. Record metadata: Store the original URL, capture timestamp, settings, output location, and relevant response status alongside the file.
  6. Recover visibly: Retry transient failures according to a bounded policy, alert on repeated failures, and record a failed or missed run instead of silently treating it as archived.

Runnable local example in Python

This example uses ScreenshotNeo’s one-call screenshot endpoint and writes a WebP file and JSON sidecar for each URL. It demonstrates local retention; replace the local destination with your object-storage client when deploying a shared archive. Create an API key and review the ScreenshotNeo API documentation for available settings and response headers.

import hashlib
import json
import os
import re
import time
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

import requests

API_URL = "https://api.screenshotneo.com/v1/shot"
API_KEY = os.environ["SCREENSHOTNEO_API_KEY"]
ARCHIVE_DIR = Path("website-archive")
URLS = ["https://example.com", "https://www.python.org"]


def safe_name(url):
    parsed = urlparse(url)
    base = re.sub(r"[^a-zA-Z0-9_-]+", "-", parsed.netloc + parsed.path).strip("-")
    return base[:100] or "page"


def capture(url):
    captured_at = datetime.now(timezone.utc)
    response = requests.get(
        API_URL,
        params={"access_key": API_KEY, "url": url},
        timeout=90,
    )
    response.raise_for_status()

    # Refuse to store an unexpected response as an image.
    content_type = response.headers.get("Content-Type", "").lower()
    if not content_type.startswith("image/"):
        raise ValueError(f"Expected an image response, got {content_type!r}")

    stamp = captured_at.strftime("%Y%m%dT%H%M%SZ")
    folder = ARCHIVE_DIR / safe_name(url)
    folder.mkdir(parents=True, exist_ok=True)
    image_path = folder / f"{stamp}.webp"
    image_path.write_bytes(response.content)

    metadata = {
        "url": url,
        "captured_at": captured_at.isoformat(),
        "settings": {"output": "webp", "capture": "default"},
        "content_type": content_type,
        "bytes": len(response.content),
        "sha256": hashlib.sha256(response.content).hexdigest(),
        "page_verdict": response.headers.get("X-Page-Verdict"),
        "billed": response.headers.get("X-Billed"),
        "file": str(image_path),
    }
    image_path.with_suffix(".json").write_text(
        json.dumps(metadata, indent=2), encoding="utf-8"
    )
    return metadata


if __name__ == "__main__":
    for url in URLS:
        for attempt in range(3):
            try:
                result = capture(url)
                print(json.dumps(result))
                break
            except (requests.Timeout, requests.ConnectionError) as exc:
                if attempt == 2:
                    raise
                time.sleep(2 ** attempt)

Install the dependency with python -m pip install requests and set SCREENSHOTNEO_API_KEY in the process environment. Run the script from a scheduler at the interval that fits your archive requirements. The example retries connection and timeout errors only; it does not retry every HTTP error, since authentication, parameter, and plan errors generally need a configuration fix first.

cURL: save one capture

curl -fG "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o example.webp

Node.js: save one capture

import { writeFile } from "node:fs/promises";

const q = new URLSearchParams({
  access_key: process.env.SCREENSHOTNEO_API_KEY,
  url: "https://example.com",
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const contentType = res.headers.get("content-type") ?? "";
if (!contentType.startsWith("image/")) {
  throw new Error(`Expected image response, got ${contentType}`);
}
await writeFile("example.webp", Buffer.from(await res.arrayBuffer()));

These single-capture examples are useful for checking credentials and output handling. For a recurring archive, put URL scheduling, unique filenames, metadata, retry limits, alerts, and retention around the request.

4. Choose settings that make captures comparable

Capture settings determine what a later reader can see. Store the settings used on every run; otherwise, a difference in viewport or render mode can look like a change to the website.

  • Viewport or device: Use the same dimensions or device preset for comparisons. A mobile viewport and desktop viewport show different layouts.
  • Viewport or full page: Full-page images show content below the fold. For long or lazy-loaded pages, ensure the provider waits for content to render before capture.
  • Format: Use a supported format that your review and archival tools can open. Record it in metadata.
  • Rendering conditions: Where supported, standardize locale, timezone, authentication, cookies, and other page state that can alter the result.
  • Dynamic pages: Choose a wait condition appropriate to the site, such as waiting for a selector or a brief delay. Network idle may never occur on pages with persistent connections.
  • Cache behavior: If the provider offers caching, understand whether a request returns a new render or a previous result. Record cache or page status when provided.

For ScreenshotNeo, the capture API supports full-page capture with lazy images loaded; CSS-selector element capture; dark mode; 12 device presets and custom viewports; retina scale; PDF; custom CSS and JavaScript; click-before-capture; selector hiding; selector, delay, or network-idle waits; request and resource blocking; custom headers, cookies, user agent, authorization, timezone, geolocation, and transparent backgrounds. It also supports resizing and caching with a chosen TTL. See the docs for parameter details. Avoid changing capture conditions during a visual time series unless the change is intentional and documented.

5. Store files and metadata for later retrieval

Use a stable archive key that identifies the page and the capture time, while keeping the unmodified URL in metadata. For example:

archive/{url-id}/{capture-utc-timestamp}.webp
archive/{url-id}/{capture-utc-timestamp}.json

A useful metadata record can include:

  • Original URL and, if relevant, final URL after redirects.
  • Capture timestamp in UTC and the scheduler’s intended run time.
  • Viewport or device, full-page setting, output format, and other material capture settings.
  • Provider, request/job identifier, completion status, and useful response headers.
  • Object-storage key or local path, file size, and a SHA-256 checksum.
  • Retry count or error details when the run did not produce a usable file.

ScreenshotAPI.to recommends retaining a timestamp and SHA-256 checksum alongside an archived capture as an integrity aid. A checksum can help detect whether the file later changed; it does not establish when the page existed, who captured it, or whether a screenshot is legally admissible. ScreenshotAPI.to.

For object storage, configure access controls, lifecycle retention, and backup or replication behavior to fit your requirements. Keep API credentials out of filenames and metadata records. Protect the metadata index as carefully as the files, since it is what makes a large archive searchable.

6. Make the archive reliable

Website rendering depends on the target site, network, and capture provider, so a successful scheduled job should mean more than “the process ran.” Define a successful capture as a completed request that returned the expected output, passed basic validation, and was persisted with its metadata.

  • Use bounded retries: Retry transient network errors and selected server-side failures with backoff. Avoid endless retries and avoid retrying permanent errors unchanged.
  • Make writes idempotent: Use a stable run identifier or timestamped key so retries do not overwrite unrelated captures.
  • Validate outputs: Check HTTP status, content type, non-empty bytes, and any provider-specific completion or verdict fields.
  • Keep failure records: Persist the URL, scheduled time, error class, and attempt count even when no screenshot was produced.
  • Alert on patterns: Alert on repeated failures or a growing backlog, rather than treating each isolated timeout as proof the whole archive is unavailable.
  • Reconcile expected runs: Compare scheduled URLs with completed records so gaps are discoverable.
  • Test restore and retrieval: Confirm that older files can still be located, downloaded, and matched to their metadata.

If the provider offers asynchronous jobs or signed webhooks, follow its documented job lifecycle and verify webhook signatures before accepting completion events. Do not assume that starting a job means the archive file is ready.

7. Performance, reliability, and cost

Archive volume grows with the number of URLs, capture frequency, page length, and output size. A rough request estimate is URLs × captures per URL per day × days; this is a planning formula, not a benchmark. Full-page pages and high-resolution output can create larger files, take longer to render, and consume more storage and transfer.

Keep a queue for large URL lists so that a temporary provider or target-site slowdown does not stop all work. Set concurrency with the provider’s request limits and your own storage capacity in mind. If many URLs are due at once, spread requests across a schedule where your use case permits.

Compare total cost across capture requests, plan requirements, storage, transfer, workflow execution, and time spent operating the system. Verify API availability, minimum capture interval, request limits, retention window, output formats, authentication, and export behavior against current provider documentation. These details can change, and the vendor material cited here does not provide a neutral performance or reliability comparison.

8. Troubleshooting common failures

Symptom Likely cause Fix
401 or 403 response Missing, invalid, or insufficiently entitled API credentials. Check the key, authentication method, account plan, and whether API access is enabled. Keep secrets in environment variables or a secret manager.
400 response Malformed URL or unsupported parameter/value. URL-encode the target, use documented parameter names, and test one request before scheduling it.
429 response Request rate or plan limit reached. Reduce concurrency, add backoff, and verify the provider’s current limits and permitted schedule.
Timeout or failed load Slow target, network issue, or page that does not reach the selected wait condition. Use a suitable timeout and wait strategy; retry transient failures with a cap and record the miss.
Blank or incomplete image Capture began before content rendered, page requires interaction/authentication, or content is lazy-loaded. Supply required cookies or headers where permitted, wait for a meaningful selector or content load, and enable full-page/lazy-image behavior where available.
HTML or JSON saved as an image Error response body was written without checking status or content type. Check HTTP status and response Content-Type before writing bytes; preserve the error body separately for diagnosis if appropriate.
Duplicate or overwritten files Filename lacks a unique capture timestamp or job identifier. Use a UTC timestamp or unique run ID in the object key and make retries idempotent.
Archive has unexplained gaps Scheduler missed runs, failures were swallowed, or asynchronous jobs were not collected. Track expected runs, persist failure records, alert on gaps, and reconcile job completion with stored files.
Visual differences that are not site changes Viewport, locale, page state, cache, or timing changed between captures. Keep capture settings consistent and store them with each file.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. A GET request returns a PNG, JPEG, WebP, or PDF, and the API can fit the capture step in the archive pipeline. See the API docs for configuration options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include X-Page-Verdict and X-Billed headers.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month with no card.

10. Frequently asked questions

How do I automatically take screenshots of a website?

Put a screenshot API request behind a scheduler or create a monitor in a service that supports recurring captures. Save each result with its URL, UTC timestamp, settings, and status.

Can I schedule website screenshots with an API?

Yes. Some services expose monitor and scheduling APIs; with a capture-only API, use a cron job, queue, or workflow to run requests on your schedule. Confirm any plan-specific minimum interval first.

How do I save website screenshots over time?

Write each capture to durable storage under a timestamped key and store a metadata record that points to it. Set and review a retention policy for both the file and its metadata.

How do I archive full-page website screenshots?

Use a provider’s full-page option, keep viewport and output settings consistent, and ensure the capture waits for lazy-loaded content. Confirm whether the provider supports very long pages and what output formats apply.

Does a screenshot checksum prove when the page was captured?

No. A checksum can show whether the saved file’s bytes changed after the checksum was created. It does not independently prove capture time, origin, or legal admissibility.

Sources