ScreenshotNeo

BlogUse cases

Web Archiving Use Cases

Learn why organizations archive websites, what WARC preserves, where crawlers fail, and how to design a defensible capture workflow.

By the ScreenshotNeo team1 October 202610 min read

Web archiving creates a time-specific record of online content so people can access, verify, research, quote and reuse it after the live site changes. A sound program captures the page, its context and linked assets; records provenance and limitations; stores the result in durable formats such as WARC; and applies retention, access and integrity controls appropriate to the risk.

Why archive a website?

A live URL is a moving target. Pages are edited, deleted, redesigned, moved behind logins or made dependent on services that later disappear. An archive gives you a reference point: what was published, where it appeared, when it was captured and how its components related to one another.

  • Evidence and accountability: preserve notices, policies, decisions, emergency information and public commitments.
  • Research and scholarship: compare versions, study public discourse, trace links and cite a stable representation.
  • Legal and regulatory work: retain guidance, disclosures, terms and transactions that may later be challenged or changed.
  • Institutional memory: preserve event sites, online exhibitions, born-digital publications and community material.
  • Business continuity: retain a restorable snapshot after equipment failure, migration or catastrophe.
  • Change tracking: detect when important pages, assets or site relationships change.

The National Archives and Records Administration (NARA) explains that web content can meet the definition of a federal record when it documents agency organization, functions, policies, decisions, procedures, essential transactions or legal and financial rights. Its guidance also emphasizes that trustworthy records support legal and internal business needs. See NARA guidance on managing web records. Your jurisdiction, retention schedule and legal counsel determine what must be kept and whether a capture is sufficient for a particular proceeding.

Who uses web archiving?

Government and public-sector organizations

Agencies archive public notices, grant and procurement information, policy pages, election material, emergency updates and other communications. Capture plans should identify the responsible office, retention period, access restrictions and the risk of a later challenge to authenticity or completeness.

An archived page can preserve wording, publication context, linked documents and a capture time. It is evidence to review, not an automatic guarantee of admissibility. Document the capture process, controls, custody and known gaps, and consult counsel about the applicable rules.

Researchers, journalists and historians

Archives let researchers compare versions, follow campaigns and cite content after the live page changes. The UK National Archives describes long-term access and reuse of online knowledge as a central purpose of web archiving.

Libraries, museums and universities

Heritage organizations preserve institutional history, online exhibitions, event sites, digital publications and community resources. ISO/TR 14873:2013 addresses statistics and quality issues for web archiving across libraries, archives, museums, research centers and heritage foundations.

Businesses and product teams

Organizations preserve release announcements, documentation, customer-facing terms, campaign pages and a site map before a redesign. Event-based captures are useful for launches, policy changes, incidents, elections and litigation holds; periodic crawls suit stable content.

What should a web archive contain?

Define the record before choosing a crawler. A page-only screenshot may be useful evidence, but it does not preserve the complete web object. For each capture, decide whether you need:

  • HTML pages and their original URLs.
  • Images, style sheets, scripts, fonts, PDFs and downloadable files.
  • Audio, video, captions and player metadata.
  • HTTP headers, redirects, status codes, timestamps and content types.
  • Forms, query parameters, embedded frames and link relationships.
  • Authentication context, privacy restrictions and excluded paths.
  • A crawl manifest or site map showing scope, depth, exclusions and failures.
  • Provenance: who captured it, with which software and configuration, and what transformations occurred.

NARA transfer guidance calls for permanent web records to maintain original links, functionality and data integrity. The Library of Congress recommends recording the archiving institution, capture dates and times, and differences between archived functionality and the live site. See the Library of Congress Web Archives program for its collection and access approach.

WARC, ARC and WACZ formats

WARC (Web ARChive) is the main preservation format in the official guidance reviewed. It stores records such as responses and metadata in a structured container and is designed for replay and exchange. NARA identifies WARC 1.0 in permanent-record transfer guidance. The Library of Congress describes WARC with record-at-a-time GZIP compression as a preferred format.

ARC is an older archival container still found in established collections, including ARC_IA variants. WACZ is a packaged distribution format used by some modern workflows. The format does not by itself prove that a capture is complete or authentic; provenance, fixity, scope and operating controls still matter.

Format or artifact Best use Watch for
WARC Preservation, exchange and replay Requires a crawler and storage workflow that records metadata and failures
ARC/ARC_IA Compatibility with older collections May need conversion or legacy replay tooling
WACZ Packaged delivery and sharing in supported workflows Confirm tool and replay compatibility
HTML/assets folder Small internal snapshot or review copy Not a preservation container; links, scripts and provenance may be incomplete
Screenshot or PDF Visual reference and quick review Does not preserve the full resource graph or interactive behavior

A defensible web-archiving workflow

  1. Assess risk and ownership. Identify the activity, audience, legal or operational risk, required retention and record owner. Use reliability, authenticity, integrity and usability as design criteria.
  2. Write a scope. List domains, paths, URL patterns, linked assets, depth, languages, authenticated areas, exclusions and privacy constraints.
  3. Choose capture frequency. Use event captures for elections, emergencies, releases, policy changes and litigation holds. Use periodic crawls for stable sites. Increase frequency when change or risk increases.
  4. Choose the capture method. Select a crawler or managed service that handles JavaScript, redirects, assets, authentication and WARC output when those requirements matter.
  5. Record provenance. Keep the source URL, institution, start and end times, software version, configuration, scope, exclusions, HTTP outcomes and transformation notes.
  6. Preserve structure. Keep a crawl manifest or site map that shows relationships among pages and components. Record broken links and blocked resources instead of silently omitting them.
  7. Store durably. Use replicated storage, documented fixity or integrity checks, controlled access and a tested restore or replay procedure.
  8. Review quality. Sample pages, assets, redirects, dates and replay behavior. Compare high-risk pages with the live source and document differences.
  9. Publish or restrict deliberately. Apply privacy, copyright, authentication and embargo rules. Make the archiving institution and capture dates visible to users.

DIY snapshot for a small site

The following example creates a reviewable HTML snapshot with Playwright. It is useful for a small, permissioned site or for validating scope before adopting a WARC workflow. It is not a substitute for a standards-based crawl when you need complete replay, durable provenance or legal records.

Python and Playwright

import asyncio
import json
from datetime import datetime, timezone
from pathlib import Path
from playwright.async_api import async_playwright

URL = "https://example.com"
OUT = Path("snapshot")

async def main():
    OUT.mkdir(exist_ok=True)
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page()
        started = datetime.now(timezone.utc).isoformat()
        response = await page.goto(URL, wait_until="networkidle", timeout=90_000)
        html = await page.content()
        (OUT / "index.html").write_text(html, encoding="utf-8")
        manifest = {
            "url": URL,
            "captured_at": started,
            "status": response.status if response else None,
            "title": await page.title(),
            "note": "HTML review snapshot; assets and interactive behavior may be incomplete"
        }
        (OUT / "manifest.json").write_text(json.dumps(manifest, indent=2), encoding="utf-8")
        await browser.close()

asyncio.run(main())
python -m pip install playwright
playwright install chromium
python snapshot.py

Node.js and Playwright

import { chromium } from "playwright";
import { mkdir, writeFile } from "node:fs/promises";

const url = "https://example.com";
await mkdir("snapshot", { recursive: true });
const browser = await chromium.launch();
const page = await browser.newPage();
const response = await page.goto(url, { waitUntil: "networkidle", timeout: 90000 });
await writeFile("snapshot/index.html", await page.content());
await writeFile("snapshot/manifest.json", JSON.stringify({
  url,
  captured_at: new Date().toISOString(),
  status: response?.status() ?? null,
  title: await page.title(),
  note: "HTML review snapshot; assets and interactive behavior may be incomplete"
}, null, 2));
await browser.close();
npm install playwright
npx playwright install chromium
node snapshot.mjs

Quick static mirror with cURL

curl --location --fail --remote-name-all \
  --create-dirs --output-dir snapshot \
  https://example.com/

For a real archival collection, use a crawler that writes WARC and preserves response metadata, linked resources and crawl logs. Keep the manifest beside the collection and record every exclusion or failure.

What crawlers cannot capture reliably

  • Streaming media: manifests, DRM, adaptive segments and rights restrictions can prevent complete capture.
  • Databases and deep-web content: records generated only after a query, form submission or private session may not be discoverable.
  • Highly interactive applications: client-side state, WebSockets, maps, editors and personalization may replay differently.
  • Third-party services: analytics, ads, chat, embeds and APIs can disappear or change independently.
  • Authentication and consent: a crawler may lack credentials or may be blocked by bot defenses.
  • Large downloads: timeouts, rate limits and storage limits can create partial captures.

Record the gap in the manifest. Supplement the crawl with source-system exports, screenshots, PDFs, database extracts or other evidence when the risk assessment requires it.

How to compare archiving approaches

Criterion Questions to ask
Completeness Are JavaScript, linked assets, redirects, media and authenticated areas handled?
Replay fidelity Can users see the page as it appeared at capture time, and are differences disclosed?
Standards Does the workflow produce WARC or a compatible package?
Metadata Are timestamps, scope, software, HTTP outcomes and provenance retained?
Governance Can you enforce retention, access control, legal holds and deletion rules?
Durability Are there replicated copies, fixity checks and restore tests?
Operations Who schedules crawls, reviews failures, handles credentials and maintains replay?
Cost What are storage, bandwidth, compute, staffing and compliance costs over the retention period?

Or skip the browser setup

For a clean visual record of a page, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF. Cookie and consent banners, newsletter popups and chat widgets are removed before capture; bot checks, blank pages, failed loads, timeouts and cache hits are not billed, with the result identified by response headers.

See the ScreenshotNeo API documentation for the complete option set, including full-page capture with lazy images, CSS-element capture, device presets, dark mode, retina scale, PDF paper and margin controls, custom CSS and JavaScript, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture and usage reporting.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Use the screenshot as a visual supplement to a WARC collection when you need durable web preservation. ScreenshotNeo’s MCP server also lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf.

There are 1,000 free screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Troubleshooting

Symptom Likely cause Fix
Page is blank JavaScript did not finish, a bot check appeared or the origin failed Capture after network idle, log the final URL and status, retry with a permitted browser profile, and record the failure.
Images or styles are missing Relative URLs, blocked third-party resources or an incomplete asset crawl Preserve response URLs and headers, include linked assets in scope, and document blocked hosts.
Interactive content does not replay State depends on APIs, WebSockets, personalization or a database Capture the relevant state with a controlled session and supplement it with exports or screenshots.
Capture stops partway through Timeout, rate limit, storage limit or an unexpectedly large resource Set bounded retries, throttle requests, raise limits where appropriate and keep a failure log.
Archive cannot be trusted later Missing provenance, fixity, custody or scope documentation Store manifests, hashes, software/configuration details, access logs and retention decisions with the collection.
ScreenshotNeo response is not a clean image Target returned a bot check, blank page or failed load Inspect X-Page-Verdict and X-Billed; these outcomes are not billed, so fix the source or capture settings before retrying.

Performance, reliability and cost

  • Performance: scope narrowly, avoid recrawling unchanged content, use event-based captures for high-risk changes and parallelize only within the origin’s rate limits.
  • Reliability: retry transient failures, keep immutable manifests, replicate storage and test replay or restoration periodically.
  • Storage: linked assets, media and repeated versions dominate long-term cost. Deduplicate only when provenance and replay remain intact.
  • Governance: retention, privacy, copyright, credentials and legal holds can cost more than raw storage. Assign an owner for review and deletion decisions.
  • ScreenshotNeo billing: only clean shots are billed; bot checks, blank pages, timeouts, failed loads and cache hits cost nothing. Plans include Free (1,000/month), Starter ($5/3,000), Growth ($15/15,000), Pro ($39/60,000), Scale ($99/250,000) and Business ($249/1,000,000). Yearly billing gives two months free.

FAQ

Is a screenshot the same as a web archive?

No. A screenshot records appearance at one moment. A web archive can preserve responses, assets, metadata, links and replay context in a format such as WARC.

How often should a site be captured?

Set the interval from change rate and risk. Capture around known events and use periodic crawls for stable pages; there is no universal interval.

No. Legal sufficiency depends on jurisdiction, rules, custody, provenance, integrity and the facts of the matter. Consult counsel.

Can a crawler preserve a private database?

Only when the workflow is authorized and can access the records. Query-generated, authenticated and dynamic data often require source-system exports in addition to web capture.

When is ScreenshotNeo useful?

Use it when you need a clean visual capture or PDF quickly, especially when consent banners, popups, chat widgets, failed loads or AI-agent access would otherwise add browser work.