ScreenshotNeo

BlogHow-to

How to Archive a Website for Long-Term Access

Preserve a website with clear scope, WARC or WACZ exports, metadata, multiple copies and replay tests.

By the ScreenshotNeo team1 October 20264 min read

Short answer: define the pages and interactions in scope, capture them while available, export WARC or WACZ, retain capture metadata and multiple copies, then replay and inspect the result. A screenshot or saved homepage is one view, not automatically a complete archive.

The Library of Congress Web Archiving Program calls websites ephemeral and at risk. Its guidance prefers WARC and accepts WACZ.

1. Define the scope

Record seed URLs, subdomains, page families, interactions, capture dates and exclusions. A seed may be one page, document, subdomain or domain. Decide whether “website” means selected pages, a public site, or a specific workflow.

  • List home, articles, products, downloads and legal pages.
  • Identify menus, tabs, infinite scroll, forms and media that require interaction.
  • Save the URL list and reason for selection.

2. Choose a capture method

Need Option Check
One public page Public web archive save feature Linked assets, export and current rules.
Interactive browsing ArchiveWeb.page Active capture, pending URLs, WARC/WACZ export and replay.
Many pages Crawler or institutional workflow Scope, depth, JavaScript, metadata, storage and replay.
Visual evidence ScreenshotNeo Use alongside an archive; an image is not HTTP history.

3. Capture an interactive site

  1. Install ArchiveWeb.page and start a session.
  2. Open each seed while capture is active.
  3. Perform the interactions to preserve.
  4. Wait for pending URLs and assets before navigating; see the capture guide.
  5. Export WARC or WACZ, then replay it locally.

Inspect representative pages, links, images, downloads and interactions. Record scope, timestamps, tool version and omissions.

4. Preserve WARC/WACZ and metadata

WARC is the preservation-oriented, non-proprietary choice; WACZ is an accepted Webrecorder package. Keep the export unchanged and store a sidecar record:

{"capture_started":"2026-10-01T12:00:00Z","capture_ended":"2026-10-01T12:30:00Z","seeds":["https://example.com/"],"scope":"public pages linked from the seed","format":"WARC or WACZ","known_omissions":[]}

Replace example values with your actual capture facts. Do not call it a complete clone unless your scope and inspection support that claim.

5. Make a site easier to preserve

  • Use accessible, standards-based HTML and ordinary links.
  • Provide a comprehensive sitemap and stable URIs.
  • Prefer open formats and expose important content without login where possible.

These practices reduce obstacles but cannot guarantee flawless capture.

6. Store and replay-test

  1. Keep archive files and metadata together.
  2. Maintain multiple managed copies in separate locations; an external drive is only one copy.
  3. Replay a representative sample and log failures.
  4. Recapture important content while the source remains available.

Replay shows captured responses, not every future third-party service. Distinguish archived content from the live site.

Or skip the browser setup

For dated visual records, ScreenshotNeo returns PNG, JPEG, WebP or PDF from one GET request. Pair it with WARC/WACZ when you need images or PDFs.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture it accepts consent banners and removes 60+ known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, with X-Page-Verdict and X-Billed headers explaining the result. Options include full-page lazy-image capture, CSS-element shots, dark mode, 12 device presets or custom viewports, retina scale, PDF controls, custom CSS/JavaScript, click and wait actions, blocking, custom headers/cookies/user agent/Authorization, timezone, geolocation, transparency, resizing, TTL caching, signed links, async jobs, bulk capture, usage API, OpenAPI and an MCP server with take_screenshot, get_page_info and capture_pdf.

Free: 1,000 shots/month with no card. Paid plans start at $5 for 3,000; yearly billing gives two months free. Create a free ScreenshotNeo account.

Limitations and troubleshooting

Problem Cause Fix
Missing links/images Pending assets or narrow scope Wait, revisit, expand scope and recapture.
Interactive content absent Action was never captured Perform it during an active session and verify replay.
Streaming/video missing Rich media and remote services may not be preserved Log the omission and preserve a permitted download separately.
Only homepage exists Homepage mistaken for a crawl Define seeds and page families, then crawl or browse them.
Replay looks wrong Uncaptured dependency or live third-party service Compare logs and recapture while available.
ScreenshotNeo not billed Bot check, blank page, timeout, failed load or cache hit Read verdict headers; adjust waits, blocking or access.

Performance, reliability and cost

  • Broad crawls take longer as scope and assets grow; start with representative pages.
  • Wait for pending resources, keep logs, replay-test and retain multiple copies.
  • ScreenshotNeo plans: Free 1,000; $5/3,000; $15/15,000; $39/60,000; $99/250,000; $249/1,000,000. Only clean shots are billed.

FAQ

Is a PDF a website archive?

No. It is a rendered document view; WARC/WACZ preserves web context.

How often should I recapture?

Base frequency on change rate and deadlines; retain each timestamp and scope.

Can I archive a private site?

Only with authorization; document login-dependent and missing content.

Does ScreenshotNeo replace WARC?

No. Use it for clean snapshots and PDFs alongside preservation-oriented capture.

Checklist: scope; seed list; interactions; WARC/WACZ; metadata; omissions; multiple copies; replay inspection.