How to Archive a Website for Long-Term Access
Preserve a website with clear scope, WARC or WACZ exports, metadata, multiple copies and replay tests.
Short answer: define the pages and interactions in scope, capture them while available, export WARC or WACZ, retain capture metadata and multiple copies, then replay and inspect the result. A screenshot or saved homepage is one view, not automatically a complete archive.
The Library of Congress Web Archiving Program calls websites ephemeral and at risk. Its guidance prefers WARC and accepts WACZ.
1. Define the scope
Record seed URLs, subdomains, page families, interactions, capture dates and exclusions. A seed may be one page, document, subdomain or domain. Decide whether “website” means selected pages, a public site, or a specific workflow.
- List home, articles, products, downloads and legal pages.
- Identify menus, tabs, infinite scroll, forms and media that require interaction.
- Save the URL list and reason for selection.
2. Choose a capture method
| Need | Option | Check |
|---|---|---|
| One public page | Public web archive save feature | Linked assets, export and current rules. |
| Interactive browsing | ArchiveWeb.page | Active capture, pending URLs, WARC/WACZ export and replay. |
| Many pages | Crawler or institutional workflow | Scope, depth, JavaScript, metadata, storage and replay. |
| Visual evidence | ScreenshotNeo | Use alongside an archive; an image is not HTTP history. |
3. Capture an interactive site
- Install ArchiveWeb.page and start a session.
- Open each seed while capture is active.
- Perform the interactions to preserve.
- Wait for pending URLs and assets before navigating; see the capture guide.
- Export WARC or WACZ, then replay it locally.
Inspect representative pages, links, images, downloads and interactions. Record scope, timestamps, tool version and omissions.
4. Preserve WARC/WACZ and metadata
WARC is the preservation-oriented, non-proprietary choice; WACZ is an accepted Webrecorder package. Keep the export unchanged and store a sidecar record:
{"capture_started":"2026-10-01T12:00:00Z","capture_ended":"2026-10-01T12:30:00Z","seeds":["https://example.com/"],"scope":"public pages linked from the seed","format":"WARC or WACZ","known_omissions":[]}
Replace example values with your actual capture facts. Do not call it a complete clone unless your scope and inspection support that claim.
5. Make a site easier to preserve
- Use accessible, standards-based HTML and ordinary links.
- Provide a comprehensive sitemap and stable URIs.
- Prefer open formats and expose important content without login where possible.
These practices reduce obstacles but cannot guarantee flawless capture.
6. Store and replay-test
- Keep archive files and metadata together.
- Maintain multiple managed copies in separate locations; an external drive is only one copy.
- Replay a representative sample and log failures.
- Recapture important content while the source remains available.
Replay shows captured responses, not every future third-party service. Distinguish archived content from the live site.
Or skip the browser setup
For dated visual records, ScreenshotNeo returns PNG, JPEG, WebP or PDF from one GET request. Pair it with WARC/WACZ when you need images or PDFs.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture it accepts consent banners and removes 60+ known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, with X-Page-Verdict and X-Billed headers explaining the result. Options include full-page lazy-image capture, CSS-element shots, dark mode, 12 device presets or custom viewports, retina scale, PDF controls, custom CSS/JavaScript, click and wait actions, blocking, custom headers/cookies/user agent/Authorization, timezone, geolocation, transparency, resizing, TTL caching, signed links, async jobs, bulk capture, usage API, OpenAPI and an MCP server with take_screenshot, get_page_info and capture_pdf.
Free: 1,000 shots/month with no card. Paid plans start at $5 for 3,000; yearly billing gives two months free. Create a free ScreenshotNeo account.
Limitations and troubleshooting
| Problem | Cause | Fix |
|---|---|---|
| Missing links/images | Pending assets or narrow scope | Wait, revisit, expand scope and recapture. |
| Interactive content absent | Action was never captured | Perform it during an active session and verify replay. |
| Streaming/video missing | Rich media and remote services may not be preserved | Log the omission and preserve a permitted download separately. |
| Only homepage exists | Homepage mistaken for a crawl | Define seeds and page families, then crawl or browse them. |
| Replay looks wrong | Uncaptured dependency or live third-party service | Compare logs and recapture while available. |
| ScreenshotNeo not billed | Bot check, blank page, timeout, failed load or cache hit | Read verdict headers; adjust waits, blocking or access. |
Performance, reliability and cost
- Broad crawls take longer as scope and assets grow; start with representative pages.
- Wait for pending resources, keep logs, replay-test and retain multiple copies.
- ScreenshotNeo plans: Free 1,000; $5/3,000; $15/15,000; $39/60,000; $99/250,000; $249/1,000,000. Only clean shots are billed.
FAQ
Is a PDF a website archive?
No. It is a rendered document view; WARC/WACZ preserves web context.
How often should I recapture?
Base frequency on change rate and deadlines; retain each timestamp and scope.
Can I archive a private site?
Only with authorization; document login-dependent and missing content.
Does ScreenshotNeo replace WARC?
No. Use it for clean snapshots and PDFs alongside preservation-oriented capture.
Checklist: scope; seed list; interactions; WARC/WACZ; metadata; omissions; multiple copies; replay inspection.


