How to Download an Entire Website From the Wayback Machine
Use the Wayback CDX index to find archived URLs, select captures, download them repeatably, and understand what a local copy cannot restore.
Short answer: Use the Wayback Machine’s CDX index to enumerate captures for a domain, choose the timestamps and URL scope you need, then download each capture through its replay URL with a script or bulk downloader. Keep a manifest containing the original URL, capture timestamp, status, MIME type, and local filename. This creates a local collection of material the Archive captured and can still serve; it does not guarantee a complete, functioning clone of the original site.
Save Page Now submits an individual page. It does not crawl that page’s outlinks into a whole-site export. For site-scale retrieval, start with the Wayback CDX index, then retrieve the records it returns.
Define what “entire website” means
Choose the scope before downloading. These are different projects:
| Goal | What to select | Result |
|---|---|---|
| One historical version | One date or narrow date range, usually one capture per URL | A smaller, more coherent snapshot |
| All indexed pages | Every distinct URL in a domain or path | A broad collection across dates |
| Historical corpus | Multiple timestamps for each URL | More evidence of how the site changed, with more files and storage |
| One section | A host or path such as example.com/docs/* |
Less noise and a manageable download |
“Everything” can mean every URL, every capture of every URL, or one selected version per URL. State that choice in your project notes and manifest.
How the Wayback workflow works
- Enumerate. Query CDX for captures matching a domain, host, or path.
- Filter. Keep the MIME types, status codes, dates, and URL patterns relevant to your goal.
- Select. Choose one timestamp per URL for a point-in-time copy, or retain multiple timestamps for a historical set.
- Retrieve. Request each capture through a raw replay URL using the
id_modifier. - Record. Write a manifest mapping every local file to its original URL and archive timestamp.
- Inspect. Open representative HTML, images, stylesheets, scripts, and documents. Retry transient failures and record permanent misses.
The CDX record describes an indexed capture. It is an index entry, not proof that every dependency exists or that server-side behavior can be reconstructed.
Discover captures with CDX
A basic query returns capture records for one URL:
curl -G 'https://web.archive.org/cdx/search/cdx' \
--data-urlencode 'url=example.com/' \
--data-urlencode 'output=json' \
--data-urlencode 'fl=timestamp,original,mimetype,statuscode,digest,length' \
--data-urlencode 'filter=statuscode:200' \
--data-urlencode 'collapse=digest'
For a domain-wide query, use a wildcard:
curl -G 'https://web.archive.org/cdx/search/cdx' \
--data-urlencode 'url=example.com/*' \
--data-urlencode 'output=json' \
--data-urlencode 'fl=timestamp,original,mimetype,statuscode,digest,length' \
--data-urlencode 'filter=statuscode:200' \
--data-urlencode 'collapse=urlkey'
Useful fields include:
| Field | Use |
|---|---|
timestamp |
Capture time in YYYYMMDDhhmmss form |
original |
URL requested when the capture was made |
mimetype |
Helps distinguish HTML, images, CSS, JavaScript, PDFs, and other resources |
statuscode |
HTTP status recorded for the capture |
digest |
Content fingerprint useful for collapsing duplicate content |
length |
Recorded capture size |
If the original URL contains its own query string, pass it as a separately URL-encoded parameter. For large result sets, use CDX pagination rather than requesting an unbounded response.
Limit dates and result size
curl -G 'https://web.archive.org/cdx/search/cdx' \
--data-urlencode 'url=example.com/*' \
--data-urlencode 'from=20190101' \
--data-urlencode 'to=20191231' \
--data-urlencode 'output=json' \
--data-urlencode 'fl=timestamp,original,mimetype,statuscode,digest,length' \
--data-urlencode 'filter=statuscode:200' \
--data-urlencode 'collapse=urlkey' \
--data-urlencode 'page=0' \
--data-urlencode 'pageSize=1000'
Pagination parameter names and limits can change. Check the current CDX documentation before running a very large export and continue until a page contains no records.
Download captures with Python
This script paginates CDX results, keeps one record per original URL, downloads raw replay responses, and writes a CSV manifest. It is intentionally conservative: it sleeps between requests, creates parent directories, and continues after individual failures.
#!/usr/bin/env python3
import csv
import hashlib
import json
import os
import re
import sys
import time
from pathlib import Path
from urllib.parse import quote, urlparse
import requests
DOMAIN = sys.argv[1] if len(sys.argv) > 1 else "example.com"
OUT = Path(sys.argv[2] if len(sys.argv) > 2 else "wayback-download")
FROM = os.getenv("WAYBACK_FROM", "")
TO = os.getenv("WAYBACK_TO", "")
PAGE_SIZE = 500
DELAY_SECONDS = 0.5
OUT.mkdir(parents=True, exist_ok=True)
session = requests.Session()
session.headers.update({"User-Agent": "wayback-research-downloader/1.0"})
fields = ["timestamp", "original", "mimetype", "statuscode", "digest", "length"]
manifest_path = OUT / "manifest.csv"
with manifest_path.open("w", newline="", encoding="utf-8") as manifest_file:
writer = csv.DictWriter(manifest_file, fieldnames=fields + ["local_path", "download_status"])
writer.writeheader()
page = 0
while True:
params = {
"url": f"{DOMAIN}/*",
"output": "json",
"fl": ",".join(fields),
"filter": "statuscode:200",
"collapse": "urlkey",
"page": page,
"pageSize": PAGE_SIZE,
}
if FROM:
params["from"] = FROM
if TO:
params["to"] = TO
response = session.get("https://web.archive.org/cdx/search/cdx", params=params, timeout=60)
response.raise_for_status()
payload = response.json()
if not payload:
break
# CDX JSON normally starts with a header row.
header = payload[0]
rows = [dict(zip(header, row)) for row in payload[1:]] if header == fields else [dict(zip(fields, row)) for row in payload]
if not rows:
break
for record in rows:
timestamp = record["timestamp"]
original = record["original"]
parsed = urlparse(original)
safe_host = parsed.netloc or DOMAIN
safe_path = re.sub(r"[^A-Za-z0-9._-]+", "_", parsed.path.strip("/")) or "index"
suffix = hashlib.sha1(original.encode("utf-8")).hexdigest()[:12]
local_path = OUT / safe_host / f"{timestamp}_{safe_path}_{suffix}"
replay_url = f"https://web.archive.org/web/{timestamp}id_/{original}"
status = "ok"
try:
if not local_path.exists():
local_path.parent.mkdir(parents=True, exist_ok=True)
download = session.get(replay_url, timeout=90)
download.raise_for_status()
local_path.write_bytes(download.content)
else:
status = "already-present"
except requests.RequestException as exc:
status = f"error: {exc}"
writer.writerow({**record, "local_path": str(local_path), "download_status": status})
manifest_file.flush()
time.sleep(DELAY_SECONDS)
page += 1
print(f"Finished. Manifest: {manifest_path}")
Run it like this:
python3 -m pip install requests
WAYBACK_FROM=20180101 WAYBACK_TO=20181231 python3 download_wayback.py example.com archive-output
The script chooses one record per URL because of collapse=urlkey. Remove that parameter when you deliberately need multiple captures, then change the local naming logic so each timestamp remains distinct.
Download with Node.js
import fs from "node:fs/promises";
import path from "node:path";
const domain = process.argv[2] || "example.com";
const params = new URLSearchParams({
url: `${domain}/*`,
output: "json",
fl: "timestamp,original,mimetype,statuscode,digest,length",
filter: "statuscode:200",
collapse: "urlkey",
page: "0",
pageSize: "500"
});
const cdx = await fetch(`https://web.archive.org/cdx/search/cdx?${params}`);
if (!cdx.ok) throw new Error(`CDX request failed: ${cdx.status}`);
const payload = await cdx.json();
const header = payload[0];
const rows = payload.slice(1).map(row => Object.fromEntries(header.map((key, i) => [key, row[i]])));
await fs.mkdir("wayback-download", { recursive: true });
const manifest = [];
for (const record of rows) {
const replay = `https://web.archive.org/web/${record.timestamp}id_/${record.original}`;
const file = path.join("wayback-download", `${record.timestamp}-${encodeURIComponent(record.original)}`);
const response = await fetch(replay);
if (!response.ok) {
manifest.push({ ...record, local_path: file, download_status: `error ${response.status}` });
continue;
}
await fs.writeFile(file, Buffer.from(await response.arrayBuffer()));
manifest.push({ ...record, local_path: file, download_status: "ok" });
}
await fs.writeFile("wayback-download/manifest.json", JSON.stringify(manifest, null, 2));
For a real export, add pagination, retries with backoff, a delay between requests, and safe filesystem names. The Python example includes those operational safeguards.
Understand replay URLs
A raw capture URL has this shape:
https://web.archive.org/web/TIMESTAMPid_/ORIGINAL_URL
TIMESTAMP is the CDX timestamp. ORIGINAL_URL is the escaped or unescaped URL returned by CDX. The id_ modifier requests the captured response without the Wayback navigation toolbar, which is preferable when saving bytes to disk. Preserve the original URL separately in your manifest because replay URLs are implementation details of retrieval.
Choose files and avoid accidental duplicates
- Use
mimetype:text/htmlfilters when you only need pages. - Keep CSS, JavaScript, fonts, and images when your goal is visual inspection.
- Use
collapse=digestto reduce identical content across URLs, orcollapse=urlkeyto keep one capture per URL. - Do not collapse when you are studying changes over time.
- Keep query-string URLs distinct when the query changes page content.
- Store the CDX response or a normalized copy beside the downloaded files for auditability.
Storage, performance, and reliability
Estimate storage from the index
Sum the CDX length values for the records you selected, then add room for duplicate files, manifests, retries, and filesystem overhead. The Archive’s guidance is to save downloads to a location you choose; it does not publish one universal drive size for every site. An external drive is an optional way to add capacity.
Control request rate
Large exports involve many requests. Use pagination, a modest delay, connection timeouts, and retries for transient network errors. Keep downloads resumable by skipping files that already exist and flushing the manifest after each record.
Make failures visible
Record HTTP failures, empty responses, and missing captures instead of silently dropping them. Compare the manifest’s selected-record count with the number of successful files. Inspect a sample from each MIME type and date range.
Expect incomplete reconstruction
A capture can be missing, partial, or missing dependencies. Dynamic features, databases, login flows, and server-side code are not guaranteed to return as a working service. Treat the result as an archived collection bounded by what was captured and retrievable.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| CDX returns no rows | The URL pattern has no indexed captures, or the date/filter is too narrow | Try the bare domain, remove filters, and widen the date range. |
| Only one page downloads | You used Save Page Now or queried one exact URL | Query a wildcard such as example.com/* and paginate the CDX results. |
| JSON parsing fails | The response is an error page, rate-limit response, or non-JSON output | Print the status and content type, slow down, and check the request URL and parameters. |
| Files contain a toolbar or wrapper | The replay URL omitted the raw-capture modifier | Use /web/TIMESTAMPid_/ORIGINAL_URL. |
| Images or CSS are missing | Those dependencies were never captured, were filtered out, or have different timestamps | Query dependency MIME types, retrieve their own CDX records, and accept that some resources may not exist. |
| Too many duplicate files | Multiple captures contain identical bytes | Use collapse=digest or deduplicate by digest after downloading. |
| Some captures fail intermittently | Transient network or archive response errors | Retry with exponential backoff, preserve failed rows in the manifest, and rerun only missing files. |
| Local links do not work | Pages reference absolute URLs or server routes that are not present locally | Keep the original URL mapping, rewrite links in a separate post-processing step, and do not assume application behavior can be restored. |
Or skip the browser setup
If your goal is a clean screenshot of a current page rather than a historical download, ScreenshotNeo provides a website screenshot API and MCP server. One request returns a PNG, JPEG, WebP, or PDF; it does not retrieve Wayback captures or replace CDX for archival downloading.
See the ScreenshotNeo API documentation for options. A minimal request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing result.
- An MCP server lets Claude, Cursor, and other MCP clients call
take_screenshot,get_page_info, andcapture_pdf. - The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account to try it.
FAQ
Can Save Page Now download a whole site?
No. It saves the submitted page and included resources; it does not start a crawl of outlinks.
Should I download one capture or every timestamp?
Use one capture per URL for a coherent snapshot. Keep multiple timestamps when change history matters.
Does a downloaded archive become a working website?
Not necessarily. It is a local collection of retrievable captures. Dynamic services, databases, authentication, and missing dependencies may not function.
How do I resume a large download?
Write a manifest as you go, skip existing files, and rerun the script for failed or missing rows.
Where should I store the files?
Use any location with enough capacity for the selected records and overhead. Estimate from CDX lengths rather than assuming a fixed site size.


