ScreenshotNeo

BlogHow-to

How to Recover a Website From the Wayback Machine

Recover an old website from Wayback captures, rebuild missing pages and assets, fix links, and avoid the archive’s limits.

By the ScreenshotNeo team1 October 20267 min read

Short answer: The Wayback Machine can provide historical copies of pages and some assets, but it is not a general-purpose website backup or restoration service. Recover a site by finding the best captures for each URL, downloading usable HTML and assets, rebuilding the site on hosting you control, and checking every link and interactive feature.

Internet Archive says its terms do not cover backups for the general public and that it can no longer pack up lost sites as a recovery service. Treat captures as source material for a reconstruction, and keep your own repository and backups after the rebuild.

1. What the Wayback Machine can and cannot recover

Usually recoverable Often incomplete or unavailable
Archived HTML for captured URLs Pages that were never crawled
Images, CSS and downloads saved with a page Images blocked by robots.txt, exclusions or crawl gaps
Historical URL paths and navigation clues Server-side code, databases and private content
Static layout and copy Forms, login flows, JavaScript interactions and APIs

The archive’s Save Page Now feature saves the entered page, including images and CSS, but does not save outlinks or start a whole-site crawl. A dynamic page that depends on the original host will not retain its original functionality. See Internet Archive’s Wayback guidance, Save Pages documentation and General Information.

2. Before you start: collect the original site map

  1. Write down every known hostname: example.com, www.example.com, staging hosts and old subdomains.
  2. Collect URL paths from old emails, search results, analytics exports, source repositories, DNS history and printed links.
  3. List important asset paths such as /images/, /css/, /js/, PDFs and downloads.
  4. Confirm that you own the content or have permission to reuse it. An archived copy is not automatically proof that you may republish the material.
  5. Create a working directory and a manifest recording the original URL, capture timestamp, local filename, HTTP status and notes.

3. Find the best capture for each URL

Search the exact URL first

Enter the old address in the Wayback Machine. Inspect the timeline and calendar, then open captures from several dates. Test all meaningful variants:

  • http:// and https://
  • www and non-www
  • with and without a trailing slash
  • index files such as /index.html
  • alternate paths, language prefixes and old extensions

A newer capture is not automatically better. Prefer the date that contains the complete HTML, images, stylesheets and downloads you need.

Check assets and important child pages

Open the page source and record every referenced stylesheet, script, image, font and download. Search each asset URL separately in Wayback. Repeat this for navigation targets, contact pages, policies, product pages and documents. A home page can be present while a critical stylesheet or image is missing.

Use the capture timestamp in archived URLs

Wayback links commonly use a timestamped form such as https://web.archive.org/web/20200102123456id_/https://example.com/page. The id_ view requests the archived resource without the replay toolbar, which is useful when downloading files. Keep the original URL in your manifest even when the download URL contains a timestamp.

4. Save HTML and assets without losing paths

For a small site, save pages manually from the browser and download missing assets one by one. Preserve the original directory structure so relative links continue to work:

recovered-site/
  index.html
  about/index.html
  css/site.css
  js/site.js
  images/logo.png
  downloads/catalog.pdf

When a saved page contains replay URLs, edit them to point at your local copies. Check both root-relative links (/css/site.css) and document-relative links (../css/site.css). Keep a copy of the untouched archived HTML before making replacements.

Download a single archived resource with cURL

curl -L "https://web.archive.org/web/20200102123456id_/https://example.com/css/site.css" -o recovered-site/css/site.css

Use -L for redirects, but inspect the response before trusting the file. A missing capture may return an archive error page instead of the asset you expected.

Record response headers and failures with Python

import requests
from pathlib import Path

url = "https://web.archive.org/web/20200102123456id_/https://example.com/"
out = Path("recovered-site/index.html")
out.parent.mkdir(parents=True, exist_ok=True)

r = requests.get(url, timeout=60)
r.raise_for_status()
if "text/html" not in r.headers.get("content-type", ""):
    raise RuntimeError(f"Unexpected content type: {r.headers.get('content-type')}")
out.write_bytes(r.content)
print(r.status_code, r.headers.get("content-type"), out)

Save a page with Node.js

import { writeFile } from "node:fs/promises";

const archivedUrl = "https://web.archive.org/web/20200102123456id_/https://example.com/";
const res = await fetch(archivedUrl, { redirect: "follow" });
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
await writeFile("recovered-site-index.html", Buffer.from(await res.arrayBuffer()));

5. Rebuild the site on a new host

  1. Put the recovered files in a local web server or static-site project.
  2. Recreate the original navigation and directory paths before redesigning anything.
  3. Replace replay URLs with local or new production URLs.
  4. Restore fonts, image dimensions, responsive CSS and print styles where available.
  5. Replace forms with a current form endpoint and add spam protection.
  6. Replace scripts that called the old API, analytics account or third-party service.
  7. Deploy to hosting you control, then configure DNS and HTTPS.

Run a crawler against the local copy and production host. Check for 404s, redirect loops, mixed-content warnings, missing media, malformed downloads and links that still point to web.archive.org. Add redirects from valuable old paths to their rebuilt equivalents, and preserve slugs where possible.

Test mobile and accessibility behavior

  • Resize the viewport through common mobile and desktop widths.
  • Test keyboard navigation, focus states, heading order, image alternative text and color contrast.
  • Verify that menus work without the original JavaScript dependencies.
  • Check that PDFs and downloads have descriptive names and usable MIME types.

6. Handle missing pages and broken features

Symptom Likely cause Practical fix
Image shows an archive placeholder The image was never captured or is excluded Search other dates and URL variants; recover from your own backups or replace it with a rights-cleared asset.
CSS is unstyled Stylesheet URL is missing or still points to the old host Search the exact stylesheet path in Wayback, download it, and update the link.
Fonts fail Font files were not archived or cross-origin rules block them Use a licensed replacement and update @font-face.
Form submits nowhere Original server-side endpoint is gone Implement a new endpoint and test validation, mail delivery and abuse controls.
JavaScript errors appear Missing libraries, APIs or inline assumptions about the old host Inspect the console, pin available dependencies and rewrite calls to current services.
Only the home page exists Outlinks were not captured or the crawl had gaps Search each path directly, inspect sitemap clues and use your own records to rebuild missing pages.
Replay redirects unexpectedly The archive selected another capture or protocol Open the calendar, choose a specific timestamp and preserve the intended original URL.

7. Verify the recovered site

  • Content: compare headlines, pricing, legal text and downloads with your source records.
  • Assets: confirm every image, stylesheet, script, font and PDF loads from your host.
  • URLs: test canonical URLs, redirects, trailing slashes and old inbound links.
  • Behavior: test forms, search, menus, embeds, login boundaries and error pages.
  • Operations: enable version control, automated backups, HTTPS renewal and monitoring.
  • Rights: remove material you do not own or cannot republish.

8. Performance, reliability and cost considerations

Wayback recovery is manual and capture-dependent. Downloading many assets can take time, and repeated replay requests may produce different captures. Work from a manifest, cache files locally, and make the rebuild reproducible in version control. Keep the archive URLs and timestamps as provenance notes, but serve visitors from your own host.

Do not assume that an archived snapshot is a complete backup. For recurring preservation or whole-site collection, investigate a dedicated web-archiving service. Save Page Now is a one-page capture tool, not a crawler for an entire site.

9. Or skip the browser setup

If you need a current screenshot of a recovered page while checking layouts or documenting evidence, ScreenshotNeo provides a website screenshot API. It accepts one GET request and returns PNG, JPEG, WebP or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups and chat widgets are removed; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get started.

10. FAQ

Can Internet Archive restore my whole domain?

No. It provides historical captures, not a general-public backup and restoration service.

Why is a page visible but its images missing?

The images may have different URLs, a different capture date, robots exclusions or a crawl gap. Search each asset URL independently.

Can I make archived forms work?

Usually you must build a new form endpoint because archived JavaScript and server-side integrations cannot connect to the original host.

Is a Wayback screenshot authoritative evidence?

Not automatically. For legal or evidentiary use, consult Internet Archive’s legal and affidavit procedure and preserve capture details.

How do I prevent another loss?

Keep source code in version control, maintain host backups, export databases and periodically verify that restoration procedures work.