How to Build Your Own Wayback Machine for Web Page Archiving
Build a self-hosted web archive with ArchiveBox, choose useful capture formats, and learn how to replay, back up, and troubleshoot saved pages.
A practical way to build a personal web archive is to run ArchiveBox on a machine you control, add URLs, and keep its capture files on persistent storage. ArchiveBox can save several representations of a page, including HTML, SingleFile, screenshots, PDF, WARC, and extracted content. The right mix depends on whether you need a quick visual reference, searchable files, or data suitable for replay and transfer.
This creates a useful personal collection, not a replacement for the Internet Archive’s Wayback Machine. A capture may miss resources, logged-in content, scripts, or behavior that only appears after interaction. Start small, inspect what was saved, and verify replay before relying on it.
1. Decide what your archive needs to preserve
Write down the URLs and how often they should be saved. Occasional bookmarks, a scheduled collection, and a whole-site crawl have different storage and maintenance costs. ArchiveBox supports adding individual URLs and importing sources such as bookmarks, browser history, and feeds; consult its current repository and installation guidance for the supported workflows and commands.
| Need | Useful capture choice | Limitation to check |
|---|---|---|
| See how a page looked | PNG screenshot or PDF | These preserve an appearance, not the original interactive behavior. |
| Keep a page as one portable file | SingleFile | Embedded resources may be incomplete or behave differently from the live page. |
| Keep source files for inspection | HTML and related assets | External resources can change or disappear; local JavaScript behavior can vary by capture method. |
| Preserve crawl transactions for later replay | WARC, or a WACZ package when the workflow supplies one | Replay depends on the captured material and a compatible viewer. |
Decide whether you need static evidence or interactive replay. In ArchiveBox’s documented viewing behavior, only its wget and DOM capture methods execute archived JavaScript when viewed; other methods produce static output. A screenshot can be the clearest record of appearance, while a WARC-oriented workflow is more suitable when you need crawl data and replay context.
2. Install ArchiveBox and keep its data persistent
ArchiveBox currently recommends Docker Compose on its homepage. Use the project’s current quickstart for the exact commands and version, since installation steps can change. The general setup is:
- Choose a host where the archive can keep running and where you control the stored files.
- Follow the current Docker Compose quickstart to start the local service.
- Confirm the application’s data directory is on persistent storage, not only inside a disposable container.
- Keep the archive private unless you have deliberately configured access and reviewed current security guidance.
ArchiveBox also describes Docker and Python-based installation routes, a web interface, CLI, REST API, webhooks, browser extension, and filesystem access. Treat the official installation material as authoritative for what is available in your version.
3. Add a small set of URLs first
Start with a few pages that represent your real collection: a static article, a page with images, and a page that depends on scripts or interaction. Add a URL using the method supported by your installed ArchiveBox version. The repository documents URL input and scheduled imports; check the current instructions for exact syntax.
For each test URL, record the original URL and capture date. Add a note about why it matters or what you expect the archive to preserve. That small amount of context makes it easier to find the right snapshot later and spot a failed or incomplete capture.
Inspect the result before expanding the collection. Check whether the main text, images, styling, and links are present. If the page requires a login, personalization, a click, or a scrolling feed, record that limitation; a URL alone may not reproduce the visitor’s state.
4. Choose outputs and understand replay
ArchiveBox documents multiple outputs, including original HTML/CSS/JavaScript, SingleFile, PNG screenshots, PDF, WARC, metadata, media, and source-code clones. Available output depends on configuration and installed capture tools, so check your version’s documentation before planning around a particular format.
- Screenshot or PDF: convenient for visual reference and sharing. It records appearance rather than page behavior.
- SingleFile: convenient for carrying a page as one file, when its capture is suitable.
- HTML and assets: useful for inspecting saved content and links, with the caveat that referenced resources or scripts may not be self-contained.
- WARC: stores web crawl transactions. It can be useful in workflows designed for replay and exchange.
- WACZ: packages WARC data with descriptive context and a page index. Browsertrix documents export and import of archived items as WACZ, and names ReplayWeb.page as a browser viewer.
WARC and WACZ are not interchangeable terms: WACZ is a package that can include WARC files plus context and indexing information. Browsertrix’s concepts documentation describes these archived-item and replay concepts. After exporting or moving a package, open it in a compatible viewer and check the pages you care about.
5. Organize, back up, and maintain the collection
Keep the application data on storage you control. ArchiveBox documents local and remote storage possibilities, including S3/B2 and network storage. Choose storage based on your own collection and retention plan; there is no universal capacity figure because pages differ in media, capture frequency, and retained formats.
- Keep an independent backup of the archive data, ideally in a separate location from the running host.
- Preserve capture dates, original URLs, and useful notes.
- Periodically restore a sample from backup and verify that the files and replay workflow still work.
- Review whether scheduled imports are still collecting the sources you intended.
- Review access controls before exposing a public archive or its URL-add interface.
This is practical maintenance advice for a self-hosted, file-based collection. No particular setup guarantees long-term preservation, so verify that your own files remain readable and that your backup can be restored.
6. When browser-based capture is a better fit
Browsertrix captures pages through a real browser and supports archived-item workflows. Consider it when a browser-driven capture and WACZ export/import are central to your process. Its documentation cautions that complex social-media sites can be difficult to archive. As a result, test pages that rely on login state, personalization, dynamic feeds, bot defenses, or interaction instead of assuming a capture will preserve them.
Compare tools by how you collect URLs, whether captures are static or browser-driven, which formats are available, how you organize and replay results, where files live, and how much operational upkeep you can handle. Neither workflow guarantees a complete copy of every page.
7. Troubleshoot incomplete captures
| Symptom | Likely reason | What to try |
|---|---|---|
| Images or styles are missing | The relevant resources were not captured, or the page references remote assets that are unavailable during replay. | Inspect the available capture outputs and dependencies. Try another supported capture format and verify the result locally. |
| The saved page looks right but buttons do nothing | The output is static, or the capture method does not execute archived JavaScript during viewing. | Use a method whose replay behavior fits the need; ArchiveBox documents JavaScript execution during viewing for wget and DOM captures only. |
| A social page is blank or partial | Complex social-media sites are difficult to archive; content may depend on login, scripts, or changing feeds. | Capture a representative page, inspect it immediately, and keep a screenshot or PDF as a visual record when useful. |
| A capture is missing after a restart | Data may have been written to temporary or non-persistent container storage. | Check the configured data location and Docker Compose volume mapping, then confirm new files survive a restart. |
| A public archive exposes more than intended | Visibility or add-access settings may be too permissive. | Review the current ArchiveBox configuration and security documentation before making the service public. |
| An imported WACZ does not replay as expected | The package may be incomplete, the viewer may not support the workflow, or the required page data may not be present. | Confirm the export completed, use a compatible replay tool such as ReplayWeb.page, and inspect the package’s page index and contents. |
Performance, reliability, and cost
Capture time and storage depend on the page, number of formats, media, and collection schedule. Whole-site or frequent collection increases storage and processing needs compared with saving a few URLs. Measure the growth of your own archive over a representative period and set retention and backup expectations from that data; no fixed disk estimate applies to every collection.
Reliability comes from persistent storage, independent backups, and periodic restore checks. A successful capture command does not prove that every resource or behavior was preserved. Review important captures and retain an additional format when a visual snapshot and replayable source serve different needs.
ArchiveBox is open source and self-hosted; budget for the host, storage, backups, and maintenance that your deployment needs. Optional local storage can be an existing disk or storage you already operate. The research sources establish no required hardware model or universal operating cost.
Or skip the browser setup
For a screenshot of a page, ScreenshotNeo provides a website screenshot API and MCP server. It does not replace a web archive: use an archive workflow when you need retained crawl data or replay. A screenshot can be a straightforward visual record.
One GET request returns an image or PDF. See the ScreenshotNeo API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card.
FAQ
Does a personal archive preserve a website exactly?
No. A capture preserves what the chosen method collected. Missing resources, dynamic behavior, login state, and site changes can affect what you can view later.
Can I save only a few pages instead of crawling a site?
Yes. ArchiveBox documents one-at-a-time URL input as well as scheduled source imports. Start with the smallest collection that meets your need.
Should I keep both a screenshot and a replayable capture?
If appearance and later exploration both matter, keeping different formats can serve those separate purposes. Verify each output against your use case.
Can I make the archive public?
A public deployment needs deliberate access and security configuration. Review current project guidance before exposing the archive or its add interface.


