How to Save a Website with ArchiveBox Using a URL List
Save a list of URLs with ArchiveBox using standard input, Docker, or Docker Compose. Learn what gets stored, how to control crawl depth, and how to troubleshoot imports.
To save the URLs in a text file with ArchiveBox, run archivebox add < urls_to_archive.txt from your ArchiveBox data directory. ArchiveBox reads the URLs from standard input and creates a local snapshot for each submitted page. You can also pipe the file into the command, or pass standard input through Docker or Docker Compose.
1. Prepare ArchiveBox and your URL list
Install and initialize ArchiveBox using the official Quickstart for your setup. The Quickstart currently documents macOS and Ubuntu on amd64 or arm64, and Docker on Linux and macOS; check that page for current release support before installing on another platform.
Make a plain text file with one URL per line. For example, save this as urls_to_archive.txt:
https://example.com/
https://www.iana.org/domains/reserved
https://archive.org/
Use complete URLs, including https:// when the site supports it. Keep the file in a location you can read from the shell, and change into the directory where your ArchiveBox collection is initialized before adding pages.
2. Import the list from a terminal
Run this from the ArchiveBox data directory:
archivebox add < urls_to_archive.txt
Shell input redirection sends the file contents to ArchiveBox through standard input. The equivalent pipe form is:
cat urls_to_archive.txt | archivebox add
Both forms are documented in the ArchiveBox Usage guide. The commands submit the URLs in the file; they do not by themselves ask ArchiveBox to crawl every link found on those pages.
3. Run the same import with Docker
If you use the plain Docker workflow, mount the current directory as /data and keep standard input open with -i:
docker run -v "$PWD:/data" -i archivebox/archivebox:dev add < urls_to_archive.txt
For Docker Compose, the documented form is:
docker compose run --rm -T archivebox add < urls_to_archive.txt
Here -T disables Compose’s pseudo-TTY allocation so standard input can be read as supplied. These commands assume your Compose service is named archivebox and that you run them from the directory containing the Compose configuration. Use the installation instructions for your collection if its service name or data mount differs.
4. Check the saved snapshots
ArchiveBox stores snapshots as files in the local collection. The Quickstart points to the archive directory for browsing saved data and also describes the web interface. Depending on the configured extractors and what a particular page permits, a snapshot may contain outputs such as HTML, PDF, PNG, JSON, or WARC. Do not expect every format for every URL.
After an import, inspect the collection or open its web interface using the command and address described in the Quickstart for your installation. Preserve the full collection directory: ArchiveBox recommends backing it up before upgrading.
Input formats and crawl scope
Plain URL lists and URL-bearing files
A hand-written text file is not the only supported input. The Usage guide demonstrates feeding RSS or XML, Netscape-format bookmarks, JSON bookmark exports, and other text containing URLs to archivebox add. For instance, an exported bookmark file can be passed as standard input:
archivebox add < ~/Downloads/bookmarks_export.html
ArchiveBox parses URLs from these inputs; a bookmark export is not treated as a guarantee that every linked page can be captured. Use the official Usage guide for the current supported input examples and syntax.
Do you want to follow links?
The basic URL-list import submits the URLs you supplied. To expand the run to links one hop away, the Usage guide documents adding --depth=1:
archivebox add --depth=1 < urls_to_archive.txt
This expands the scope and can add many more pages. Leave the option off when you only want the listed URLs. Review the Usage guide’s depth documentation before choosing a deeper crawl.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
archivebox: command not found |
ArchiveBox is not installed in this shell, or its executable is not on the shell’s PATH. | Follow the Quickstart for your installation method. If you use Docker, run the documented container command instead of expecting a host executable. |
| The command reports that the collection is not initialized or the expected configuration is missing. | You ran the import outside the collection’s data directory, or initialization has not been completed. | Initialize the collection following the Quickstart, then run archivebox add in the directory used for that collection. |
| Docker starts, but no URLs are imported or standard input appears empty. | The container invocation did not preserve standard input, or input redirection points to the wrong file. | Use -i with the documented plain Docker command or -T with Compose. Check the filename and working directory, and confirm the list is not empty. |
| A line in the list is skipped or fails to produce a useful snapshot. | The line may not contain a valid page URL, the site may be unavailable, or the page may resist or block automated access. | Check that the URL is complete and reachable from the machine running ArchiveBox. Inspect that URL’s result in the collection; capturing a URL does not guarantee every extractor can produce output. |
| Only the listed pages appear, not pages linked from them. | The default import did not request recursive link discovery. | Add --depth=1 if one-hop link crawling is intended, and account for the expanded scope. |
| The collection is missing after a Docker run. | The container’s data directory may not be mounted to the host location you expected. | Check the -v mount or Compose volume configuration and inspect the mounted collection path. Keep using the same persistent data directory for later imports. |
Performance, reliability, and storage
Import time depends on the number of URLs, the pages’ response behavior, and the extractors enabled in your installation. A large list is usually easier to diagnose when divided into smaller batches, especially if one URL behaves unusually. Avoid adding --depth=1 unless you want the extra pages: a crawl can substantially increase the amount of work and stored data.
ArchiveBox saves local snapshot files, but the available formats and completeness vary by page and extractor. Keep a backup of the full collection before upgrades, as the project documentation recommends. The cited ArchiveBox documentation does not specify a universal storage estimate or preservation guarantee, so size a backup from your own collection rather than relying on a generic per-page figure.
Or skip the browser setup
If your goal is a screenshot or PDF rather than a locally managed preservation collection, ScreenshotNeo provides a website screenshot API and MCP server. Its one-call API returns an image or PDF for a URL. See the ScreenshotNeo API documentation for parameters and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Replace YOUR_API_KEY with your key. The Node.js example uses Bun’s file writer to save the response; with Node.js, use await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))) after the fetch.
- Cookie banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers identify the page verdict and whether the request was billed.
- An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
- 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 screenshots; all features are on every plan.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
FAQ
Can I put multiple URLs on one line?
For a plain text list, use one URL per line so the input is easy to inspect and correct.
Does ArchiveBox save a complete copy of every site?
No import can guarantee a complete copy of every page. Snapshot files and formats depend on the page and the extractors configured.
Should I use ArchiveBox or a screenshot API?
Use ArchiveBox when you want a self-hosted collection of snapshots. Use a screenshot API when your task is to request an image or PDF for a URL without managing the capture browser yourself.


