ScreenshotNeo

BlogHow-to

How to Restore an ArchiveBox Collection from a Backup

Restore ArchiveBox by reconnecting the complete collection directory, including its SQLite index and archived files. Verify the deployment, permissions, and snapshots before reopening access.

By the ScreenshotNeo team4 October 20268 min read

Short answer: Restore the complete ArchiveBox collection directory, then point ArchiveBox at that directory using the same deployment style or a compatible one. The SQLite index and archived files belong together. A static HTML or JSON index export is not a complete collection backup. The ArchiveBox README describes collection state as a single folder containing the database, content, configuration, and logs.

1. Identify the backup you need

Before changing a deployment, preserve the original backup and make a working copy. The exact copy procedure depends on your operating system, backup tool, and ArchiveBox version. ArchiveBox does not document one universal restore command or a consistency procedure that applies to every backup method.

A full collection backup should include the collection root and its persistent contents. Look for:

  • index.sqlite3, the collection’s SQLite index.
  • ArchiveBox.conf, if your collection uses it.
  • The archived content tree. The current README describes snapshot data under data/archive/users/; deployment layouts may present the archive directory at the collection root.
  • Other persistent directories present in your deployment, such as personas/, sonic/, and logs/.

Do not mistake /tmp/archivebox for the collection: the Docker deployment documentation identifies it as runtime state. Back up the persistent data directory instead. The exact directory layout can vary with ArchiveBox version and deployment; preserve the directory as a unit rather than selecting files by name.

2. Restore the directory without overwriting your only copy

  1. Choose a destination directory for the restored collection. Keep the source backup untouched.
  2. Copy or extract the complete collection into that destination, preserving nested paths and file metadata where your backup tool supports it.
  3. Inspect the restored root. Confirm that the SQLite database and archived content tree are present, along with any configuration and other persistent folders used by this collection.
  4. If the backup’s database consistency is uncertain, do not delete the source or assume the copy is usable. Retain it and consult guidance matching the ArchiveBox version and backup method that created it.

The following is an illustrative shell copy for a directory-based backup on a Unix-like system. Replace both paths with your own. It is not an ArchiveBox restore command; use the restore or extraction procedure supplied by your backup tool when appropriate.

cp -a /path/to/backup/collection /path/to/restore/collection

If your backup is an archive file, extract it so the restored destination contains the collection directory itself—not an accidental extra nesting level or only selected files. Check the paths before starting the service.

3. Point ArchiveBox at the restored collection

Docker Compose

In Docker, the host directory containing the restored collection must be mounted at the container’s data path. The current Docker deployment example uses ./data:/data. For example, if the restored collection is in /srv/archivebox-restored, the relevant Compose volume mapping can look like this:

services:
  archivebox:
    volumes:
      - /srv/archivebox-restored:/data

Use the rest of your existing Compose configuration, including the image version and any environment settings required by your deployment. Confirm that the left side names the restored host directory and the right side is the application’s data path. Start the service using your normal Compose workflow after checking the mapping.

Docker run

For a deployment managed with docker run, mount the restored host directory at /data:

docker run -v /srv/archivebox-restored:/data YOUR_ARCHIVEBOX_IMAGE

YOUR_ARCHIVEBOX_IMAGE is a placeholder, not a literal image name. Use the image and version appropriate to your installation, along with the options your deployment needs.

Non-Docker installation

For a non-Docker installation, select or enter the restored collection directory using the method appropriate to your installation and version. ArchiveBox documents that Docker and non-Docker installations can share a data directory format, but that does not remove version-specific migration considerations.

Do not point a fresh initialization at an empty directory and treat that as a restore. The current Docker image initializes automatically at startup; it cannot recover archived files that were not placed in the mounted collection directory.

4. Check file ownership and access

A container may be unable to use files whose numeric owner or group does not match the service identity. The Docker entrypoint uses the collection owner’s numeric UID and GID; deployment documentation describes PUID and PGID overrides for filesystems that require specific IDs.

  • Check that the ArchiveBox process can read the database and snapshot files.
  • Check that it can write where the deployment needs to update collection state.
  • If your mounted filesystem needs specific numeric IDs, configure the documented PUID and PGID values for your deployment.
  • Avoid broad permission changes such as making private archive contents world-readable. Preserve the collection’s intended access controls.

5. Start the service and verify the recovered collection

Once the mount or data directory and ownership are correct, start ArchiveBox using your normal deployment procedure. The project README documents archivebox version; the Docker deployment documentation documents archivebox status. Run the command supported by your installation, for example:

archivebox version
archivebox status

These commands are documented checks; the sequence below is a practical verification checklist, not a quoted official recovery checklist:

  1. Confirm the service starts without database, path, or permission errors.
  2. Open the collection in the UI or inspect it through the supported interface for your version.
  3. Check several known snapshots, including older and newer entries if available.
  4. Verify that expected snapshot files are present in the filesystem, such as HTML, screenshots, media, WARC files, or repository outputs where those extractors were used.
  5. Compare a few known items against the backup inventory. A page in the index does not guarantee every expected extractor output survived.

Keep the original backup until you have verified the records and files you need.

6. Know what an index export can and cannot restore

ArchiveBox can export its index with archivebox list --html, archivebox list --json, or archivebox list --csv. These exports can help browse or recover index information, but they are not documented as full collection backups: they do not replace index.sqlite3 and the snapshot files.

The README notes that exports are not paginated and that relative paths require the export to remain alongside the archive folder if you want those paths to work. Treat an export as a useful companion to a backup, not a substitute for the collection directory.

7. Handle older collections carefully

Current releases organize snapshots under data/archive/users/. Older collections may use legacy timestamp directories. The current README documents archivebox update --migrate-only for legacy timestamp directories:

archivebox update --migrate-only

Use this only after identifying the collection layout and checking the instructions for the version you are restoring to. Preserve the untouched backup before migration. Database migrations and upgrade behavior are version-dependent, so do not assume that the newest image can safely open every older backup without consulting the matching release guidance.

8. Moving between Docker and non-Docker

ArchiveBox states that Docker and non-Docker installations share a data directory format, which supports moving a collection between those installation styles. For Docker, mount the restored directory at the container data path. For a non-Docker installation, select that directory through the installation’s supported mechanism. In either direction, check numeric file ownership, the collection’s layout, and version-specific migration guidance before opening it.

Common restore problems

Symptom Likely cause What to check or do
ArchiveBox appears empty or creates a new collection The service is pointed at an empty or incorrect directory, or the mount’s host path is wrong. Verify that the restored host directory is mounted at the actual data path (commonly /data in Docker) and that index.sqlite3 is at the expected collection root.
Snapshots appear in the index but saved files are missing Only the database or an index export was restored; the archive tree was omitted or incomplete. Restore the complete collection directory, including the archived content tree, from the backup.
Snapshot files exist but the collection is missing or unusable The archive tree was copied without its index, or the database is missing or inconsistent. Restore the matching database and files from the same backup set. Keep the source backup intact if consistency is uncertain.
Permission denied or database cannot be opened The container’s numeric user or group does not have access to the mounted files. Check ownership and read/write access; configure documented PUID/PGID overrides when required by the mounted filesystem.
Startup reports migration or schema errors The collection is from a different or older ArchiveBox version. Stop and consult the instructions for the source and target releases. Keep an untouched backup; determine whether the legacy layout needs the documented migrate-only operation.
Exported links do not open snapshots The export’s relative paths no longer point to the archive folder, or only the export file was moved. Keep the export alongside the archive folder as documented, or use the complete collection in ArchiveBox.
Temporary files were backed up, but the collection is absent The backup captured runtime state such as /tmp/archivebox instead of persistent collection data. Locate and restore the persistent collection directory. Runtime state is not a replacement for the database and archived content.

Performance, reliability, and privacy notes

  • Restore time: The amount of data to copy is determined by the collection, and a large archive can take time to copy or extract. The cited project materials do not provide a universal restore-time benchmark. Avoid interrupting the backup tool’s operation and verify that the copy completed.
  • Consistency: A directory copy is useful only if the database and files represent a usable backup together. The official guidance reviewed does not prescribe a consistency procedure for every backup tool. Keep a pristine backup and use instructions for the method and version that produced it when database consistency is in doubt.
  • Version compatibility: Moving the data format between Docker and non-Docker is supported, but older layouts and database migrations require release-specific care. Check version guidance before migration.
  • Access control: Collections can contain private URLs, cookies, session tokens, and private archived content. Recheck filesystem permissions, user access, and network exposure before making the restored service available.

Or skip the browser setup

If your goal is to capture a current webpage rather than recover an existing ArchiveBox collection, ScreenshotNeo returns a screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, and failed loads are never billed.
  • An MCP server lets AI agents use screenshot tools.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Is there a special ArchiveBox restore command?

The official materials reviewed do not describe a universal restore command. Restore the complete collection directory and reconnect the deployment to it.

Can I restore from an HTML or JSON export alone?

No. An index export is useful for browsing or index information, but it is not the database and archived files that make up the complete collection.

Can I use a Docker backup with a non-Docker installation?

The project says both installation styles share a data directory format. Check the target installation’s version requirements, permissions, and any legacy migration needs.

Should I make the restored archive public?

Only if that matches your intended access policy. Recheck permissions and network exposure because archived content may include private material and session data.