ArchiveBox disk space usage: how to find and remove large captures
Find what is using space in ArchiveBox, locate large snapshots with standard shell tools, and remove captures through the supported CLI.
To find what is using space in ArchiveBox, first locate the data directory, then measure its archive/ tree and drill into snapshot directories with ordinary filesystem tools such as du. ArchiveBox does not document a built-in command that sorts snapshots by disk size. To remove a known snapshot, identify the exact URL and use archivebox remove --yes URL; avoid deleting its directory directly because ArchiveBox’s index and files are related state.
ArchiveBox stores its main SQLite index at index.sqlite3 by default, alongside configuration and the archive/ tree containing snapshot directories and extractor output. In current documented layouts, snapshots are sharded beneath archive/users/<user>/snapshots/.... Check the layout and command syntax for your installed version using archivebox help and the ArchiveBox Usage documentation.
1. Find the data directory and the filesystem that holds it
Before measuring anything, establish where ArchiveBox actually writes data. The path may be set by OUTPUT_DIR, a configuration file, a container volume, or a host mount. In Docker, the path inside the container and the corresponding host path are different views of the same storage. A full container root filesystem reading does not necessarily mean the archive volume is full, and the reverse can also happen.
- Identify the ArchiveBox data/output directory for your installation.
- Check whether
archive/is on its own bind mount, network filesystem, or separate disk. - Run the measurements on the host or mounted filesystem that owns the data. In a container, inspect both the container path and its host volume mapping.
- Make sure your account can read the directories. Permission errors can make a size report incomplete.
The project’s documented directory example includes the index, configuration, and archive output tree. Do not assume every installation has the same path or exact directory layout.
2. Measure ArchiveBox storage with shell tools
The commands below are general operating-system tools, not ArchiveBox features. Replace /path/to/data with your actual data directory.
Linux and macOS
# Overall data directory and archive outputs
du -sh /path/to/data /path/to/data/archive
# One level below archive/, largest entries first (GNU/Linux)
du -sh /path/to/data/archive/* 2>/dev/null | sort -h
# macOS: summarize immediate child directories, then sort by size
du -sh /path/to/data/archive/* 2>/dev/null | sort -h
The leading space before du is harmless in common shells, but you can omit it. The macOS and GNU tools differ in some options; the examples use options available in their typical installations. If wildcard expansion omits hidden entries, or there are too many children for the shell to expand, measure a specific subtree at a time.
Drill down through snapshot directories
Current documented snapshot layouts are sharded by user, date, and domain, so a single archive/* level may not show individual snapshots. Measure progressively deeper directories:
# Inspect one level at a time; adapt each path to your layout
du -sh /path/to/data/archive/users/* 2>/dev/null | sort -h
du -sh /path/to/data/archive/users/USER/snapshots/* 2>/dev/null | sort -h
Continue into the date and domain directories until you reach snapshot UUID directories. A broad scan can take time on large collections or network storage. Restrict it to likely areas when possible.
For a GNU/Linux host with GNU du, this command lists directories below a selected tree and sorts them by reported size:
find /path/to/data/archive -type d -print0 \
| xargs -0 du -s 2>/dev/null \
| sort -n \
| tail -50
This is a general shell pipeline. It may report nested directories as well as snapshots, and hard links or filesystem compression can make apparent size differ from physical space reclaimed. Treat its output as a way to find candidates, not as an ArchiveBox-generated report. On macOS, the available find, xargs, and du options vary; the simple one-level du -sh commands are more portable.
Understand what the largest files belong to
Snapshot directories can contain an index file and outputs from different extractors. The Usage documentation gives examples such as wget/warc/, ytdlp/media/, and git/. Media downloads can dominate space. The ArchiveBox repository gives a very broad estimate of roughly 1 GB to 50 GB per 1,000 snapshots, with video/audio saving and the YTDLP_MAX_SIZE limit among the factors. This is a project estimate, not a per-snapshot promise; actual usage depends heavily on your content and settings. See the ArchiveBox project repository.
3. Identify the exact snapshot before deleting it
Size alone is not enough to choose a deletion target. Match a candidate directory to the intended snapshot using ArchiveBox’s list or UI and confirm the URL or snapshot identifier. Review the snapshot details before proceeding. If the capture matters, make a backup and verify that it can be restored.
Deleting a snapshot is irreversible through the documented UI flow. Also, removing its snapshot output may not erase all traces of the URL: imported source lists under sources/, operational logs under logs/, and an external search backend can retain related data. If your goal is privacy erasure rather than reclaiming archive space, account for those stores separately and follow the retention rules that apply to your archive.
4. Remove a known snapshot through ArchiveBox
For a known URL, the Security Overview documents this command:
archivebox remove --yes 'https://example.com/article'
Replace the example with the exact URL for the snapshot you intend to remove. The command deletes matching Snapshot rows and schedules their directories for cleanup through ArchiveBox’s normal state-machine path. The legacy --delete flag is accepted for CLI compatibility but does not change that behavior. Check archivebox help remove or the relevant CLI help for your installed version, since command syntax can change. See the ArchiveBox Security Overview.
You can also use the UI’s Delete action where available; the Usage documentation says it removes a snapshot and its archive results and cannot be undone. Do not use rm -rf on a snapshot directory as the normal removal method. Manual filesystem deletion can leave the index and files out of sync. Reserve it for a version-specific recovery procedure, with a backup and a verified database state.
After removal, remeasure the archive tree and check the owning filesystem’s free space. Cleanup may be scheduled through the application path, and remote storage may report reclaimed space differently or after a delay. Confirm that ArchiveBox’s non-root user has permission to remove files on the mounted storage.
5. Reduce future storage growth
Disable extractors you do not need
ArchiveBox documents disabling unused extractors as one way to reduce storage. Decide which outputs your archive needs: media, Git repositories, or other extractor results can add substantial data. Turning an extractor off reduces the outputs you preserve, so make that choice based on your archival requirements rather than treating it as a universal cleanup switch.
Choose storage placement deliberately
The ArchiveBox storage guidance recommends keeping the SQLite index on reliable local storage; bulk archive outputs can be placed on a larger HDD or suitable remote filesystem. That can trade lower-cost capacity for slower access and more storage administration. For Docker, NFS, SMB, or FUSE setups, confirm UID/GID mappings or ACLs let the ArchiveBox non-root user create and remove files. See Setting Up Storage.
Use retention only as an explicit deletion policy
DELETE_AFTER can remove Crawls, Snapshots, ArchiveResults, and Process rows with their on-disk outputs after the configured duration. The most specific setting takes precedence across global, persona, crawl, and snapshot levels. The current Configuration page says 0, an empty value, or None disables automatic deletion by default: ArchiveBox does not delete anything unless asked. Retention is destructive and irreversible, so define the intended scope, test your backup/recovery process, and review the exact configuration for your version before enabling it. See ArchiveBox Configuration.
Consider filesystem compression or deduplication carefully
The project mentions filesystem compression and deduplication approaches such as ZFS/BTRFS, fdupes, or rdfind. These are system-level options, not ArchiveBox cleanup controls. Their results depend on the data and filesystem, and deduplication tools do not understand ArchiveBox’s application state. Use them only if you can operate and maintain the chosen storage setup safely.
Or skip the browser setup
If the task is capturing a website before it enters your ArchiveBox workflow, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It returns a PNG, JPEG, WebP, or PDF from one request, with options for full-page capture, selectors, viewport and device settings, custom CSS or JavaScript, waits, headers and cookies, caching, and more. Check the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. Sign up for free and get 1,000 screenshots a month with no card.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
du reports less space than the filesystem appears to use |
Deleted files may still be held open by a process; filesystem metadata, snapshots, hard links, or reserved space can also affect reported usage. | Check the owning filesystem’s own usage tools and whether a process still has deleted files open. On remote storage, inspect usage at the server or mount that owns the data. |
| The archive path looks unexpectedly small | You measured the container filesystem instead of the mounted host path, or used the wrong data directory. | Confirm OUTPUT_DIR, Docker volume mappings, and the mount source; measure the host-side path too. |
| Some directories show permission errors | The reporting user cannot traverse all snapshot directories, or remote storage permissions do not match ArchiveBox’s user. | Use an account with read access for measurement. Check non-root UID/GID mappings and ACLs for cleanup operations. |
| A directory is large but you cannot identify its URL | You have reached a UUID or shard directory without matching it to the ArchiveBox record. | Use ArchiveBox’s list or UI to identify the snapshot, then confirm its URL or identifier before removal. |
| The removal command is rejected or behaves differently | CLI syntax or behavior differs in your installed release. | Run archivebox help and the relevant subcommand help; consult documentation for that version before acting. |
| The index entry disappears but disk space does not immediately change | Directory cleanup may be scheduled, the wrong mount was measured, or the filesystem is remote or retaining open/deleted data. | Recheck the application cleanup state, the actual archive mount, and free space reported by the storage owner. |
| Space fills again quickly | Media or other extractors are saving outputs, new snapshots continue to accumulate, or the archive volume is smaller than the workload needs. | Review enabled extractors and YTDLP_MAX_SIZE, measure growth by subtree, and set retention only if its destructive policy fits your needs. |
Performance, reliability, and cost notes
- Large scans take time: recursive directory walks can be slow on a large archive, HDD, or network mount. Start at the top level and narrow the scan.
- Keep the index reliable: ArchiveBox guidance favors reliable local storage for SQLite, while bulk archive outputs can use a suitable HDD or remote filesystem.
- Reported size is not always reclaimed size: compression, hard links, filesystem snapshots, and remote storage behavior affect the relationship between directory totals and free space.
- There is no universal cost per capture: content and extractor choices vary. ArchiveBox’s broad storage estimate is not a budget guarantee. Estimate from your own collection and leave capacity headroom.
- Protect against accidental loss: use the application removal path for known records, and keep a tested backup when captures or index history matter.
Frequently asked questions
Does ArchiveBox have a built-in largest-captures report?
The consulted official documentation does not describe a per-snapshot size report sorted by size. Use filesystem tools to find large directories, then map candidates back to ArchiveBox records.
Will removing a snapshot erase every record of its URL?
Not necessarily. Imported source files, logs, and an external search backend can retain related information even after snapshot outputs are removed.
Does ArchiveBox automatically delete old captures?
Automatic deletion is disabled by default. DELETE_AFTER can enable destructive retention behavior; check its scope and your installed version before setting it.


