ScreenshotNeo

BlogHow-to

How to Schedule Recurring Website Captures with ArchiveBox

Schedule recurring ArchiveBox imports with interval aliases or cron expressions, understand the built-in orchestrator, and fix common scheduling issues.

By the ScreenshotNeo team4 October 20266 min read

Use archivebox schedule to create a recurring import. For example, archivebox schedule --every=daily --depth=1 https://example.com/feed.xml schedules a daily import. Current ArchiveBox versions store schedules in the database, and a running archivebox server uses its built-in orchestrator to queue work when schedules are due. You generally do not need a separate scheduler container or host cron for this current workflow. See the current ArchiveBox scheduling documentation.

1. Check your ArchiveBox version and collection

Run scheduling commands from the ArchiveBox collection context, using the same installation or Docker Compose setup that serves the collection. Scheduling flags and architecture have changed across versions, so check the command help before copying an older cron recipe:

archivebox --version
archivebox schedule --help

Current documentation describes database-backed schedules and a global orchestrator. Older ArchiveBox 0.7.4 documentation describes external task schedulers such as cron, at, or systemd. Treat those as version-specific historical instructions rather than the current default. [Current schedule guide; ArchiveBox 0.7.4 docs]

2. Add a recurring import

Choose a URL, feed, or other import path supported by your installed version. This daily feed example limits link following to depth 1:

archivebox schedule --every=daily --depth=1 https://example.com/feed.xml

You can use interval aliases such as minute, hour, day, week, month, year, daily, weekly, monthly, and yearly. The docs also show cron expressions. This example schedules a run every six hours:

archivebox schedule --every='0 */6 * * *' https://example.com/feed.xml

Use --depth=1 when you want the starting feed or page and its directly linked pages included. Decide the crawl scope deliberately, and use the URL allow and deny list configuration documented by ArchiveBox when you need to constrain which linked pages are followed. A schedule cannot guarantee that a third-party site will permit access or render successfully.

3. Inspect and manage schedules

List configured schedules and their state with:

archivebox schedule --show

The current schedule command also documents these management options:

Option Purpose
--show Show configured schedules.
--run-all Immediately enqueue enabled schedules.
--clear Clear schedules.
--foreground Run the global orchestrator outside archivebox server.

Confirm exact option behavior with archivebox schedule --help for the installed version before using management flags, especially --clear.

4. Keep the scheduler running

In the current server-managed setup, the running archivebox server process monitors enabled schedules. When one is due, the orchestrator creates a queued crawl; the queued job then follows the normal processing path. A one-shot archivebox add does not also sweep unrelated scheduled jobs. If you are not running the server, the docs provide --foreground to run the global orchestrator outside it.

The current project documentation says a running server removes the need for host cron, user crontabs, or a separate archivebox_scheduler container. The current Docker Compose flow likewise does not require a dedicated scheduler service. [Scheduled Archiving; ArchiveBox repository]

5. Docker Compose examples

To create a schedule with a one-off Compose command, the documented pattern is:

docker compose run --rm archivebox schedule --every=weekly --depth=1 https://example.com/feed.xml

When the main server is already running, its orchestrator handles future due runs. The image repository also documents adding a schedule through the running service:

docker compose exec archivebox archivebox schedule --add --every=day --depth=1 'https://example.com/feed.xml'
docker compose exec archivebox archivebox schedule --show

These examples reflect documented syntax from current project and image sources; use the syntax matching your installed image version. [Schedule guide; Image and project docs]

6. Storage, reliability, and operating cost

ArchiveBox can save several representations of a snapshot, including original HTML, CSS and JavaScript, single-file HTML, screenshots, PDF, WARC, titles, article text, favicons, and headers. A recurring crawl therefore grows the archive according to the sites and capture outputs involved; the documentation does not provide a universal storage estimate, so monitor your own collection rather than assuming a fixed size per run. Snapshot contents are stored as ordinary files. [ArchiveBox 0.9.71 documentation]

The Docker deployment guide says persistent data includes the database and archive. It advises keeping the database and browser profiles on reliable local storage and describes mounting remote storage for /data/archive. An external drive is optional for local capacity; it is not required to schedule jobs. [Docker deployment guidance]

  • Choose a cadence that matches how often the source changes and how much archive growth you can maintain.
  • Keep the server or foreground orchestrator running for due schedules to be noticed.
  • Limit crawl depth and URL scope to avoid archiving more pages than intended.
  • Expect network failures, access restrictions, and rendering differences on third-party sites; scheduling only queues capture work.

7. Troubleshooting

Symptom Likely cause What to do
schedule rejects an option or alias The installed version uses different command syntax. Check archivebox --version and archivebox schedule --help; follow docs for that version.
The schedule exists, but no new crawl appears The running server/orchestrator may not be active, or the next due time has not arrived. Confirm the server is running, inspect archivebox schedule --show, and use the version’s documented --run-all behavior to enqueue enabled schedules immediately if appropriate.
A one-shot import runs, but scheduled work does not archivebox add processes that import and does not sweep unrelated schedules. Keep the server orchestrator running or invoke the documented global orchestrator flow.
Compose command works on one deployment but not another Compose service names, image versions, or supported flags differ. Check the service name in your Compose file and consult docs for the installed image; compare run --rm and exec examples only where supported.
Some pages are absent from a capture The crawl depth or allow/deny rules exclude them, or the site did not allow or render the request. Review depth and URL rules, then inspect the resulting snapshot and the site’s behavior. Do not assume every linked page is capturable.
The archive consumes more disk than expected Recurring runs save multiple snapshot formats and the collection accumulates over time. Review enabled capture outputs and crawl scope, monitor archive storage, and plan persistent storage capacity.

Or skip the browser setup

If you need a clean screenshot of a page rather than a durable, multi-format ArchiveBox archive, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Does the schedule capture the site at an exact wall-clock time?

Use the interval or cron format supported by your installed version and inspect the schedule state. The documented flow describes due schedules being queued by the orchestrator; it does not promise that capture processing finishes at the scheduled instant.

Do I need an external drive or remote storage?

No. Those are storage choices as the archive grows, not scheduling prerequisites. Follow the deployment guidance for where to keep database, browser profile, and archive data.

Can I use ArchiveBox scheduling to guarantee a complete site copy?

No. Scheduling repeats an import configuration. Crawl depth, URL rules, site behavior, and successful access determine what is captured.