ScreenshotNeo

BlogHow-to

ArchiveBox API: How to Add URLs and Retrieve Capture Status

Add URLs to ArchiveBox with its CLI or local Python interface, authenticate to the REST API, and inspect snapshot records without assuming undocumented status fields.

By the ScreenshotNeo team4 October 20269 min read

To add a URL to ArchiveBox, use the documented archivebox add CLI or its local Python interface. For REST integrations, first inspect the API schema served by your own ArchiveBox installation at /api/v1/docs: the exact REST route and payload for adding a URL, and a universal capture-completion status field, are not established by the documentation covered here. To inspect records, the documented API guide shows GET /api/v1/core/snapshots.

ArchiveBox’s REST API has been available since v0.8.0, but the project labels it alpha. Treat routes, fields, and behavior as version-specific. Confirm them in your running instance’s docs before building an integration. The example address in the guide is http://api.archivebox.localhost:5797/api/v1/docs; use your own configured host and port. ArchiveBox’s project repository and its installation-specific API documentation are the sources to consult for version details.

1. Find the API schema for your installation

Open http://YOUR_ARCHIVEBOX_HOST:YOUR_PORT/api/v1/docs in a browser, replacing the host and port with those configured for your server. The interactive page describes the routes and schemas exposed by that installation.

  1. Check the ArchiveBox version and confirm the API documentation page is available.
  2. In the docs, locate the authentication route and the snapshot routes.
  3. Look for an operation that creates or adds a snapshot. Confirm its HTTP method, path, request body, required permissions, and response schema.
  4. To track a submission, inspect the returned response and the snapshot schema for lifecycle fields. Confirm their meanings in the deployed version before treating any field as “capture complete.”
  5. Try the documented request against a noncritical URL and verify the result in the ArchiveBox UI or with the snapshot listing call below.

Do not assume a route such as /api/v1/core/snapshots accepts POST requests just because it lists snapshots. A listing operation does not establish an add operation, and the docs reviewed here do not define one universal completion field.

2. Authenticate with a token

The official authentication guide describes creating a token in the Admin UI or obtaining one by posting a username and password to /api/v1/auth/get_api_token. Store the token as a secret and send it in the recommended bearer header.

curl -X POST 'http://YOUR_ARCHIVEBOX_HOST:YOUR_PORT/api/v1/auth/get_api_token' \\
  -H 'Content-Type: application/json' \\
  -d '{"username":"YOUR_USERNAME","password":"YOUR_PASSWORD"}'

Use the address for your installation. Read the token from the response according to the schema shown by its API docs; do not paste credentials into a shared script or commit them to source control.

The guide also documents X-ArchiveBox-API-Key for deployments where a reverse proxy consumes the bearer header. Query-string API keys are discouraged: anyone who obtains the URL may obtain the key and potentially perform API actions.

3. Add a URL with the documented CLI

For a local workflow, the CLI offers a direct documented way to submit URLs:

archivebox add 'https://example.com'

You can also pass URLs through standard input or a file:

echo 'https://example.com' | archivebox add
cat urls_to_archive.txt | archivebox add
archivebox add < urls_to_archive.txt

The CLI documentation also describes importing formats including RSS, XML, Netscape bookmarks, and text containing URLs. Its --depth=1 option can include one-hop outlinks:

archivebox add --depth=1 'https://example.com'

Use depth deliberately: following outlinks can expand the work beyond the URL you initially supplied. If you need an HTTP integration rather than a process running on the ArchiveBox host, identify and validate the add operation in /api/v1/docs first.

4. Add a URL through the local Python interface

For code running in the ArchiveBox environment with access to its data directory and Python installation, the usage docs show this pattern. It initializes Django, calls the add function, and prints the crawl ID and snapshot IDs.

import os
from pathlib import Path

DATA_DIR = Path("~/archivebox/data").expanduser()
os.chdir(DATA_DIR)

from archivebox.config.django import setup_django
setup_django(check_db=True)

from archivebox.cli.archivebox_add import add
crawl, snapshots = add(urls=["https://example.com"], index_only=True)
print(crawl.id, list(snapshots.values_list("id", flat=True)))

This is a local Python-library workflow, not a REST recipe. It needs the ArchiveBox environment and data directory. The documented example uses index_only=True; do not interpret its return values as proof that a full capture has completed.

5. List snapshots with the REST API

The authentication guide demonstrates listing snapshots with a bearer token. This is a runnable cURL request; replace the host, port, and token.

curl -X GET 'http://YOUR_ARCHIVEBOX_HOST:YOUR_PORT/api/v1/core/snapshots?limit=10' \\
  -H 'accept: application/json' \\
  -H 'Authorization: Bearer YOUR_API_TOKEN'

The response lets you inspect snapshot records. The documented example does not establish that a particular field means capture finished, nor does it specify whether adding a URL is synchronous. Inspect the actual response and schema from your version.

Python: retrieve the snapshot listing

This example uses requests to make the documented GET call. Install the dependency with python -m pip install requests, then set the host and token in your environment.

import os
import requests

base_url = os.environ["ARCHIVEBOX_BASE_URL"].rstrip("/")
token = os.environ["ARCHIVEBOX_API_TOKEN"]

response = requests.get(
    f"{base_url}/api/v1/core/snapshots",
    params={"limit": 10},
    headers={
        "Accept": "application/json",
        "Authorization": f"Bearer {token}",
    },
    timeout=30,
)
response.raise_for_status()
print(response.json())

Set ARCHIVEBOX_BASE_URL to the instance origin, for example http://api.archivebox.localhost:5797. The timeout is a client-side limit, not a statement about server performance.

Node.js: retrieve the snapshot listing

With a Node.js version that provides global fetch, set the same environment variables and run:

const baseUrl = process.env.ARCHIVEBOX_BASE_URL?.replace(/\\/$/, "");
const token = process.env.ARCHIVEBOX_API_TOKEN;

if (!baseUrl || !token) {
  throw new Error("Set ARCHIVEBOX_BASE_URL and ARCHIVEBOX_API_TOKEN");
}

const url = new URL(`${baseUrl}/api/v1/core/snapshots`);
url.searchParams.set("limit", "10");

const response = await fetch(url, {
  headers: {
    Accept: "application/json",
    Authorization: `Bearer ${token}`,
  },
});

if (!response.ok) {
  throw new Error(`ArchiveBox returned HTTP ${response.status}: ${await response.text()}`);
}

console.log(await response.json());

6. Determine capture status safely

Use the snapshot listing to inspect records, but do not infer a completion guarantee from the fact that a record exists. The reviewed API documentation does not define a universal status field, polling interval, or synchronous-versus-asynchronous contract.

  1. Read the live snapshot schema and the add operation’s response schema in your instance docs.
  2. Identify any lifecycle or status field and confirm its documented values and transition behavior for your installed version.
  3. If the add response returns an identifier, use the documented read operation for that record if one is provided.
  4. Choose polling behavior based on the documented lifecycle and your own operational limits. Stop when the documented terminal state is reached, or surface a timeout for manual investigation.
  5. For local operational checks, the installation guide suggests archivebox list and archivebox status. These help inspect snapshots and collection health; they are not documented as equivalents of a REST status field.

A robust client should distinguish “record found,” “capture still processing,” “capture completed,” and “capture failed” only when the installed API documents those distinctions. If it does not, report the record and its observed fields without inventing lifecycle semantics.

7. Choose REST, CLI, or Python

Method Best fit What it needs Version caution
REST API A separate service or script that communicates over HTTP Instance URL, token, and routes verified in /api/v1/docs The REST API is labeled alpha; check the deployed schema.
CLI Local shell automation, imports, and one-off additions Access to the ArchiveBox command and its configured environment Use the CLI docs for available options and installed version behavior.
Python interface Code running alongside ArchiveBox with access to its Python environment and data directory ArchiveBox installation, data directory, Django initialization The Python API is described as beta; confirm compatibility.

For an HTTP client, REST is the natural boundary, but only after confirming the add route and response on the actual server. For a single-host job, CLI avoids having to discover an undocumented REST payload. Python is useful when a local integration needs the documented library workflow and can share the ArchiveBox environment.

8. Troubleshooting

Symptom Likely cause What to do
/api/v1/docs is missing or unreachable Wrong host or port, API version mismatch, proxy routing, or the service is unavailable. Use the configured instance address, confirm the installed version, and check the deployment’s routing and service health.
Token request fails Incorrect credentials, wrong instance URL, or request body differs from the live schema. Check the auth operation in the instance docs and use the documented JSON fields. Prefer creating a token in the Admin UI when appropriate.
Authenticated request returns unauthorized Missing, expired, malformed, or incorrectly transmitted token. Send Authorization: Bearer TOKEN exactly as shown; verify the token and target instance. If a proxy consumes this header, check whether the documented X-ArchiveBox-API-Key alternative applies.
Snapshot listing works but URL submission does not A listing route does not imply a creation route or POST payload. Find the add operation in the deployed OpenAPI page. Verify method, path, payload, permissions, and response before sending requests.
A snapshot record exists but completion is unclear The listing response alone does not define a universal completion field. Inspect the live schema and lifecycle documentation. Use archivebox list or archivebox status for local checks, without treating them as REST status equivalents.
Python import or initialization fails The script is outside the ArchiveBox Python environment, uses the wrong data directory, or Django has not been initialized. Run in the installed environment, change to the correct data directory, and follow the documented setup_django(check_db=True) sequence.
Some URLs do not produce expected archives Target availability, network access, or an ArchiveBox extractor/capture issue may be involved. Inspect the relevant snapshot and installation logs using the tools available in your deployment. Avoid assuming a REST status value explains the cause unless the schema documents it.

9. Performance, reliability, and security notes

  • Keep work bounded. Imports and depth-based crawling can involve more URLs than a single submission. Use depth only when following outlinks is intended.
  • Design clients for uncertainty. Because REST is alpha and the live schema is instance-specific, validate responses and handle unexpected fields or HTTP errors. Do not treat undocumented values as stable contracts.
  • Do not equate a timeout with failure. A client timeout says the client stopped waiting; it does not by itself establish whether the server created a record or continued processing. Check the instance using documented read operations before retrying blindly.
  • Protect credentials. Keep tokens out of source control, logs, and shared URLs. Prefer the bearer header; use the documented API-key header only when the proxy setup calls for it.
  • Budget local resources. ArchiveBox stores preserved snapshot files. Monitor storage and collection health with the local operational commands documented for your installation.
  • Plan compatibility around the installed version. Check the instance docs after upgrades and validate the routes and fields your integration depends on.

Or skip the browser setup

If your goal is a screenshot rather than a self-hosted web archive, ScreenshotNeo provides a website screenshot API. The one-call request below returns an image; see the ScreenshotNeo API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say which outcome occurred. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

FAQ

Can I use ArchiveBox’s API to add a URL?

Use the add operation shown by your running instance’s /api/v1/docs. The exact REST route and payload were not verifiable in the documentation reviewed for this guide, so no endpoint is assumed here. The documented CLI alternative is archivebox add 'https://example.com'.

Does a snapshot record mean the capture has finished?

Not by itself. Confirm the lifecycle fields and their meanings in the schema for your installed version.

Where do I find ArchiveBox’s API docs?

Open /api/v1/docs on your ArchiveBox instance. The example host in the official guide is http://api.archivebox.localhost:5797; deployments use their own address.

Is the Python example an API request?

No. It is a local integration that initializes ArchiveBox’s Django environment and calls its Python add function.

Sources