How to Bulk Archive a Sitemap with ArchiveBox
Extract sitemap URLs into a reviewable list, then import them into ArchiveBox with stdin. Includes sitemap indexes, Docker, troubleshooting, and a screenshot option.
To bulk archive a sitemap with ArchiveBox, turn its page locations into a plain-text file with one URL per line, then run archivebox add < urls.txt from your ArchiveBox data directory. ArchiveBox accepts URL lists through stdin and documents support for RSS, XML, Netscape, and other import formats, but its docs do not promise that every sitemap shape or sitemap index will be parsed automatically. Extracting and reviewing the URLs first is the most predictable approach. ArchiveBox Usage documentation · Quickstart
1. Find the sitemap
Check the site’s robots.txt for a sitemap declaration, or use the sitemap URL you already know. A sitemap index contains references to child sitemap files; those child files contain the page URLs. Expand the index and collect the page-level <loc> values if the goal is to archive the listed pages.
Do not assume the sitemap index itself is the complete page list. Also, importing a sitemap URL as an ordinary page does not necessarily mean ArchiveBox will expand every entry. ArchiveBox documents XML/feed imports generally, but compatibility with each sitemap variant can depend on the installed version and parser plugins.
2. Extract URLs to a text file
Use a sitemap-aware XML parser or script that handles XML namespaces and sitemap indexes. Write only the page locations to urls.txt, one absolute URL per line. Preserve the URLs as published and review the resulting list for duplicates, staging URLs, unwanted sections, or URLs outside the scope you intend to preserve.
https://example.com/docs/
https://example.com/blog/article-a
https://example.com/blog/article-b
If you already have an XML file and the ArchiveBox parser for your installed version is confirmed to handle its format, you can test direct ingestion on a small sample. When uncertain, use the text list: it is easy to inspect and filter before submission.
3. Import the URL list with ArchiveBox
Run the command from the ArchiveBox data directory (or use the equivalent configured path for your installation):
archivebox add < urls.txt
The official quickstart also shows these stdin forms:
cat urls.txt | archivebox add
# Plain Docker: mount your ArchiveBox data directory at /data.
cat urls.txt | docker run --rm -v "$PWD:/data" -i archivebox/archivebox:dev add
# Docker Compose: -T allows stdin to be piped to the service.
cat urls.txt | docker compose run --rm -T archivebox add
For Docker, include -i so the container receives piped stdin. For Compose, the documented invocation uses -T. Confirm the mounted path is the actual ArchiveBox data directory, rather than an unrelated working folder. See the ArchiveBox Quickstart.
4. Verify what was archived
Importing links submits them to ArchiveBox; it does not guarantee every target will be captured successfully. Review the command output and the resulting archive for omissions or failures. Check a sample of saved pages, especially if the sitemap is large or the pages matter for compliance, research, or long-term preservation.
Pages can fail because they were removed, redirect unexpectedly, require authentication, block automated retrieval, or time out. Keep the original sitemap and generated URL list so you can compare intended inputs with archived results and retry selected URLs later.
Choose direct XML import or an extracted list
| Input path | Use when | Check |
|---|---|---|
| Direct XML/feed through stdin | Your installed ArchiveBox version and parser support the exact input format. | Test representative entries and confirm index expansion and resulting pages. |
| One-URL-per-line text file | You need predictable, reviewable inputs or parser support is unclear. | Review scope, duplicates, and URL formatting before importing. |
ArchiveBox’s usage guide says RSS, XML, Netscape, and other supported import formats can be piped to stdin. That does not establish universal support for every sitemap namespace, index structure, or file size. Check the installed version’s help and parser options when relying on direct XML ingestion. ArchiveBox Usage · Repository overview
Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| The sitemap appears as one archived page instead of many. | The input was treated as a page URL or its XML structure was not expanded by the active parser. | Extract the page <loc> values into urls.txt and use archivebox add < urls.txt. |
| Only some pages appear in the archive. | Some targets may be removed, gated, blocked, redirected, or otherwise unreachable; successful capture is not guaranteed. | Compare output with the input list, inspect failures, and retry applicable URLs individually. |
| Docker receives no URLs. | stdin was not forwarded into the container. | Use -i with docker run and mount the correct data directory. |
| Compose does not pass the piped list. | The service invocation is using an unsuitable TTY mode. | Use the documented docker compose run --rm -T archivebox add form. |
| Nested sitemap pages are missing. | Only the sitemap index was processed; child sitemaps were not expanded. | Fetch the child sitemaps, extract their page locations, and combine them into the URL list. |
| The import contains unwanted pages or duplicates. | The sitemap includes multiple sections or overlapping child files. | Normalize and review the extracted URL list before importing; retain a copy for audit and retry work. |
Performance, reliability, and cost considerations
Large sitemaps mean many independent page captures. The reviewed ArchiveBox documentation does not specify a universal batch size, throughput, or timeout for sitemap imports, so avoid assuming a fixed safe limit. For a recurring or very large archive, process a representative subset first, monitor resource use and failures in your own deployment, and proceed in manageable batches if needed.
Keep the source sitemap and a dated copy of the extracted list. That makes later comparisons and retries possible when the sitemap changes or individual captures fail. If the source requires a logged-in session or blocks automated retrieval, submission alone will not make the page accessible; address access using the ArchiveBox setup appropriate to your installation and verify the saved result.
Or skip the browser setup
ArchiveBox is for preserving pages in your own archive. If your immediate job is to turn URLs into screenshots or PDFs, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its cookie and consent handling accepts banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with page verdict and billing information in response headers. AI agents can use its MCP tools take_screenshot, get_page_info, and capture_pdf.
One GET request captures a URL. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
The Node.js example uses Bun’s file writer to save the response. In Node.js without Bun, save the response body using your preferred filesystem method.
ScreenshotNeo includes full-page capture with lazy images loaded, CSS selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, HTML/CSS capture, custom CSS and JavaScript, click-before-capture, selector hiding, wait conditions, request and resource blocking, headers, cookies, user agent and authorization, timezone and geolocation, transparent background, image resizing, configurable caching, signed image links, async jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names also work with those used by other screenshot APIs. Free includes 1,000 shots/month without a card; paid plans start at $5 for 3,000. Every feature is on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Can ArchiveBox archive every URL in a sitemap index automatically?
Do not assume so. Expand child sitemaps and verify the imported pages unless your installed parser has been confirmed to handle that index format.
Does adding a URL mean the page was successfully saved?
No. Inspect the command output and archive; access restrictions, redirects, removed pages, or retrieval failures can prevent a successful capture.
Should I use ArchiveBox or a screenshot API?
Use ArchiveBox when you want a self-hosted web archive. Use a screenshot API when the deliverable is an image or PDF capture rather than an ArchiveBox archive.


