ScreenshotNeo

BlogComparisons

ArchiveBox vs. SingleFile: differences in how they save web pages

ArchiveBox manages an organized collection of web captures; SingleFile saves a page as portable HTML. Compare their outputs, workflows, limits, and best uses.

By the ScreenshotNeo team4 October 20268 min read

Short answer: ArchiveBox is for collecting URLs and organizing their captures in a self-hosted archive. SingleFile is for saving a page as one portable, self-contained HTML file. ArchiveBox can also produce a SingleFile snapshot, so you can use both: ArchiveBox manages the collection, while SingleFile is one of the saved representations.

Choose SingleFile for a quick offline copy of a page you are viewing. Choose ArchiveBox when you want to import many URLs, keep snapshots organized, or produce several formats for each page. Neither guarantees that a complex, interactive site will behave exactly as it does online.

1. The main difference

Question ArchiveBox SingleFile
What is it for? Building and managing a self-hosted collection of URL snapshots. Saving a web page as a self-contained HTML file.
How do you start? Import URLs into an archive through its CLI, web interface, or supported sources. Save the current page, selected content, frames, or multiple tabs using its browser extension; it also has a CLI.
Where do captures go? Into snapshot folders in the local archive. Usually to the browser’s configured downloads location, with other destinations configurable.
What formats can you keep? Depending on configured capture methods and installed dependencies, outputs can include HTML captures, PDF, PNG screenshots, WARC, rendered DOM, article text, and extracted media or metadata. Primarily one self-contained HTML file; the project also documents packaging options such as self-extracting ZIP.
Is it suited to an ongoing collection? Yes. URL imports, organization, and scheduled workflows support collection building. It supports multiple-tab and automatic capture, but its central output is an individual saved page.

The best-fit descriptions are based on the documented features of each project, not a controlled head-to-head test. Capture results vary with the site, enabled methods, and local configuration.

2. How ArchiveBox saves web pages

ArchiveBox takes URLs as input and stores capture results in snapshot folders inside a local archive. It can run multiple capture methods for a URL, giving you different representations rather than one canonical file. Its documented outputs include a SingleFile HTML snapshot, wget-style clone and WARC data, rendered DOM, PDF, PNG screenshot, article text, title, favicon, response headers, and extracted media where supported.

The exact files depend on which extractors are enabled, whether their dependencies are installed, and how the target site behaves. The presence of an output type in the feature list does not mean every method succeeds for every page.

URLs can be added manually or imported from sources such as bookmarks, browser history, RSS, and link-saving services. ArchiveBox offers a CLI and a self-hosted web interface; scheduled ingestion can support an archive that grows over time. See the ArchiveBox project and its documentation for current setup and configuration details.

3. How SingleFile saves web pages

SingleFile captures a page’s resources into one HTML file. Its browser extension can save the current tab, selected content or a frame, several tabs, or pages automatically. A CLI is also available for command-line capture. The resulting file is convenient to move, store, and open without managing a directory of related assets.

SingleFile removes scripts by default because they can change the rendered page and often depend on online services. You can change options to keep scripts, but that does not guarantee that interactive controls will work offline. Some resources also depend on request headers such as Referer, and browser security rules can restrict access to particular domains or content.

SingleFile documents browser extensions and CLI usage in its project repository. Check its current release notes for supported browsers and installation instructions, since compatibility can change.

4. What happens to scripts, media, and interactivity?

SingleFile

  • Scripts are removed by default. This favors a static saved representation, but page features that rely on JavaScript may not work offline.
  • Keeping scripts is an option, not a promise of offline functionality. Scripts may still call remote services or expect the original site environment.
  • Some domains are protected from extension access by the browser. Security restrictions can also prevent capturing representations of canvas images or video elements.
  • Images or other resources may depend on request headers such as Referer. A resource that appears in the live page may not be available in the saved copy.

ArchiveBox

  • Different capture methods preserve different aspects of a page. A screenshot records appearance; HTML-oriented outputs preserve markup or resources; PDF is useful for a document-like copy; WARC is an archival container format.
  • Multiple formats give you alternatives, but they do not guarantee faithful replay of a dynamic site.
  • Results depend on site behavior, enabled extractors, and installed components. Inspect captures if completeness matters.

For either tool, test representative pages before relying on the output for research, evidence, or long-term preservation.

5. Which one should you choose?

  • Choose SingleFile if you want to save a page you are viewing as a portable standalone HTML file, or capture a small number of tabs without setting up an archive. This is a recommendation based on its extension and file-focused workflow.
  • Choose ArchiveBox if you want to collect many URLs, import or schedule sources, organize snapshots locally, or retain several representations such as screenshots, PDFs, and WARC alongside HTML. This is a recommendation based on its documented workflow and outputs.
  • Use both if you want ArchiveBox to manage the collection and also want a SingleFile representation among the outputs. ArchiveBox lists a SingleFile snapshot as one of its capture outputs.
  • For formal preservation, start with your institution’s required formats, validation, access, and retention process. SingleFile says it is not professional web-archiving software and points to WARC-based tools. ArchiveBox’s WARC output may be relevant, but its availability alone does not establish that an installation meets a preservation policy.

6. A practical workflow for choosing and checking captures

  1. Decide what you need to keep. If one portable page file is enough, try SingleFile. If you need an organized collection or multiple formats per URL, try ArchiveBox.
  2. Pick representative pages. Include a simple article, a page with lazy-loaded images, and any page with login, media, or interactive behavior that matters to your use case.
  3. Capture and inspect the result. Open the saved HTML or PDF, check that images and key text are present, and note which interactive features no longer work.
  4. For an ArchiveBox collection, confirm configuration. Check which extractors are enabled and whether their dependencies are installed. Inspect the snapshot folder rather than assuming every listed output was created.
  5. For preservation work, validate against requirements. Confirm the required file formats, metadata, integrity checks, and access process with the relevant archive policy.

7. Troubleshooting

Symptom Likely cause What to try
SingleFile cannot run on a page The browser blocks extensions on that protected domain. Check whether the domain is restricted by the browser. If it is, use an allowed capture route or a different page source where permitted.
An image is missing from a SingleFile save The resource may require a Referer header or be blocked by browser security rules. Compare the live resource with the saved file, and check the project FAQ for applicable browser restrictions and options.
A canvas image or video frame is absent Browser security restrictions can prevent access to those representations. Do not assume the extension can extract protected canvas or video content. Consider whether a screenshot or another capture method meets the need.
The saved page looks right but controls do not work Scripts are removed by default, or the page depends on remote services. Use the static copy for reading. If you enable scripts, treat the result as unverified until you check its offline behavior.
An expected ArchiveBox output is missing The relevant extractor may be disabled, a dependency may be missing, or that site capture may have failed. Review the installation and extractor configuration, then inspect the snapshot folder and logs for that capture.
ArchiveBox output exists but does not replay the live page The saved representation cannot reproduce all dynamic behavior or external services. Choose the output that matches your need: screenshot or PDF for appearance, HTML for page content, or WARC where an archival format is required. Verify the result.

8. Performance, reliability, and storage considerations

The dossier contains no controlled comparison of capture speed, fidelity, or success rates, so there is no sound basis here for saying one is faster or more reliable. Results depend on the site, page weight, browser or capture dependencies, network conditions, and configuration.

SingleFile’s one-file output is straightforward to move and organize, but the saved file can contain embedded resources. ArchiveBox can produce several files and representations for each URL, which offers flexibility and also means collection size and maintenance depend on enabled capture methods and what the pages contain. If storage matters, measure your own representative captures and decide which formats are necessary.

For reliability, keep the original URL and capture date with your records, inspect important snapshots, and maintain backups of the local archive. For any formal preservation workflow, validate the outputs and retention process against the applicable requirements.

9. ScreenshotNeo as an alternative for screenshot captures

ArchiveBox and SingleFile are useful for offline copies and archives. If your immediate task is to get a rendered screenshot from a URL through an API, try ScreenshotNeo first. It is a website screenshot API and MCP server; it is not a replacement for an organized archival collection or a self-contained HTML file.

Or skip the browser setup

One GET request returns an image or PDF. See the ScreenshotNeo API documentation for parameters and configuration.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners are accepted and removed before capture; newsletter popups and chat widgets are removed too. Each step can be turned off.
  • Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify page verdict and billing status in headers.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

10. Frequently asked questions

Can ArchiveBox save a page as one HTML file?

Yes. Its documented outputs include a SingleFile HTML snapshot, alongside other capture formats. Whether it is produced depends on the configured capture methods and successful capture.

Is SingleFile suitable for professional web archiving?

SingleFile’s own FAQ says it is not intended as professional web-archiving software and points to WARC-based tools. Match the tool and validation process to your preservation requirements.

Can I use the saved page as proof of what a website showed?

A saved file records a capture, but these tools’ documented features do not by themselves establish evidentiary completeness or authenticity. If that matters, define a capture and validation process that meets your requirements.

Do these tools work equally well on every website?

No universal success or fidelity guarantee is established by the project feature lists. Site behavior, browser restrictions, dependencies, and capture settings all affect the result; test the pages that matter to you.