How to Enable ArchiveBox’s Single-File HTML Captures
Enable ArchiveBox’s SingleFile extractor for one URL or ongoing runs, find the saved singlefile.html, and troubleshoot configuration issues.
To make a one-off ArchiveBox capture with SingleFile, run archivebox add --plugins=singlefile 'https://example.com'. ArchiveBox uses headless Chrome and SingleFile to produce a standalone singlefile.html in the snapshot folder. The chrome plugin is pulled in automatically as a requirement.
1. Run SingleFile for one URL
Run this from the ArchiveBox installation or collection where you normally add URLs:
archivebox add --plugins=singlefile 'https://example.com'
Replace the example URL with the page you want to archive. The --plugins value selects extractors for this add run. To run another extractor alongside SingleFile, use a comma-separated list:
archivebox add --plugins=singlefile,favicon 'https://example.com'
The exact combination depends on your archiving workflow. Choosing singlefile automatically includes its required chrome plugin.
2. Find and inspect the saved HTML
ArchiveBox writes the SingleFile artifact as singlefile.html under the snapshot’s output folder. Snapshot outputs are ordinary files in that folder. Open the file in a browser or inspect it with your usual file tools. A standalone HTML snapshot is useful when you want one saved document rather than relying on the original site being available.
If the file is missing, first confirm that SingleFile was selected for the run and that the add operation completed. Then check the relevant ArchiveBox output and configuration as described below. Installation dependencies may also matter: SingleFile is an optional extractor dependency, and archivebox install resolves extractor dependencies.
3. Enable SingleFile for ongoing runs
The command-line flag applies to the current add run. To use SingleFile repeatedly, set PLUGINS=singlefile at the configuration scope that applies to the Crawl, Snapshot, or Persona, or choose the extractor through the admin Add form or API.
ArchiveBox documents several configuration routes:
- Use
archivebox config. - Set the value in the collection’s
ArchiveBox.conf. - Provide it through environment variables.
These configuration methods also work with Docker. Check for existing scope-specific settings before changing process defaults: more specific persisted configuration can take precedence, and changing an environment variable does not silently rewrite an existing Crawl configuration. Set the value at the scope that owns the runs you want to affect.
Choose the right method
| Method | When to use it | Scope |
|---|---|---|
archivebox add --plugins=singlefile URL |
A single add operation or an explicit one-off choice | That add run |
PLUGINS=singlefile in persisted configuration |
Repeated captures governed by that configuration | The relevant Crawl, Snapshot, or Persona, subject to configuration precedence |
| Admin Add form or API extractor selection | Adding through ArchiveBox’s interface or API | The submitted add operation and its applicable stored settings |
4. Check installation and configuration
- Confirm you are running the
archivebox addcommand in the intended ArchiveBox collection. - For a one-off capture, include
--plugins=singlefileon that add command. - For ongoing captures, inspect the configuration at the Crawl, Snapshot, or Persona scope that applies to them.
- If the extractor dependency is unavailable, run the documented dependency setup command
archivebox installfor your installation. - After the run, inspect the snapshot output folder for
singlefile.html.
ArchiveBox configuration behavior and plugin selection can evolve. For exact syntax and supported settings for your installed release, consult the ArchiveBox Configuration wiki and the ArchiveBox repository.
5. Keep browser-extension capture separate
ArchiveBox’s browser extension has a separate local-capture route. Its README says that saving local HTML copies requires the optional SingleFile extension to be installed and connected. That is distinct from selecting ArchiveBox’s server-side extractor with --plugins=singlefile; installing the browser extension does not replace configuring the server-side run.
See the ArchiveBox browser extension README for that extension workflow.
6. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
No singlefile.html appears |
SingleFile was not selected for that run, the run did not finish, or output is being checked in the wrong snapshot folder. | Check the command or Add form/API selection, confirm the add completed, and inspect the snapshot’s output folder. |
| A one-off run works, but later runs do not save SingleFile output | The command-line plugin selection only applies to that add run. | Set PLUGINS=singlefile at the relevant persisted scope or select the extractor for each interface/API add operation. |
| Changing an environment variable has no effect on an existing Crawl | Stored scope-specific configuration may override process defaults. Environment changes do not silently update existing Crawl configuration. | Inspect and update the configuration for the Crawl, Snapshot, or Persona that governs those runs. |
| ArchiveBox cannot run the extractor because a dependency is missing | SingleFile is an optional extractor dependency and may not have been installed. | Use archivebox install to resolve extractor dependencies, then run the add operation again. |
| The browser extension does not save a local HTML copy | The optional SingleFile browser extension may not be installed or connected. | Install and connect that extension as its README describes. This is separate from the server-side ArchiveBox plugin. |
| Other extractors stopped running after enabling SingleFile | The configured plugin list may select only SingleFile. | Specify the desired extractors as a comma-separated list, such as singlefile,favicon. |
7. Performance, reliability, and storage considerations
SingleFile renders pages with headless Chrome, so capture depends on the browser-based extraction path and the page being available to load. Page complexity and the resources it uses can affect how long a capture takes and how large its saved HTML becomes; the documentation reviewed here does not specify a universal duration or size.
For repeated use, persisted configuration makes extractor selection part of the applicable ArchiveBox scope. Check that scope and its precedence when a run behaves differently from a one-off command. The resulting artifact is stored as an ordinary file in the snapshot folder; plan storage according to the pages you retain. No fixed storage or performance figure is implied here.
8. Or skip the browser setup
If your goal is to get a screenshot or PDF of a page rather than preserve a SingleFile HTML archive, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request can return an image or PDF. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Try it with a free ScreenshotNeo account.
9. FAQ
Does SingleFile create a PDF?
No. The documented artifact for this extractor is singlefile.html, an HTML snapshot.
Does the one-off command also change future Crawl settings?
No. The --plugins flag selects plugins for the current add run. Configure the relevant persisted scope for ongoing runs.
Is the browser extension required for server-side SingleFile extraction?
The browser extension’s optional SingleFile extension is for its separate local-capture route. The ArchiveBox server-side extractor is selected through ArchiveBox’s plugin configuration.
Can I run SingleFile with other extractors?
Yes. Pass a comma-separated plugin list, for example --plugins=singlefile,favicon.


