How to Configure ArchiveBox to Save Pages as PDF
Enable ArchiveBox’s Chrome-based PDF capture, archive a page, find the output, and troubleshoot common configuration issues.
To have ArchiveBox save a rendered web page as a PDF, enable its PDF plugin and make sure Chrome or Chromium is available. In your ArchiveBox data directory, run:
archivebox config --set PDF_ENABLED=true
Then archive a page:
archivebox add 'https://example.com'
Look in that page’s snapshot output for output.pdf. The PDF plugin uses the shared Chrome session to print the rendered page. The plugin reference lists PDF capture as enabled by default, so setting it explicitly is useful when you want to confirm or persist the intended configuration.
1. Enable PDF capture
Choose the configuration method that matches how broadly the setting should apply:
| Method | Example | Scope |
|---|---|---|
| Persisted setting | archivebox config --set PDF_ENABLED=true |
Stores the setting in the ArchiveBox configuration. |
| Configuration file | Set PDF_ENABLED=true in ArchiveBox.conf |
Persists in the data directory’s configuration. |
| Environment variable | PDF_ENABLED=true archivebox add 'https://example.com' |
Convenient for one process or command. |
The documented aliases are SAVE_PDF and USE_PDF. Prefer the canonical PDF_ENABLED name for new configuration. Environment variables seed process-level defaults; persisted settings at more specific scopes can take precedence. In particular, changing an environment variable later does not silently replace configuration already stored for an existing Crawl.
2. Make sure Chrome is available
ArchiveBox’s page-to-PDF plugin requires Chrome or Chromium. The official Docker image or the project’s dependency installation can supply the browser dependencies. If ArchiveBox cannot locate the browser, check these settings:
CHROME_ENABLEDcontrols Chrome integration.CHROME_BINARYselects the browser executable when it is not found at the expected path.CHROME_RESOLUTIONsets the browser viewport and falls back to the sharedRESOLUTIONsetting.
The browser renders the page before printing it, so the result depends on what the page can load and display in that Chrome session.
3. Archive a page and locate the PDF
- Run the configuration command from the relevant ArchiveBox data directory, or use the environment-variable form for a one-off run.
- Add the page with
archivebox add 'https://example.com'. - Open the snapshot output for that URL and check for
output.pdf.
The PDF plugin prints a rendered HTML page. If the URL itself returns a PDF file, ArchiveBox handles that as a static-file download instead; that is a different capture path and does not depend on page printing being enabled.
4. Tune PDF generation when needed
| Setting | What it controls | Documented default or fallback |
|---|---|---|
PDF_TIMEOUT |
Time allowed for PDF generation. | 60 seconds. |
PDF_RESOLUTION |
Resolution used for PDF generation. | 1440,2000; falls back to RESOLUTION. |
PDFTOPPM_BINARY |
Poppler utility used to render the first-page card thumbnail. | Configure the executable path if needed. |
PDFTOPPM_BINARY is for the first-page thumbnail, not for creating the PDF. Chrome produces the PDF. Increase the PDF timeout only when pages need longer to render or the PDF operation is timing out; check general timeout and Chrome-specific settings when the failure points to those stages.
5. Troubleshoot missing or failed PDFs
| Symptom | Likely cause | What to check |
|---|---|---|
No output.pdf in the snapshot |
PDF capture is disabled or a more specific stored setting overrides the value you set. | Inspect the effective configuration and the Crawl or Snapshot settings for the run. Set PDF_ENABLED=true at the applicable scope. |
| Chrome launch or browser error | Chrome integration is disabled, or the browser executable is unavailable. | Check CHROME_ENABLED, install the required browser dependency, and set CHROME_BINARY if the executable is at a nonstandard path. |
| PDF generation times out | The page or browser takes longer than the configured allowance. | Review PDF_TIMEOUT (documented default: 60 seconds) and relevant general or Chrome-specific settings. |
| The URL downloads a PDF but page printing does not occur | The URL already serves a PDF document rather than an HTML page. | Check the static-file capture output. Direct PDF downloads use the staticfile plugin. |
| Thumbnail is missing but the PDF exists | The Poppler thumbnail renderer may be missing or misconfigured. | Check PDFTOPPM_BINARY. This affects the first-page card thumbnail, not PDF creation. |
Configuration sources and precedence can vary across ArchiveBox versions. When a change appears ignored, inspect the configuration for the installed version and the settings stored with the specific Crawl or Snapshot that produced the output.
6. Performance, reliability, and storage considerations
- Rendering time: PDF generation needs a browser render, so slow pages and pages that wait on external resources can take longer. Adjust
PDF_TIMEOUTwhen the observed failure is a timeout. - Browser availability: Chrome or Chromium must be installed and usable by the ArchiveBox process. A configured executable path must point to the browser available in that environment.
- Output verification: Check the snapshot for
output.pdfafter the archive operation. A successful URL add alone does not establish that the PDF plugin produced an output. - Cost: The supplied ArchiveBox documentation establishes configuration and dependencies, but provides no cost or performance benchmarks for this task. Plan for the storage used by generated PDFs and the browser resources needed to render pages.
7. Or skip the browser setup
If your goal is to capture a page as a PDF without managing ArchiveBox’s browser configuration, ScreenshotNeo provides a website screenshot API and MCP server. Its PDF endpoint accepts a URL in one GET request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -d format=pdf -o page.pdf
Cookie banners, popups, and chat widgets are removed before capture. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free.
Frequently asked questions
Is PDF capture enabled by default?
The PDF plugin reference lists PDF_ENABLED as true by default. Set it explicitly if you want the persisted configuration to state that choice.
Does PDF capture save the original HTML source as a PDF?
No. It prints the page rendered in Chrome. A URL that directly returns a PDF follows ArchiveBox’s static-file capture path.
Does a missing thumbnail mean the PDF failed?
Not necessarily. The PDF and its first-page card thumbnail have separate dependencies. Check for output.pdf first; PDFTOPPM_BINARY concerns the thumbnail.


