ScreenshotNeo

BlogHow-to

How to Configure ArchiveBox to Save Pages as PDF

Enable ArchiveBox’s Chrome-based PDF capture, archive a page, find the output, and troubleshoot common configuration issues.

By the ScreenshotNeo team4 October 20264 min read

To have ArchiveBox save a rendered web page as a PDF, enable its PDF plugin and make sure Chrome or Chromium is available. In your ArchiveBox data directory, run:

archivebox config --set PDF_ENABLED=true

Then archive a page:

archivebox add 'https://example.com'

Look in that page’s snapshot output for output.pdf. The PDF plugin uses the shared Chrome session to print the rendered page. The plugin reference lists PDF capture as enabled by default, so setting it explicitly is useful when you want to confirm or persist the intended configuration.

1. Enable PDF capture

Choose the configuration method that matches how broadly the setting should apply:

Method Example Scope
Persisted setting archivebox config --set PDF_ENABLED=true Stores the setting in the ArchiveBox configuration.
Configuration file Set PDF_ENABLED=true in ArchiveBox.conf Persists in the data directory’s configuration.
Environment variable PDF_ENABLED=true archivebox add 'https://example.com' Convenient for one process or command.

The documented aliases are SAVE_PDF and USE_PDF. Prefer the canonical PDF_ENABLED name for new configuration. Environment variables seed process-level defaults; persisted settings at more specific scopes can take precedence. In particular, changing an environment variable later does not silently replace configuration already stored for an existing Crawl.

2. Make sure Chrome is available

ArchiveBox’s page-to-PDF plugin requires Chrome or Chromium. The official Docker image or the project’s dependency installation can supply the browser dependencies. If ArchiveBox cannot locate the browser, check these settings:

  • CHROME_ENABLED controls Chrome integration.
  • CHROME_BINARY selects the browser executable when it is not found at the expected path.
  • CHROME_RESOLUTION sets the browser viewport and falls back to the shared RESOLUTION setting.

The browser renders the page before printing it, so the result depends on what the page can load and display in that Chrome session.

3. Archive a page and locate the PDF

  1. Run the configuration command from the relevant ArchiveBox data directory, or use the environment-variable form for a one-off run.
  2. Add the page with archivebox add 'https://example.com'.
  3. Open the snapshot output for that URL and check for output.pdf.

The PDF plugin prints a rendered HTML page. If the URL itself returns a PDF file, ArchiveBox handles that as a static-file download instead; that is a different capture path and does not depend on page printing being enabled.

4. Tune PDF generation when needed

Setting What it controls Documented default or fallback
PDF_TIMEOUT Time allowed for PDF generation. 60 seconds.
PDF_RESOLUTION Resolution used for PDF generation. 1440,2000; falls back to RESOLUTION.
PDFTOPPM_BINARY Poppler utility used to render the first-page card thumbnail. Configure the executable path if needed.

PDFTOPPM_BINARY is for the first-page thumbnail, not for creating the PDF. Chrome produces the PDF. Increase the PDF timeout only when pages need longer to render or the PDF operation is timing out; check general timeout and Chrome-specific settings when the failure points to those stages.

5. Troubleshoot missing or failed PDFs

Symptom Likely cause What to check
No output.pdf in the snapshot PDF capture is disabled or a more specific stored setting overrides the value you set. Inspect the effective configuration and the Crawl or Snapshot settings for the run. Set PDF_ENABLED=true at the applicable scope.
Chrome launch or browser error Chrome integration is disabled, or the browser executable is unavailable. Check CHROME_ENABLED, install the required browser dependency, and set CHROME_BINARY if the executable is at a nonstandard path.
PDF generation times out The page or browser takes longer than the configured allowance. Review PDF_TIMEOUT (documented default: 60 seconds) and relevant general or Chrome-specific settings.
The URL downloads a PDF but page printing does not occur The URL already serves a PDF document rather than an HTML page. Check the static-file capture output. Direct PDF downloads use the staticfile plugin.
Thumbnail is missing but the PDF exists The Poppler thumbnail renderer may be missing or misconfigured. Check PDFTOPPM_BINARY. This affects the first-page card thumbnail, not PDF creation.

Configuration sources and precedence can vary across ArchiveBox versions. When a change appears ignored, inspect the configuration for the installed version and the settings stored with the specific Crawl or Snapshot that produced the output.

6. Performance, reliability, and storage considerations

  • Rendering time: PDF generation needs a browser render, so slow pages and pages that wait on external resources can take longer. Adjust PDF_TIMEOUT when the observed failure is a timeout.
  • Browser availability: Chrome or Chromium must be installed and usable by the ArchiveBox process. A configured executable path must point to the browser available in that environment.
  • Output verification: Check the snapshot for output.pdf after the archive operation. A successful URL add alone does not establish that the PDF plugin produced an output.
  • Cost: The supplied ArchiveBox documentation establishes configuration and dependencies, but provides no cost or performance benchmarks for this task. Plan for the storage used by generated PDFs and the browser resources needed to render pages.

7. Or skip the browser setup

If your goal is to capture a page as a PDF without managing ArchiveBox’s browser configuration, ScreenshotNeo provides a website screenshot API and MCP server. Its PDF endpoint accepts a URL in one GET request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -d format=pdf -o page.pdf

Cookie banners, popups, and chat widgets are removed before capture. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free.

Frequently asked questions

Is PDF capture enabled by default?

The PDF plugin reference lists PDF_ENABLED as true by default. Set it explicitly if you want the persisted configuration to state that choice.

Does PDF capture save the original HTML source as a PDF?

No. It prints the page rendered in Chrome. A URL that directly returns a PDF follows ArchiveBox’s static-file capture path.

Does a missing thumbnail mean the PDF failed?

Not necessarily. The PDF and its first-page card thumbnail have separate dependencies. Check for output.pdf first; PDFTOPPM_BINARY concerns the thumbnail.

Sources