Best Webpage Archiving Tools That Save Both PDF and Screenshot
Compare tools for saving webpages as PDFs, screenshots, offline HTML, or replayable archives, and choose the right format for your needs.
If you need both a readable PDF and a visual screenshot of a webpage, ArchiveBox is the clearest documented match among the tools covered here: its documentation lists both PDF and PNG screenshot outputs. Choose a different tool if you need to replay an interactive browsing session or want a single offline HTML file instead.
The right archive depends on what you need to preserve. A PDF is convenient to read, print, or share; a screenshot records how a page looked; WARC or WACZ can preserve browsing data for replay; and a self-contained HTML file keeps a page and its resources together. These formats serve different purposes and are not interchangeable.
1. Which tool saves a webpage as both a PDF and a screenshot?
ArchiveBox is the documented fit for both PDF and screenshot output. Its documentation lists screenshot PNG, PDF, original HTML/CSS/JS, single-file HTML, and WARC among webpage outputs. It is a self-hosted application with CLI, REST API, webhooks, and web interface options. Output availability depends on the configured extractors and their dependencies, so check the setup that applies to your installation.
ArchiveBox is a better fit if you want to keep several representations of a page in an archive you operate or control. Its official site describes a self-hosted approach and notes that Docker bundles dependencies to simplify upgrades. That control comes with setup and maintenance work compared with saving a page from a browser extension.
The ArchiveBox Browser Extension is related, but it is not the same thing as the full ArchiveBox application: its documentation describes local full-page PNG screenshots and MHTML snapshots. The main ArchiveBox documentation is the source for its PDF output. Browser and capture-setting support can affect which extension formats are available.
2. Pick the archive format before picking the tool
| Format | What it preserves well | Useful when | Main limitation |
|---|---|---|---|
| A paginated, readable document | You want to read offline, print, or share a document | It does not preserve the page as an interactive website | |
| PNG screenshot | A visual rendering of the page at capture time | You need a visual record or want to compare appearances | Long pages can be awkward to read as one image; a screenshot is not a replayable page |
| WARC or WACZ | Captured web data intended for archival and replay workflows | You need to revisit a recorded browsing session or its states | It is a more specialized archive workflow than opening a PDF or image |
| Single-file HTML | A webpage and its resources packaged into one HTML file | You want a page copy you can open offline in a browser | It is neither a PDF nor a screenshot, and dynamic behavior may not be preserved |
| Original HTML/CSS/JS | Source resources gathered by the archive workflow | You need files for inspection or an archival collection | Saving source files alone does not guarantee the original page will run identically later |
For evidence or a visual comparison, keep the screenshot alongside the PDF and record the source URL and capture date in your own notes. A PDF is easier to read and print; a screenshot makes visual layout differences easier to spot. If you need to reproduce an interactive state, look at WARC/WACZ-oriented capture rather than expecting a static file to recreate it.
3. Compare the tools by the job you need done
| Tool | Best suited to | Documented workflow and outputs | Tradeoffs |
|---|---|---|---|
| ArchiveBox | People who want a self-hosted collection and multiple output types | Its documentation lists PNG screenshot, PDF, HTML, WARC, and other outputs. It offers CLI, REST API, webhooks, and a web interface. | Requires more setup and administration than a one-click saver. Output depends on configured extractors and dependencies. |
| ArchiveWeb.page | Recording an interactive browsing session, including state that matters to the capture | Browser extension or desktop app; exports WARC and WACZ; captured data stays local unless shared and can be replayed offline. | The reviewed official page presents interactive archiving and WARC/WACZ rather than PDF plus screenshot as its core output pair. |
| SingleFile | Making a quick self-contained offline page copy | Packages a page and its resources into one HTML file. Offers browser extensions and a CLI for Windows, macOS, and Linux; its official page lists Chrome, Edge, Firefox, and Safari. | A single HTML file is not a PDF or screenshot. Do not assume highly dynamic interactions will be preserved. |
| ArchiveBox Browser Extension | Sending pages to ArchiveBox and making local browser captures | Documentation describes local full-page PNG screenshots and MHTML snapshots. Its listed browser coverage includes Chromium browsers, Firefox-based browsers, Edge, and Safari/iOS Safari. | Extension output support depends on browser support and settings. The extension’s documentation does not mean every browser captures every format, and the PDF output is documented by the main ArchiveBox project. |
These are conditional recommendations based on documented formats and workflows, not a hands-on performance ranking.
4. Choose based on fidelity, privacy, and maintenance
For a PDF and screenshot in one archive
Start with ArchiveBox and check its current documentation for the extractors and dependencies required by your installation. Decide whether you will use its web interface, CLI, REST API, or webhooks to submit pages. If using the browser extension, check its format support separately from the main application’s outputs.
For a page that must be replayed
Choose ArchiveWeb.page when the interaction or browsing state is part of what you need to preserve. It records interactive browsing sessions and exports WARC/WACZ; its official page says captures are stored locally unless shared and can be replayed offline. It is not the direct choice when your required deliverables are specifically a PDF and a screenshot.
For a simple offline copy
Choose SingleFile when an HTML file that bundles page resources is enough. Its browser extension and CLI give you more than one capture workflow, but the result is still HTML rather than a PDF or screenshot.
For privacy and control
ArchiveBox is self-hosted, so you can keep the archive on infrastructure you control; you also take responsibility for operating and maintaining it. ArchiveWeb.page says captured data stays local unless you share it. Check where your capture files are stored, who can access them, and how you will back them up before archiving sensitive pages.
For less setup
A browser extension or single-file saver is generally a lighter workflow than maintaining a self-hosted archive. ArchiveBox offers several interfaces, but its setup may involve tools such as Chrome/Chromium and wget. Its documentation says Docker bundles dependencies to make upgrades easier.
5. Save your own PDF and screenshot
The exact commands and configuration depend on the tool and its current version. For ArchiveBox, use its official documentation to install it, choose an interface, and configure the extractors and dependencies that produce the formats you need. Then follow this capture checklist:
- Submit the page URL. Use the ArchiveBox interface or workflow configured for your installation.
- Enable the relevant outputs. Confirm the screenshot PNG and PDF extractors are available and enabled; add HTML or WARC if you need those formats too.
- Allow the page to finish loading. Pages that depend on scripts, delayed content, login state, or user interaction may not be represented by a static capture as expected.
- Inspect the saved files. Open the PDF and image and verify that the content, layout, and page length meet your purpose.
- Keep useful provenance. Store the original URL and capture date with the archive record if you will need to identify or explain the saved copy later.
For an interactive session, use ArchiveWeb.page’s browser extension or desktop app and export WARC/WACZ. For a single offline HTML copy, use SingleFile’s browser extension or CLI. Consult each project’s current official instructions for installation and command syntax; the reviewed documentation does not establish one shared command that applies to all three tools.
6. Screenshot webpages with an API
If your goal is to produce a screenshot in code, ScreenshotNeo is the alternative to try first. It is a website screenshot API and MCP server from Yorker Media. A GET request returns a PNG, JPEG, WebP, or PDF. This can automate visual capture, but a screenshot API is not a replacement for a replayable WARC/WACZ archive or a self-hosted collection of original page resources.
Put your API key in a secret store or environment variable in production. The examples below use the documented endpoint and parameters. See the ScreenshotNeo API documentation for options and response details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
These minimal calls capture a page image. ScreenshotNeo also supports PDF output, full-page capture with lazy images loaded, a selected element, dark mode, device presets and custom viewports, retina scale, PDF paper size/margins/orientation/page ranges, custom CSS and JavaScript, click-before-capture, hidden selectors, wait conditions, resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent background, resizing, configurable caching, signed links for public image tags, asynchronous jobs with signed webhooks, bulk requests of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names also work with those used by other screenshot APIs, which can make switching easier. Check the docs for exact parameter names and allowed values.
7. Or skip the browser setup
Use the one-call ScreenshotNeo API when you need an automated screenshot or PDF without setting up a browser capture pipeline. The API accepts a URL and returns the requested output; the example below saves a WebP image.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. All features are available on every plan.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
8. Troubleshooting and edge cases
| Symptom | Likely cause | What to check |
|---|---|---|
| ArchiveBox saves a screenshot but no PDF, or the reverse | The relevant extractor or dependency may not be configured, or the extension and main application workflows have been conflated. | Check the main ArchiveBox documentation for PDF and PNG outputs, then verify extractor configuration. Check the extension documentation separately for its local PNG and MHTML behavior. |
| A capture omits content that appears after a click or scroll | The content depends on interaction or delayed loading. | Use a workflow that captures the needed state. For an interactive browsing record, consider ArchiveWeb.page; for a screenshot API, configure supported wait and click options. |
| A screenshot or PDF is blank or incomplete | The page may not have loaded successfully, may need time or interaction, or may rely on unavailable resources. | Open the source page in a browser, check its login and consent state, and review the capture tool’s load/wait settings and dependencies. Inspect the result rather than assuming a successful request means a complete archive. |
| SingleFile output does not look like a PDF or screenshot | SingleFile creates a standalone HTML file. | Open the saved file in a browser, or choose a tool and output specifically documented for PDF or PNG. |
| A saved page cannot reproduce the original interaction | Static PDF, PNG, and single-file HTML captures do not promise full interactive replay. | Use a WARC/WACZ workflow such as ArchiveWeb.page when replay of a recorded session is the requirement. |
| ArchiveBox setup is difficult to maintain | A self-hosted archive requires administration and dependencies. | Review its current installation guidance and consider its Docker setup, which the official site describes as bundling dependencies for easier upgrades. |
| An API request fails or saves an unexpected file | The key, URL, response status, requested format, or output filename may not match. | Check the API documentation, validate the HTTP response before saving bytes, and use a filename extension matching the requested output format. |
9. Performance, reliability, and cost
ArchiveBox gives you control over where an archive is stored, but that means you need to plan for installation, storage, dependency updates, and backup. Its output depends on extractors and their dependencies; Docker can bundle dependencies according to the official site. No comparative capture-speed or fidelity benchmarks are established by the sources used here, so choose by documented output and workflow rather than assumed speed.
ArchiveWeb.page’s local capture and offline replay workflow is useful when privacy and session fidelity matter, but plan how you will store and share exported WARC/WACZ files. SingleFile produces one HTML file, which is straightforward to keep as a standalone artifact, but the format does not provide the PDF and screenshot pair.
ScreenshotNeo charges only for clean shots; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include X-Page-Verdict and X-Billed headers to report the result. The free plan is 1,000 shots per month without a card. Paid plans are Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. These are screenshot API plans, not prices for operating ArchiveBox or the other archival tools.
10. Frequently asked questions
Does a screenshot prove what a webpage said?
A screenshot records a visual rendering, but it does not by itself preserve the underlying page resources or browsing history. Keep the URL and capture date with the file if provenance matters, and choose an archival workflow that fits your record-keeping needs.
Can I use one tool for every kind of archive?
Not reliably. PDF, screenshot, offline HTML, and replayable web archives answer different needs. Pick the outputs you need first, then confirm the tool documents them.
Should I save both a PDF and a screenshot?
That is useful when you need both a readable document and a visual reference. For interactive replay or an offline site copy, add an archive format suited to that purpose rather than expecting those two files to cover it.
Which choice needs the least ongoing administration?
A browser extension or standalone HTML saver is typically lighter to operate than a self-hosted archive. ArchiveBox provides more control and interfaces, with corresponding setup and maintenance responsibilities.
Sources and documentation
- ArchiveBox documentation for supported outputs and interfaces.
- ArchiveBox product site for its self-hosted approach and dependency notes.
- ArchiveWeb.page for interactive capture, WARC/WACZ, local storage, offline replay, and browser options.
- SingleFile for single-file HTML, browser availability, and CLI details.
- ArchiveBox Browser Extension documentation for local screenshot and MHTML capture details.
Software details can change. Check each project’s official documentation for current installation steps and supported options.
