wkhtmltopdf vs Chrome Print to PDF for Website Archiving
Compare wkhtmltopdf and Chrome Headless for saving webpages as PDFs, and learn when a PDF is not enough to archive a website.
Short answer: For a new workflow that saves individual webpages as PDFs, evaluate Chrome Headless first: its current developer documentation describes direct PDF printing. Keep wkhtmltopdf when a legacy workflow depends on its output and it still meets your requirements. Neither is a universal fidelity winner; render representative pages and inspect the PDFs.
First define “website archiving.” A PDF is a page-oriented reading copy. It does not, by itself, preserve a website’s linked pages and supporting resources for replay. The Library of Congress identifies WARC as its preferred web-archive format. If you need a replayable capture, use a web-archiving workflow and treat PDF as an additional access copy.
What each tool produces
wkhtmltopdf is a command-line HTML-to-PDF tool based on Qt WebKit. Its official downloads page lists 0.12.6 as the stable series, released June 11, 2020, and its GitHub repository was archived on January 2, 2023. It remains a possible fit for a validated legacy system, but pin the exact build and environment.
Chrome Headless prints a page through Chrome’s browser engine. The official command-line reference documents --headless --print-to-pdf, --no-pdf-header-footer, and --timeout. Those controls do not guarantee that every dynamic page has finished rendering when the PDF is produced.
Comparison: wkhtmltopdf versus Chrome Headless
| Decision | wkhtmltopdf | Chrome Headless | What to verify |
|---|---|---|---|
| Engine and maintenance | Qt WebKit; stable series 0.12.6 dates to 2020; repository archived in 2023. | Current Chrome browser engine in headless mode. | Required CSS, fonts, JavaScript, and output on real pages. |
| PDF controls | Options cover paper settings, headers and footers, document objects, and a --window-status wait condition. |
CLI documents PDF output, header/footer suppression, and a capture timeout. | Paper size, margins, scale, backgrounds, headers, footers, and page breaks. |
| Dynamic content | Wait behavior depends on the build and page. | A timeout sets a limit; it does not prove client rendering or network activity is complete. | Whether the page’s content and resources are ready before printing. |
| Security | The project warns against processing untrusted HTML and recommends sanitizing supplied HTML and JavaScript. | The cited CLI documentation does not establish a security comparison. | Use suitable process isolation and only fetch content your system is authorized to access. |
| Archiving purpose | Produces a PDF document. | Produces a PDF document. | For linked resources and replay, use a web archive such as WARC. |
There is no controlled head-to-head evidence here establishing universal fidelity, speed, or resource-use rankings. Results depend on the page, print styles, content timing, tool build, options, and environment.
Print a webpage to PDF with Chrome Headless
Install Chrome or Chromium in your environment, then run the documented command from a directory where the process can write output.pdf:
google-chrome --headless --print-to-pdf https://example.com
On systems where the executable is named chromium or chromium-browser, use that installed command name. To suppress Chrome’s generated header and footer, add the documented flag:
google-chrome --headless --print-to-pdf --no-pdf-header-footer https://example.com
To bound the capture time, set a timeout in milliseconds. This is a maximum wait, not a page-readiness condition:
google-chrome --headless --print-to-pdf --timeout=10000 https://example.com
Check the exit status and confirm the output exists and opens. For dynamic applications, a command-line timeout may not be sufficient to wait for a particular client-rendered state; establish a readiness strategy in your browser automation workflow and test it against your actual target pages.
Generate a PDF with wkhtmltopdf
With wkhtmltopdf installed, a basic URL-to-PDF command is:
wkhtmltopdf https://example.com page.pdf
Its options include paper and margin settings, header/footer configuration, document objects, and a --window-status wait option. The exact supported flags and behavior can differ across builds, including builds using patched or distribution Qt. Consult the manual for the binary you deploy and record its version and options alongside your output.
Security: the wkhtmltopdf project explicitly warns not to use it with untrusted HTML and advises sanitizing user-supplied HTML and JavaScript. Treat remote pages and generated input as untrusted unless your application has established otherwise; apply appropriate isolation and access controls.
Make a fair comparison on your pages
- Choose representative URLs: static pages, long articles, pages with print styles, and pages with client-rendered content.
- Fix the same viewport assumptions, paper size, margins, and header/footer expectations where the tools allow it.
- Capture each page with a pinned tool build and save the command, version, and relevant environment details.
- Compare text completeness, layout, fonts, images, page breaks, background colors, and whether expected dynamic content appears.
- Repeat after changing browser versions, source content, or print settings; inspect the resulting PDFs rather than assuming output stability.
This makes the choice specific to your workload. Chrome is a sensible first candidate for a new PDF workflow because its official documentation covers current headless printing. wkhtmltopdf can remain appropriate when a legacy system’s established output is required and validated.
PDF snapshot or website archive?
Ask what must survive. If the deliverable is a readable rendering of one page, PDF printing may fit. If it must include multiple pages, linked resources, and metadata for later replay, a PDF from either renderer is not a substitute for a crawler/archive workflow. The Library of Congress names WARC as its preferred web-archive format and discusses preserving supporting assets such as CSS and JavaScript.
PDF/A is a family of ISO standards for constrained PDF forms intended for long-term preservation of page-oriented documents. Choosing PDF/A addresses considerations for preserving a PDF document; it does not turn that document into a capture of a whole website.
Operational notes: performance, reliability, and cost
- Performance: the available sources do not establish comparative speed or memory use. Measure your own representative workload, including startup, rendering, and file handling.
- Reliability: pin wkhtmltopdf builds and Chrome versions, retain capture settings, and inspect representative output after upgrades. Chrome’s timeout bounds waiting but does not ensure all page content is ready.
- Reproducibility: page content, network resources, fonts, browser versions, print styles, and timing can affect output. Keep fixtures and compare generated PDFs when these inputs change.
- Cost: account for installation, version maintenance, isolation, storage, and any operational support. The sources reviewed provide no comparable pricing or benchmark figures for these tools.
Or skip the browser setup
If you need a screenshot or a PDF from a URL without managing a local browser, ScreenshotNeo provides a website screenshot API and MCP server. For API options, see the ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. This is a page capture option, not a substitute for a WARC workflow when replayable website preservation is the requirement.
Create a free ScreenshotNeo account: 1,000 screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| No PDF or output file | The command failed, executable name differs, or the process cannot write to the current directory. | Check the executable and exit status; use a writable output path and confirm the file exists. |
| Content is missing from a Chrome PDF | Client-side rendering or network resources were not ready before printing. | Use a readiness strategy appropriate to the page; test timing and resources. A maximum timeout alone does not confirm readiness. |
| Different layout between runs or hosts | Browser/build, fonts, environment, page content, or print settings differ. | Pin versions and settings, record the environment, and compare against saved fixtures. |
| Unexpected headers, footers, or page breaks | Print defaults or page print styles affect pagination. | Set the desired print options, including Chrome’s header/footer suppression where appropriate, and inspect page breaks and margins. |
| wkhtmltopdf output differs across installations | Builds can vary, including patched versus distribution Qt builds. | Check the exact build and manual, pin the binary, and validate it on representative pages. |
| Security concern with supplied markup | Untrusted HTML or JavaScript can expose the host to risk; wkhtmltopdf documents this warning explicitly. | Sanitize supplied content and run capture processes with suitable isolation and access controls. |
| A PDF cannot recreate the archived site | A PDF records a page-oriented rendering rather than the site’s harvested resources. | Use a web crawler/archive workflow that records resources and metadata, such as WARC, and retain PDF as an access copy if useful. |
FAQ
Which is better for archiving a website?
For a single-page PDF, evaluate Chrome Headless for a new workflow and retain wkhtmltopdf when a validated legacy dependency requires it. For a replayable website archive, choose a WARC-capable archiving workflow.
Does Chrome print to PDF save a whole website?
No. It prints a page to a PDF; it does not crawl and preserve the website’s linked pages and resources as a replayable archive.
Should I migrate away from wkhtmltopdf?
Assess whether its current output remains necessary, validate the replacement against your pages, and account for the archived repository and dated stable release when planning maintenance.
Is PDF/A the same as WARC?
No. PDF/A concerns preservation of page-oriented PDF documents. WARC is a web-archive format for captured resources and related records.
Sources
- wkhtmltopdf downloads and security guidance and its archived GitHub repository.
- Chrome Headless command-line documentation.
- Library of Congress guidance on WARC, PDF/A, and web archiving recommendations.
