How to Make a PDF Archive of a Webpage in Linux from the Command Line
Save a live webpage as a PDF from the Linux command line with headless Chrome, and learn when to use wkhtmltopdf, WeasyPrint, or a screenshot API.
To save a live webpage as a PDF from a Linux command line, use headless Chrome or Chromium:
google-chrome --headless --print-to-pdf=page.pdf 'https://example.com'
The Chrome executable may be named differently on your distribution. This captures the browser-rendered page, including content produced by JavaScript that runs before printing. To remove the printed URL, date, and page numbers, add --no-pdf-header-footer. See the Chrome Headless command-line reference for the current flags. For an easier hosted capture that also returns a PDF, see ScreenshotNeo.
1. Make a PDF with headless Chrome
Basic command
google-chrome --headless --print-to-pdf=page.pdf 'https://example.com'
Run it from the directory where you want the output. If your build does not accept the =page.pdf form, use the documented basic form, which writes output.pdf in the current working directory:
google-chrome --headless --print-to-pdf 'https://example.com'
Chromium builds commonly use a chromium executable name. Check what is installed and its supported flags if the command is rejected:
command -v google-chrome || command -v chromium || command -v chromium-browser
Replace the executable name in the examples with the path returned. The precise package name and executable vary across Linux distributions and installations.
Remove print headers and footers
google-chrome --headless --no-pdf-header-footer --print-to-pdf=page.pdf 'https://example.com'
The current Chrome flag is --no-pdf-header-footer. Older versions may use --print-to-pdf-no-header; consult the installed browser’s help or the official reference when the current flag is unknown.
Wait for delayed content
Some pages need a little time for scripts or delayed content. Chrome provides a maximum wait timeout and a virtual time budget:
google-chrome --headless --timeout=5000 --print-to-pdf=page.pdf 'https://example.com'
--timeout=5000 waits up to five seconds before capture, even if loading continues. For some time-dependent pages, a virtual-time budget can advance page timers while Chrome captures:
google-chrome --headless --virtual-time-budget=42000 --print-to-pdf=page.pdf 'https://example.com'
These options influence capture timing; neither proves that every image, API request, or dynamically rendered section has finished. Increase or adjust the wait for the page you need, then inspect the resulting PDF.
Keep a useful archive record
A PDF records a rendered document. It does not by itself preserve the original HTML, linked assets, interactive behavior, or network dependencies. When the capture may be used as a record, save the source URL and capture date alongside the PDF:
url='https://example.com/article'
date=$(date -u +%Y-%m-%dT%H:%M:%SZ)
printf 'URL: %s\nCaptured (UTC): %s\n' "$url" "$date" > page.txt
google-chrome --headless --no-pdf-header-footer --print-to-pdf=page.pdf "$url"
Keep page.txt with page.pdf. A screenshot or PDF is not a substitute for preserving source files when you need a complete web archive.
2. Alternatives: wkhtmltopdf and WeasyPrint
wkhtmltopdf
wkhtmltopdf is a command-line PDF converter using Qt WebKit:
wkhtmltopdf 'https://example.com' page.pdf
The project’s downloads page lists version 0.12.6 as its stable series, released in 2020. Check the rendered output on modern, script-heavy pages instead of assuming it will match a current browser. The project also warns against running it on untrusted HTML: unsanitized user-supplied HTML or JavaScript can compromise the server running it. See the project site and its downloads and security warning.
WeasyPrint
WeasyPrint is another option when you want control over HTML and CSS rendering:
weasyprint 'https://weasyprint.org' page.pdf
Its documentation notes that unsupported CSS properties can produce warnings. It also supports PDF/A variants, which are constrained formats with requirements and limitations, including restrictions around audio, video, JavaScript, color spaces, and fonts. Review the first steps and common use cases and PDF/A notes before choosing it for a preservation workflow.
| Method | Useful when | Considerations |
|---|---|---|
| Headless Chrome | You want a live page rendered by a browser with JavaScript support. | Executable names and flags vary; capture timing does not guarantee all content is ready. |
| wkhtmltopdf | You already rely on its Qt WebKit conversion workflow. | Check modern-page fidelity and heed the project’s warning about untrusted HTML. |
| WeasyPrint | You want HTML/CSS-to-PDF control or a documented PDF/A workflow. | Unsupported CSS can warn or render differently; PDF/A has format constraints. |
| ScreenshotNeo | You want a hosted screenshot or PDF API without managing a browser installation. | Requires an API key and network access; see the hosted capture section below. |
3. Choose a capture method for the page you have
For most live pages, start with Chrome because the command uses a browser rendering path and runs page JavaScript. Choose based on the output you need and the page’s behavior:
- Mostly static page: run the basic Chrome command and inspect the PDF.
- Delayed or script-rendered content: try a suitable timeout or virtual-time budget, then check that the expected section appears.
- HTML/CSS document you control: consider WeasyPrint and review any unsupported-CSS warnings.
- Existing Qt WebKit pipeline: wkhtmltopdf may fit, but verify output on the target page and isolate untrusted input.
- Repeatable hosted captures or API integration: consider a screenshot service and record its response status and billing information.
Compare browser fidelity, JavaScript and delayed-content behavior, print layout, needed PDF controls, maintenance state, and whether the workflow processes untrusted pages. No one renderer is best for every webpage.
4. Or skip the browser setup
ScreenshotNeo is a hosted website screenshot API that can return a PDF with one GET request. Its API and other options are documented at ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf
For PDF output, request the PDF format using the documented API options. The following examples use the required base request; consult the documentation for PDF parameters and other capture settings.
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.pdf", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.pdf', bytes));
Configure the request for PDF output as described in the API documentation. ScreenshotNeo accepts cookie consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.
5. Troubleshooting
| Symptom | Likely cause | What to try |
|---|---|---|
google-chrome: command not found |
Your distribution or installation uses another executable name, or Chrome is not installed. | Check command -v chromium and command -v chromium-browser; install or invoke the browser available in your environment. |
| Unknown or rejected PDF flag | The installed browser build uses a different flag spelling or version. | Check its command-line help and the Chrome Headless reference. Older builds may use --print-to-pdf-no-header for header removal. |
| PDF contains a loading screen or misses a section | The page had not rendered its delayed content at capture time. | Try --timeout=5000 or a suitable --virtual-time-budget; inspect the result because waiting does not guarantee all requests finish. |
| Unexpected page breaks or clipped layout | The page’s print styles or renderer differ from its screen layout. | Inspect the output and try a current browser renderer. For a controlled HTML/CSS source, compare with WeasyPrint and review its warnings. |
| PDF output looks unlike the live page | A different rendering engine, unsupported CSS, or script behavior changed the layout. | Use Chrome as the browser-based starting point; verify the page with the chosen renderer rather than assuming the engines are interchangeable. |
| wkhtmltopdf processes input you did not trust | Untrusted HTML or JavaScript can expose the host to compromise, according to the project warning. | Do not pass unsanitized user-supplied HTML/JS to wkhtmltopdf; use a controlled, isolated workflow. |
| API output is not the expected PDF | The request may be returning an image or an error response because output options, credentials, or target URL are wrong. | Check the API documentation, verify the key and URL, configure PDF output, and inspect response status and headers before saving the body as a PDF. |
6. Performance, reliability, and cost
A local browser capture avoids a per-request screenshot API charge but uses local CPU, memory, browser installation, and maintenance. For repeated captures, account for browser startup, page load time, concurrency, and the need to inspect failures. A longer wait can improve some delayed-page captures while increasing runtime; it still cannot confirm every dependency completed.
Reliability depends on the page as well as the renderer: network availability, redirects, consent screens, bot checks, script errors, and page changes can all affect the resulting document. Save the URL and capture timestamp when provenance matters, and inspect captures used for records. A PDF is a snapshot of rendered output, not a preserved copy of all source material.
Hosted APIs trade local browser administration for a request and service plan. ScreenshotNeo’s stated plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free. Every feature is on every plan. Only clean shots are billed; the response identifies the page verdict and billing status. Confirm the service’s current documentation for the request options needed for PDF output.
7. FAQ
Does a PDF preserve the original webpage?
No. It preserves a rendered document, not the original source files, interactions, or network dependencies. Keep source material separately if you need a complete archive.
Can Chrome guarantee that lazy-loaded content is included?
No. Timeout and virtual-time options affect when Chrome captures but do not guarantee that every page section or resource has finished loading. Check the resulting PDF.
Which renderer should I use for PDF/A?
WeasyPrint documents PDF/A variants. Review its format constraints and documentation for the specific variant and source document before relying on it.
Can I automate captures for many URLs?
Yes. A shell script can invoke a local browser repeatedly, while ScreenshotNeo provides bulk capture for up to 100 URLs per call. For either workflow, handle individual failures and keep a record of the URLs and capture outcomes.


