Best Tools and Commands for Converting HTML to PDF on Linux
Choose a Linux HTML-to-PDF renderer by page behavior, CSS needs and workflow. Get runnable commands and code for Chrome, Puppeteer, WeasyPrint and more.
For a one-off webpage, use headless Chrome or Chromium: chromium --headless --print-to-pdf=page.pdf --no-pdf-header-footer 'https://example.com/'. It prints through a browser engine, which is a good starting point when the page depends on JavaScript or modern browser behavior. For recurring application automation, use Puppeteer or Playwright. For authored HTML and CSS documents where JavaScript is not required, consider WeasyPrint. Prince is a commercial option for specialized publishing; treat wkhtmltopdf as an existing Qt WebKit-based option and verify its current fit before adopting it.
There is no universal best converter. Choose by whether the input is a live, JavaScript-driven page or a controlled document, how much paged-media styling you need, and whether you need a shell command or programmatic control. This guide covers runnable Linux examples, configuration, troubleshooting, and operational concerns.
1. Choose a renderer for your input
| Tool | Best starting point | JavaScript and browser behavior | Trade-offs to check |
|---|---|---|---|
| Chrome or Chromium headless CLI | One-off URL-to-PDF jobs | Uses a browser to render the page | Executable names vary; timing flags do not prove the application is ready |
| Puppeteer | Node.js workflows that need browser control | Automates Chromium; PDF output uses print media by default | Choose navigation and application readiness conditions deliberately |
| Playwright | Projects already using Playwright browser automation | Chromium automation with a Linux browser installation model | Provision compatible browser binaries, including headless shell when needed |
| WeasyPrint | Authored HTML/CSS documents with print styling | Do not assume full-browser JavaScript execution | Check supported CSS, Python and Pango requirements, and resource fetching |
| Prince | Specialized publishing and paged output requirements | Documents HTML/XML, CSS and JavaScript support | Commercial software; verify current licensing and evaluate with your own documents |
| wkhtmltopdf | An existing deployment that already uses it | Project describes a Qt WebKit-based renderer | Research here does not establish current maintenance or compatibility; check releases and security status |
Start with the simplest renderer that matches the page. If a URL is rendered blank because content appears only after client-side code runs, use browser automation and wait for the app’s real ready state. If a document is generated from controlled markup and should follow print CSS, a print-oriented renderer may be easier to manage. For untrusted markup, treat the renderer as a security boundary.
2. Convert a webpage with headless Chrome or Chromium
Chrome’s headless command line can save a URL directly as a PDF. The documented default filename is output.pdf; choose an explicit path to make scripts predictable. Chrome’s reference documents --print-to-pdf, --no-pdf-header-footer, --timeout, and --virtual-time-budget as controls for printing and capture timing. See Chrome’s headless command-line reference.
chromium --headless --print-to-pdf=page.pdf --no-pdf-header-footer 'https://example.com/'
Depending on the distribution and package, the executable may be named chrome, chromium, or chromium-browser. Use the installed package’s actual executable rather than assuming one alias.
Useful command options
--headless: run without opening a visible browser window.--print-to-pdf=page.pdf: write to a named PDF. Without a name, Chrome documentsoutput.pdfin the current working directory.--no-pdf-header-footer: suppress the printed URL, date, and page-number furniture when unwanted.--timeout=5000: wait up to five seconds before capture. This is a timing control, not proof that every network request or app state has settled.--virtual-time-budget=42000: advance virtual time for time-dependent page code. It does not guarantee external resources or application data are ready.
# Keep browser print headers and footers
chromium --headless --print-to-pdf=page.pdf 'https://example.com/'
# Give delayed content up to five seconds before capture
chromium --headless --timeout=5000 --print-to-pdf=page.pdf 'https://example.com/'
# Advance virtual time for pages with time-dependent code
chromium --headless --virtual-time-budget=42000 --print-to-pdf=page.pdf 'https://example.com/'
Confirm the output exists and inspect representative pages. Check that fonts, images, links, page breaks, background colors, and JavaScript-populated content appear as intended. For authenticated pages, a simple URL command may not provide the required session; use a controlled browser automation flow that handles authentication safely.
3. Use Puppeteer or Playwright in an application
When PDF creation is part of a Node.js service or automation job, browser control lets you set navigation behavior, wait for application-specific readiness, and choose print or screen media. Puppeteer’s page.pdf() uses print CSS by default and waits for fonts by default. Call page.emulateMediaType('screen') before generating the PDF if the page should use screen styles instead. Puppeteer PDF generation and the Page.pdf() API document these behaviors.
Puppeteer: runnable Node.js pattern
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com/', { waitUntil: 'networkidle2' });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
await browser.close();
}
This is an API pattern, not a guarantee that networkidle2 fits every page. Some apps keep connections open or load content after network activity has paused. Others need login, a particular selector, or a deliberate delay before capture. Wait for the application state that means the content is ready, and handle navigation failures and browser cleanup in production code.
Print CSS versus screen CSS
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-styled.pdf', printBackground: true });
Leave the media type at its default when you want the page’s print styles. Add print-specific CSS to control pagination and hide interactive elements:
@media print {
nav, .cookie-banner, .controls { display: none !important; }
h1, h2, h3 { break-after: avoid; }
}
@page {
size: A4 portrait;
margin: 15mm;
}
Playwright
Playwright is a reasonable choice when it already fits the application’s browser automation needs. Its Linux setup uses browser builds and a separate Chromium headless shell; install the browser artifacts expected by the project and deployment. Playwright documents installing only the headless shell for CI jobs that need no headed browser. See the official Playwright browser documentation. The rendering flow is similar: navigate, wait for a meaningful app-ready condition, then call the page PDF API in a supported Chromium setup.
4. Convert authored HTML with WeasyPrint
WeasyPrint provides a command line and Python API for HTML/CSS-to-PDF. It accepts a local filename or URL. It is a visual HTML/CSS renderer aimed at printing; do not assume it runs page JavaScript like a full browser. Its documentation notes that unsupported CSS properties can produce warnings, so validate your stylesheet against the target version. See WeasyPrint First Steps and its stable documentation.
weasyprint report.html report.pdf
# Render a remote URL
weasyprint 'https://example.com/report' report.pdf
# Apply a separate stylesheet
weasyprint -s print.css report.html report.pdf
For the stable documentation surfaced in this research, version 70.0 lists Python 3.10+ and Pango 1.44+ among its requirements. Recent Debian, Ubuntu, Fedora, Arch Linux, and Gentoo releases package WeasyPrint, but package names and versions depend on the distribution. Check the official instructions for your release. If you install through pip, use a virtual environment and satisfy system dependencies first rather than assuming pip alone provides them.
Set paper size and margins with CSS
WeasyPrint page geometry belongs in CSS @page, not a command-line paper-size flag:
@page {
size: A4 portrait;
margin: 15mm;
}
@media print {
.screen-only { display: none; }
h2 { break-after: avoid; }
}
Put the rules in the HTML stylesheet or pass them with -s print.css. Relative images, fonts, and stylesheets must resolve from the HTML document’s base location; check resource URLs when converting a local file.
Python API pattern
from weasyprint import HTML, CSS
HTML(filename='report.html').write_pdf(
'report.pdf',
stylesheets=[CSS(filename='print.css')],
)
5. When Prince or wkhtmltopdf fits
Prince for publishing workflows
Prince documents HTML/XML conversion with CSS and JavaScript support, command-line use, and advanced paged output capabilities. It is commercial software. If the document relies on specialized publishing features, evaluate it using representative inputs and verify current Linux licensing and product details before adoption. The available evidence does not establish current price or affiliate terms. Start with the Prince documentation.
wkhtmltopdf for an existing setup
The project’s site describes wkhtmltopdf as an open-source command-line renderer built on Qt WebKit. That description does not establish its current maintenance state or suitability for a modern application. For a new deployment, check recent releases, distribution packages, security maintenance, and output against the pages you actually need to render. See the official wkhtmltopdf site.
6. Resolve assets, fonts, authentication, and page timing
A renderer can only include content it can access and render. Before tuning page-size settings, check that the input and its dependencies are reachable:
- Choose a local HTML path or an accessible HTTP/HTTPS URL.
- Check relative image, font, and stylesheet URLs against the document’s base URL.
- Confirm that the rendering process can reach required hosts and that credentials are available where needed.
- For browser pages with client-side content, wait for a selector or app-specific ready state rather than assuming a fixed timeout is enough.
- Inspect the PDF for missing assets, unexpected page breaks, clipped content, print-color differences, and absent JavaScript-generated text.
Prince documents support for local and HTTP/HTTPS inputs and waiting for page resources; WeasyPrint accepts filenames, URLs, and file objects. Details differ by engine, so use the chosen renderer’s current documentation for its URL, font, and resource behavior.
7. Troubleshoot common conversion failures
| Symptom | Likely cause | What to do |
|---|---|---|
command not found |
The package is absent or its executable has another name. | Check the installed package and use its actual binary name, such as chrome, chromium, or chromium-browser. |
| PDF is blank or content is missing | The page renders content after initial navigation, needs authentication, or cannot load resources. | Check access and asset URLs. In browser automation, wait for the app’s actual ready state and confirm the session is available. |
| PDF lacks images, fonts, or CSS | Relative URLs resolve from an unexpected base, a host is unreachable, or a resource is blocked. | Use valid absolute URLs or the correct document base, and verify network access and font availability from the rendering environment. |
| Page is cut off or breaks badly | Screen layout is being printed without appropriate print rules, or page geometry is unsuitable. | Define @page size and margins, add print CSS, and inspect long tables, images, and headings across page boundaries. |
| Header, date, URL, or page numbers appear unexpectedly | Browser print furniture is enabled. | For Chrome CLI, add --no-pdf-header-footer. |
| Delayed content is absent | Capture occurred before the app populated the page. | Use Chrome timing flags as appropriate, or preferably wait for a page-specific selector/readiness signal in Puppeteer or Playwright. A timeout alone does not prove readiness. |
| WeasyPrint reports unsupported CSS or missing libraries | The CSS property is unsupported by that version, or system dependencies such as Pango are missing or incompatible. | Read the warning, check the versioned support and installation docs, and install the dependencies for the target distribution. |
Automation hangs or is unreliable with networkidle2 |
The page maintains network activity or readiness is unrelated to network quiet. | Choose a condition tied to the required content, such as an application selector, and set bounded timeouts and cleanup. |
8. Handle untrusted HTML safely
Do not render user-controlled HTML or CSS in a process with broad filesystem and network access. WeasyPrint’s security documentation warns that untrusted markup and styles can access local files, consume excessive resources or take a long time to render, and fetch network resources. This is a general rendering threat model: consult the selected engine’s current security guidance as well.
- Run conversions in an isolated process or container with restricted filesystem and network access.
- Limit CPU time and memory, and terminate jobs that exceed their budgets.
- Filter or restrict URL fetching where the renderer supports it, and prevent access to internal services and local secrets.
- Keep browser and renderer packages updated according to your deployment process.
- Separate trusted templates from user-supplied markup and test malicious URL and resource cases.
9. Performance, reliability, and cost planning
The research sources provide no comparable benchmark, so do not pick a renderer based on a universal speed claim. Measure representative documents in the target Linux environment. Browser engines bring browser installation and automation needs; WeasyPrint brings Python and native library dependencies. The cost of commercial publishing software such as Prince should be confirmed directly with the vendor because current pricing was not established here.
For reliability, use explicit output paths, bound navigation and render times, close browser processes in cleanup paths, and log the input, renderer version, exit status, and failure reason. Validate output PDFs on a representative sample whenever templates, fonts, browser versions, or CSS change. If batch rendering, control concurrency to fit available memory and isolate failures so one malformed document does not stall the entire queue.
10. A practical selection checklist
- Live URL and browser behavior? Begin with Chrome/Chromium CLI for a manual job; use Puppeteer or Playwright for repeated automation or custom waits.
- Controlled HTML/CSS document? Try WeasyPrint and express page geometry in
@page. - Specialized publishing? Evaluate Prince against real documents and current commercial terms.
- Existing wkhtmltopdf system? Check its current security, release, package, and rendering fit before extending the deployment.
- Untrusted input? Isolate the renderer, restrict access, and cap resource use.
- Before shipping? Inspect pages, typography, print colors, links, headers/footers, images, page breaks, and dynamic content.
11. Or skip the browser setup
If the goal is to capture a web page as an image or PDF through an API, ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API can also return a PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Replace the example target URL as needed. For a PDF response, set the documented output format option from the API docs. ScreenshotNeo removes cookie and consent banners from more than 60 known platforms, along with newsletter popups and chat widgets, before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
Frequently asked questions
How do I save a webpage as a PDF from the Linux command line?
Use Chrome or Chromium with --headless --print-to-pdf=page.pdf and the page URL. Add --no-pdf-header-footer if you do not want the browser’s print furniture.
Can I use headless Chrome to print a webpage?
Yes. Chrome documents a headless PDF command-line option. For pages whose content loads asynchronously, account for timing or use browser automation with an application-specific readiness condition.
Which option should I use for JavaScript-heavy pages?
Start with a browser-based renderer. Use Puppeteer or Playwright when you need to control navigation, wait for page state, or integrate the conversion into an application. Verify the output for the actual page.
Can WeasyPrint convert a JavaScript application?
WeasyPrint is an HTML/CSS print renderer; do not assume it executes page JavaScript as a full browser would. Use a browser engine when the required content is created by client-side JavaScript.
Where do I set A4 size and margins in WeasyPrint?
Use CSS @page rules in the document or a stylesheet supplied with -s.
Is wkhtmltopdf maintained?
The research used here establishes that the project describes it as a Qt WebKit-based command-line renderer, but does not establish its current maintenance lifecycle. Check current release and security information before choosing it for a new deployment.


