How to Convert a Web Page to PDF in Rust
Convert modern, JavaScript-heavy web pages to PDF in Rust with headless Chrome, html2pdf, wkhtmltopdf, and a hosted API.

For modern web pages, the most reliable Rust workflow is to drive headless Chromium and call its print-to-PDF operation after the page reaches a known ready state. Chromium executes JavaScript, loads current CSS, waits for web fonts and images, and supports print settings such as paper size, margins, scale, backgrounds, headers, footers, and page ranges.
For a quick local conversion, run:
chrome --headless --no-sandbox --print-to-pdf=output.pdf https://example.com/
The --print-to-pdf flag writes the rendered page to the named PDF file. Add --no-pdf-header-footer when you do not want Chrome’s generated date, URL, and page-number decorations. Chrome documents this command in its headless command reference.
1. Choose a rendering engine
| Option | Best for | Trade-offs |
|---|---|---|
| Headless Chrome or Chromium | Arbitrary live URLs, JavaScript applications, modern CSS | Requires a browser binary and more memory than an HTML-only renderer |
html2pdf |
A Rust-friendly CLI with waits and print settings | Still needs Chrome; remote URL handling may require fetching or browser control |
wkhtmltopdf |
Stable, mostly static HTML where Qt WebKit matches your output | Older WebKit behavior can differ from current browsers |
| WeasyPrint | Static HTML/CSS when a Python process is acceptable | Not a Rust crate and does not provide full browser JavaScript behavior |
Use Chromium when the source page contains client-side rendering, API calls, CSS Grid, flexbox, modern web fonts, or platform APIs. The Rust wkhtmltopdf crate is reasonable for controlled, mostly static documents. WeasyPrint is a useful separate service when the input is print-oriented HTML and JavaScript is unnecessary.
2. Install and verify Chrome
Install a Chromium or Google Chrome package appropriate for your operating system. Verify the executable path before wiring it into Rust:
which google-chrome || which chromium || which chromium-browser
google-chrome --version
In a container, keep the browser sandbox enabled when your user namespaces and permissions support it. The --no-sandbox switch is an environment-specific deployment decision; if you must use it, isolate the process and restrict its network access.
3. Minimal Rust conversion with a child process
Spawning the browser directly is the smallest dependable implementation. This example checks the exit status and confirms that Chrome created a non-empty file.
use std::{fs, io, path::Path, process::Command};
fn page_to_pdf(url: &str, output: &Path) -> io::Result<()> {
let status = Command::new("google-chrome")
.args([
"--headless",
"--disable-gpu",
"--no-sandbox",
"--no-pdf-header-footer",
"--timeout=15000",
"--print-to-pdf=page.pdf",
url,
])
.status()?;
if !status.success() {
return Err(io::Error::new(
io::ErrorKind::Other,
format!("Chrome exited with {status}"),
));
}
fs::rename("page.pdf", output)?;
let size = fs::metadata(output)?.len();
if size == 0 {
return Err(io::Error::new(io::ErrorKind::InvalidData, "empty PDF"));
}
Ok(())
}
fn main() -> io::Result<()> {
page_to_pdf("https://example.com/", Path::new("example.pdf"))
}
Use a unique temporary directory or filename for concurrent requests. The fixed page.pdf name above is intentionally simple for a single-process example; sharing it between requests creates races.
4. Wait for dynamic content before printing
A navigation can finish while a single-page application is still fetching data. Printing at that point produces a valid PDF containing incomplete content. Choose a readiness condition that matches the page:

- Load: suitable when all required content is present after the load event.
- Network idle: useful for pages that fetch data shortly after navigation, but analytics or polling can prevent a true idle state.
- Application selector: the most precise choice when the page exposes a marker such as
[data-pdf-ready]. - Bounded delay: a fallback for pages with no reliable signal. Always pair it with a maximum timeout.
Chrome supports a bounded --timeout. For deterministic virtual time, use --virtual-time-budget=42000 when page scripts need additional simulated time. The html2pdf CLI adds --wait and --wait-for values including navigation, load, and network-idle.
5. Use the Rust html2pdf CLI
html2pdf wraps the headless_chrome approach and exposes common print controls. Install it with Cargo:
cargo install html2pdf
html2pdf --wait-for network-idle --background --paper A4 \
--output page.pdf input.html
Its options include output path, landscape mode, background printing, explicit wait duration, readiness milestone, header and footer templates, paper size such as A4 or Letter, margins, scale, and page ranges. For a remote URL, either let a browser navigate directly to it or fetch the HTML first when your workflow requires local preprocessing.
A typical local-document command is:
html2pdf --wait 3000 --wait-for load --background \
--paper Letter --margin 0.5 --scale 0.95 \
--output report.pdf report.html
6. Browser-control code for print settings
When the CLI does not expose a setting you need, use a Rust Chrome DevTools Protocol binding such as headless_chrome. The exact API can vary by crate version, so pin the version in Cargo.toml and consult its current documentation. The control sequence remains the same:
- Launch or connect to Chrome.
- Open a tab and navigate to the URL.
- Wait for a lifecycle event or application selector.
- Inject print CSS or JavaScript if required.
- Call the browser print-to-PDF command with paper, margin, scale, orientation, background, and page-range settings.
- Write the returned bytes atomically and report conversion errors with the URL and browser version.
Print CSS often fixes pagination more cleanly than post-processing:
@media print {
@page { size: A4; margin: 16mm; }
.screen-only, nav, .chat-widget { display: none !important; }
.page-break { break-before: page; }
}
7. The wkhtmltopdf crate
The wkhtmltopdf crate requires a separately installed wkhtmltopdf binary. Its documentation provides build_from_html, build_from_url, and build_from_path, plus page size, orientation, margins, title, and output saving.
use wkhtmltopdf::*;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let app = PdfApplication::new()?;
let mut pdf = app.builder()
.orientation(Orientation::Landscape)
.margin(Size::Inches(0.5))
.build_from_url("https://example.com/")?;
pdf.save("example.pdf")?;
Ok(())
}
Wkhtmltopdf uses Qt WebKit. Test representative pages before selecting it for JavaScript-heavy workloads because its older engine can diverge from Chromium on CSS Grid, flexbox edge cases, web-platform APIs, and client-side rendering.
8. Production service design
Starting a browser for every request increases startup latency and memory pressure. A Rust service should keep a bounded pool of browser processes, cap concurrent tabs, recycle unhealthy instances, and use unique temporary files. The html2pdf-api project describes a thread-safe Chrome pool and settings such as CHROME_PATH, output filename, page ranges, and print configuration.
- Set navigation, readiness, and total-request timeouts.
- Limit maximum PDF size and temporary-disk usage.
- Cap concurrent conversions and queue excess work.
- Record the target URL, browser version, elapsed time, exit status, and output byte count.
- Write to a temporary file, fsync if durability matters, then rename atomically.
- Recycle a browser after crashes, repeated protocol errors, or excessive memory growth.
Security for user-supplied URLs
A renderer that fetches arbitrary URLs can become a server-side request forgery surface. Validate schemes, block localhost and private network ranges unless explicitly required, constrain redirects, disable unnecessary protocols, and run the browser in a restricted container. Do not pass untrusted strings through a shell; use Command::args so arguments are not interpreted as shell syntax.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
No such file or directory |
Chrome is not installed or is outside PATH |
Install it, set an absolute executable path, or configure CHROME_PATH. |
| Empty or missing PDF | Wrong output path, permissions, or a failed browser process | Check exit status, capture stderr, use a writable temporary directory, and verify file size. |
| PDF contains a loading spinner | Printing occurred before API data or JavaScript completed | Wait for a selector, network idle, or a bounded application-specific delay. |
| Fonts or images are missing | Resources are blocked, cross-origin credentials are absent, or the process exits too early | Inspect browser logs, provide required cookies or headers, and wait for fonts and images. |
| Headers and footers appear unexpectedly | Chrome print decorations are enabled | Add --no-pdf-header-footer or configure explicit templates. |
| Layout differs from the browser | Different viewport, device scale, print CSS, or engine | Set viewport and scale explicitly, test print media CSS, and prefer Chromium for modern pages. |
| Conversion hangs | Polling, WebSockets, ads, or a never-ending request prevents idle | Use a selector or fixed readiness rule plus a hard timeout; block unnecessary resources. |
10. Performance, reliability, and cost
No universal throughput or memory number applies: page complexity, browser version, fonts, images, JavaScript, and concurrency dominate. Measure on representative URLs. Track cold-start and warm-pool latency separately, along with PDF size, timeout rate, crash rate, and queue depth.

Pooling reduces repeated startup work, while too much concurrency causes CPU and memory contention. Keep a small bounded pool, apply backpressure, and load-test before selecting limits. Cache immutable URLs when appropriate, but include authentication, locale, viewport, and print settings in the cache key. For reproducibility, pin the browser image and Rust dependencies, and keep a small set of golden PDFs for visual regression checks.
Self-hosting costs include browser processes, container memory, storage, bandwidth, and engineering time. A managed renderer trades infrastructure work for an API bill. Compare based on your actual page mix and required isolation; the research dossier does not provide a comparable benchmark.
Or skip the browser setup
ScreenshotNeo provides a hosted screenshot and PDF API. It accepts a URL with one GET request and can render JavaScript-heavy pages without you packaging Chrome. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options. A PDF request uses the same endpoint:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-d format=pdf \
-o page.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://example.com",
"format": "pdf",
},
timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com',
format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('page.pdf', Buffer.from(await res.arrayBuffer()));
Relevant PDF controls include paper size, margins, landscape mode, and page ranges. You can also set custom headers, cookies, user agent, authorization, timezone, geolocation, custom JavaScript or CSS, selector waits, delays, network-idle waits, resource blocking, caching TTL, signed links, asynchronous jobs with signed webhooks, and bulk capture of up to 100 URLs per call. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Can Rust convert a URL without saving HTML first?
Yes. Chrome can navigate directly to the URL, and a Rust process can invoke its print-to-PDF operation. Saving HTML first is useful only when you need to rewrite or sanitize the document.
How do I include background colors?
Enable background printing in your Chrome control call or use the --background option with html2pdf. Also verify that the page’s print CSS does not remove those backgrounds.
Which engine should I use for a React or Vue application?
Use headless Chromium with an explicit readiness signal. Older WebKit-based tools may not execute the application or match its CSS layout.
How can I make page breaks predictable?
Add print rules with break-before, break-after, and break-inside, then set paper size and margins explicitly.
Should I use a browser pool for a CLI?
No. A one-off CLI can spawn Chrome per conversion. A long-running service benefits from a bounded pool and strict resource limits.


