ScreenshotNeo

BlogHTML to image & PDF

How to Convert a Webpage to PDF in Rust

Use headless Chromium from Rust to render a live webpage as a PDF, with control over readiness, print styles, page layout, and failures.

By the ScreenshotNeo team4 October 20268 min read

To convert a live webpage to PDF in Rust, use Chromium as the renderer and Rust to control it. A practical route is the Rust Playwright API: launch headless Chromium, navigate to the page, wait for the content you need, set print options, and call the PDF builder. PDF generation in this API is supported only in Chromium headless mode. You can also run Chrome Headless as a subprocess for simpler batch jobs.

Choose an implementation

Approach Best fit Trade-off
Rust Playwright API Applications that need navigation, readiness waits, and configurable print output. More browser-control code and a compatible Chromium installation.
Chrome Headless CLI Simple scripts or batch jobs that print URLs with default browser behavior. Your application must manage the process, installation, timeouts, and errors.
chromiumoxide Rust applications that want to drive Chromium through the DevTools Protocol. More direct protocol-level integration; check the selected version’s current compatibility.
html2pdf A CLI-shaped Rust option for HTML-to-PDF conversion. Its package documentation describes it as a wrapper over headless_chrome; verify release status and requirements before adopting it.

The available documentation does not establish a current maintenance or platform-support ranking among these crates. Check release history, supported platforms, and Chromium compatibility for the versions you plan to deploy.

Convert a webpage with Rust Playwright

The example below illustrates the flow using the Rust Playwright API. Pin crate versions in your project and follow their current installation instructions for the Playwright driver and compatible Chromium browser. The exact setup commands can change between releases, so verify them against the version you choose.

// Illustrative Rust Playwright flow. Adapt imports and builder methods to the
// version pinned in your Cargo.toml; see the current Rust Playwright API docs.
use playwright::Playwright;
use std::error::Error;

#[tokio::main]
async fn main() -> Result<(), Box<dyn Error>> {
    let url = std::env::args()
        .nth(1)
        .unwrap_or_else(|| "https://example.com".to_string());

    let playwright = Playwright::initialize().await?;
    let chromium = playwright.chromium();
    let browser = chromium.launcher().headless(true).launch().await?;
    let page = browser.new_page().await?;

    // Navigation completion does not guarantee every site's asynchronous
    // content is ready. Wait for a meaningful selector when the page has one.
    page.goto_builder(&url).goto().await?;
    page.wait_for_load_state(playwright::api::LoadState::NetworkIdle)
        .await?;

    // Configure the PDF builder for the crate version you use. Common controls
    // include paper format, margins, background printing, orientation, and range.
    let pdf = page.pdf_builder()
        .format("A4")
        .print_background(true)
        .build()
        .await?;
    std::fs::write("page.pdf", pdf)?;

    browser.close().await?;
    Ok(())
}

This is an API-shaped example rather than a version-pinned package recipe: Rust Playwright crate method names and setup requirements should be checked against the exact crate release in your lockfile. The Page API documents a PDF builder, print behavior, and layout controls. See the Rust Playwright Page API.

Set readiness deliberately

Choose a readiness condition based on the target site. A network-idle wait can be useful, but pages with analytics, polling, or long-lived requests may never become idle. For a page you control, waiting for a selector that appears when the report or main content is ready is often more precise. A fixed delay is a fallback for known delayed rendering, not a guarantee that content has loaded.

  1. Navigate to the URL and handle navigation errors.
  2. Wait for the main content or a known completion signal.
  3. Wait for fonts or images if their presence matters to the document.
  4. Choose print or screen media and apply paper, margin, color, and background settings.
  5. Write the returned PDF bytes to a file or stream them to your application.

Use Chrome Headless from Rust

For a small command-line workflow, Rust can spawn Chrome and let the browser print the page. Chrome documents --headless and --print-to-pdf; --no-pdf-header-footer suppresses the default header and footer. Use a process timeout and inspect the exit status so a failed capture does not look like a successful empty output.

use std::process::{Command, Stdio};
use std::time::Duration;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let url = "https://example.com";
    let output_path = "page.pdf";
    let chrome = std::env::var("CHROME_BIN").unwrap_or_else(|_| "google-chrome".into());

    let mut child = Command::new(chrome)
        .arg("--headless")
        .arg("--disable-gpu")
        .arg("--no-pdf-header-footer")
        .arg(format!("--print-to-pdf={output_path}"))
        .arg(url)
        .stdout(Stdio::null())
        .stderr(Stdio::piped())
        .spawn()?;

    let timeout = Duration::from_secs(60);
    let started = std::time::Instant::now();
    loop {
        if let Some(status) = child.try_wait()? {
            if !status.success() {
                return Err(format!("Chrome exited with {status}").into());
            }
            break;
        }
        if started.elapsed() > timeout {
            child.kill()?;
            return Err("Chrome PDF capture timed out".into());
        }
        std::thread::sleep(Duration::from_millis(100));
    }
    Ok(())
}

Production code should also collect Chrome’s stderr, validate that the output file exists and is non-empty, and clean up a partial file after failure. Chrome’s CLI documentation describes printing URLs to PDF, header/footer control, and timeout behavior: Chrome Headless documentation.

Configure PDF appearance

PDF output follows browser print behavior. Rust Playwright’s Page API documents these relevant controls:

Control What to decide
Media PDF generation uses print CSS by default. Emulate screen media first if the PDF should match the screen layout.
Paper and dimensions Use a standard format such as A4 or Letter, or explicit dimensions with units.
Margins Set margins to preserve content and allow room for headers or footers.
Orientation Choose portrait or landscape for the document’s content.
Backgrounds Enable background printing when colored sections or background graphics matter.
Page ranges Request only the pages needed for long output, where supported.
Headers and footers Use templates for page metadata when supported. Template scripts do not run, and page styles are not visible inside the templates.
Colors Browsers adjust colors for print by default. Use -webkit-print-color-adjust when exact colors are needed.

Page breaks and print-specific styles belong in the page’s CSS when you control it. For third-party pages, inspect whether their print stylesheet hides navigation, changes widths, or omits content. If screen styling is required, select screen media before producing the PDF, then verify that the result remains printable on the chosen paper size.

Use chromiumoxide or html2pdf

chromiumoxide exposes Chromium DevTools Protocol’s PrintToPdfParams. It is an option when you want to manage a Chromium session from Rust at the protocol level. You still need to launch or connect to Chromium, navigate, wait for readiness, set print parameters, and handle browser and process errors.

The html2pdf crate presents a CLI-oriented route and documents itself as a wrapper over headless_chrome. Before using it, check its current release, dependency requirements, and support for your target platform. The available research does not verify those details for a particular release.

Deployment, performance, and reliability

  • Browser availability: Ship or install a compatible Chromium binary, and account for its platform and runtime dependencies in your deployment image.
  • Startup cost: Starting a browser for every URL adds process startup overhead. For repeated captures, consider a long-lived browser with isolated pages or contexts, while managing concurrency and cleanup.
  • Concurrency: Bound the number of simultaneous pages and browser processes. Excessive parallel captures compete for memory and CPU.
  • Timeouts: Set separate limits for navigation, readiness, PDF generation, and the overall job. Kill and reap a subprocess that exceeds its deadline.
  • Version management: Pin and update the browser and automation crate deliberately. Browser changes can affect rendering and PDF output.
  • Repeatability: Dynamic content, remote fonts, ads, current timestamps, and network variation can change output. For stable reports, control the page data and wait for explicit readiness.
  • Validation: Treat a successful navigation as distinct from a valid PDF. Check errors, output size, and—where the workflow requires it—the resulting page count or content.

There is no benchmark in the cited research that establishes which Rust route is fastest. Measure against your own pages, browser version, deployment hardware, and concurrency target.

Troubleshooting

Symptom Likely cause Fix
Browser executable not found Chromium is missing or the executable path differs in the deployment environment. Install a compatible browser and configure its path, such as through an application setting or CHROME_BIN.
PDF API reports unsupported operation The browser is not running in headless Chromium mode, or the selected browser is unsupported. Use Chromium headless for the documented Playwright PDF API.
Blank or incomplete PDF Capture began before client-side rendering or asynchronous data completed. Wait for a content-specific selector or completion signal; verify the page in the same browser environment.
Network-idle wait hangs The page keeps requests open or sends recurring requests. Wait for a specific selector or use a bounded delay suited to the page, with a hard overall timeout.
PDF looks different from the browser Print CSS is active by default and may change layout or hide content. Emulate screen media before printing if screen appearance is required, then review page breaks and scaling.
Background colors or graphics are missing Background printing is disabled or print color adjustment changed the colors. Enable background printing and apply -webkit-print-color-adjust in page CSS when exact colors matter.
Headers or footers are absent or malformed Template limitations or default browser header/footer settings are involved. Use the documented header/footer options; do not rely on template scripts or page styles inside templates.
Chrome exits but the PDF is missing Invalid output path, permission issue, failed navigation, or browser error. Capture stderr, check the exit status, ensure the directory is writable, and verify the output file after completion.
Capture exceeds the job deadline Slow site, stalled requests, overloaded browser, or overly strict readiness condition. Use bounded waits, tune concurrency, and report timeout separately from a valid PDF result.

Or skip the browser setup

For a one-request webpage-to-PDF workflow, ScreenshotNeo accepts a URL and returns a PDF. See the ScreenshotNeo API documentation for the PDF parameters and response behavior.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up free.

FAQ

Can I convert an HTML string instead of a URL?

Yes, if you load or set the HTML in a Chromium page before printing. Relative assets need a usable base URL or absolute paths, and external assets must be reachable from the browser environment.

Does a PDF capture include content below the fold?

Browser PDF printing lays out the document across pages rather than taking only the visible viewport. Lazy-loaded sections may still require scrolling or another readiness step before printing.

Can I use a browser other than Chromium with Rust Playwright PDF?

The cited Rust Playwright Page API documents PDF generation as Chromium-headless-only. Use Chromium for that API path.

Which route should a small Rust service start with?

Use the CLI for a straightforward subprocess job. Choose a Rust browser API when the service needs precise readiness waits, page interaction, or per-job print configuration.