ScreenshotNeo

BlogHow-to

How to Convert HTML to PDF with an Open-Source Library

Choose WeasyPrint for document HTML/CSS, Puppeteer for JavaScript apps, and learn how to produce reliable, secure PDFs.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: use WeasyPrint when your input is primarily HTML and print CSS, and you want a Python API for reports, invoices, certificates, or other document-shaped output. Use Puppeteer when the page needs JavaScript, browser APIs, application data, or Chromium-compatible CSS. Keep wkhtmltopdf for legacy compatibility; its stable 0.12.6 series was released on June 11, 2020.

This guide shows both rendering models, complete commands and code, print CSS, asset handling, security controls, validation, troubleshooting, and a hosted option when you do not want to operate a browser.

1. Choose the renderer before writing code

Input and requirement Best starting point Reason
Static HTML/CSS, reports, invoices, certificates WeasyPrint Python-first paged-media engine with print layout controls and document features.
React, Vue, dashboards, charts, client-side data, browser APIs Puppeteer Runs a real Chromium page, including JavaScript and browser rendering.
Existing scripts tied to an older WebKit renderer wkhtmltopdf Useful for compatibility, but it is legacy software and requires strict input controls.

Classify the page by asking:

  • Is all required content present in the initial HTML, or does JavaScript fetch and render it?
  • Do you need browser APIs, canvas behavior, or Chromium CSS?
  • Do you need print features such as bookmarks, attachments, forms, PDF/A, or PDF/UA?
  • Can every image, font, stylesheet, and link be resolved from the conversion environment?

2. Convert HTML with WeasyPrint in Python

Install the library and native dependencies

WeasyPrint requires Python and native Pango-related libraries. Install those according to your operating system, then install the Python package and verify the executable:

# Create and activate a virtual environment
python -m venv .venv
. .venv/bin/activate

pip install weasyprint
weasyprint --info

The WeasyPrint installation documentation lists platform-specific native packages and verification steps. In a minimal container, install the documented Pango and related libraries in the image rather than assuming a desktop environment is present.

Minimal Python conversion

from weasyprint import HTML

HTML(string="""


  
    
    Invoice
    
  
  
    

Invoice 1001

Generated from HTML and CSS.

""").write_pdf("invoice.pdf")

Convert a file and resolve relative assets

from pathlib import Path
from weasyprint import HTML

source = Path("templates/report.html").resolve()
output = Path("build/report.pdf")
output.parent.mkdir(parents=True, exist_ok=True)

HTML(filename=str(source), base_url=str(source.parent)).write_pdf(str(output))
print(output)

base_url matters when the document contains relative paths such as images/logo.png, styles/print.css, or font URLs. An alternative is to make every asset URL absolute. For production, make the resolution rule explicit and validate that required files exist.

Use a stylesheet and print-specific rules

from weasyprint import HTML, CSS

HTML(filename="report.html", base_url=".").write_pdf(
    "report.pdf",
    stylesheets=[CSS(filename="print.css")],
)
/* print.css */
@page {
  size: A4;
  margin: 20mm 16mm 18mm;
  @bottom-right { content: counter(page) " / " counter(pages); }
}

.screen-only { display: none; }
.keep-together { break-inside: avoid; }
table { width: 100%; border-collapse: collapse; }
thead { display: table-header-group; }
tr { break-inside: avoid; }
.page-break { break-before: page; }

Check the exact CSS features your document uses against WeasyPrint’s support documentation. Advanced layout features can be unsupported or partial, so validate page breaks, tables, fonts, and generated content with representative documents.

Document features to validate

WeasyPrint can preserve hyperlinks, create bookmarks, include attachments, generate forms, and produce PDF/A or PDF/UA output. Select and validate those features against your compliance requirements rather than assuming that a visually correct page meets an accessibility or archival standard.

3. Convert JavaScript-rendered HTML with Puppeteer

Install Chromium automation

mkdir html-to-pdf && cd html-to-pdf
npm init -y
npm install puppeteer

Complete Node.js example

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({
    // In a container, you may need the system Chromium path and sandbox settings
    // required by that environment.
  });

  try {
    const page = await browser.newPage();
    await page.goto('https://example.com/report', {
      waitUntil: 'networkidle0',
      timeout: 60_000,
    });

    // Wait for application data and fonts before creating the PDF.
    await page.waitForSelector('[data-report-ready]', { timeout: 30_000 });
    await page.evaluate(() => document.fonts.ready);

    // page.pdf() uses print CSS media by default.
    await page.pdf({
      path: 'report.pdf',
      format: 'A4',
      printBackground: true,
      preferCSSPageSize: true,
      margin: { top: '18mm', right: '16mm', bottom: '18mm', left: '16mm' },
      displayHeaderFooter: false,
    });
  } finally {
    await browser.close();
  }
})();

page.pdf() uses the print CSS media type by default. If the screen stylesheet is the intended source, call await page.emulateMediaType('screen') before generating the PDF. Wait for application data and document.fonts.ready; otherwise you can capture an empty shell or fallback fonts. See the Puppeteer PDF API for the current option names.

Convert local HTML

const path = require('path');
const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch();
  try {
    const page = await browser.newPage();
    await page.goto(`file://${path.resolve('report.html')}`, {
      waitUntil: 'networkidle0',
    });
    await page.evaluate(() => document.fonts.ready);
    await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });
  } finally {
    await browser.close();
  }
})();

For local assets, confirm that the Chromium process can read the files and that CSS uses correct relative paths. For untrusted input, isolate the browser and restrict file and network access.

4. wkhtmltopdf for legacy compatibility

wkhtmltopdf \
  --page-size A4 \
  --margin-top 18mm \
  --margin-right 16mm \
  --margin-bottom 18mm \
  --margin-left 16mm \
  report.html report.pdf

wkhtmltopdf is an LGPLv3 Qt WebKit command-line tool. The project lists 0.12.6 as its stable series, released June 11, 2020. It may be appropriate when an existing workflow depends on its output, but it is not a modern browser engine. The project gives a direct security warning: “Do not use wkhtmltopdf with any untrusted HTML”; sanitize user-supplied HTML and JavaScript and isolate the process.

5. Print CSS that survives real documents

  1. Set an explicit @page size and margins.
  2. Use break-before, break-after, and break-inside for section and table control.
  3. Repeat table headings with thead { display: table-header-group; }.
  4. Keep rows, cards, signatures, and small figures together where possible.
  5. Hide navigation, cookie notices, controls, and other screen-only elements in print media.
  6. Use print-safe colors and printBackground: true in Puppeteer when backgrounds are part of the design.
  7. Load the exact fonts used in production and verify that they are embedded or otherwise available in the generated PDF.

6. Assets, URLs, fonts, and dynamic data

Most conversion failures are resource-resolution failures. Use absolute HTTPS URLs or a deliberate base URL, make private assets available with authenticated requests, and confirm that the renderer can reach DNS, certificates, and internal services. Avoid relying on a browser cache from a previous request.

For JavaScript applications, wait for a deterministic readiness signal such as [data-report-ready], not only a short sleep. If the page loads charts or images after the initial network becomes idle, wait for those specific elements as well. For WeasyPrint, remember that it renders HTML/CSS as a document and does not execute application JavaScript.

7. Security controls for untrusted HTML

  • Sanitize HTML and CSS before rendering.
  • Do not allow arbitrary filesystem URLs or unrestricted network requests.
  • Run conversion in a low-privilege container or sandbox with resource limits.
  • Set CPU, memory, page-count, and timeout limits.
  • Restrict outbound network access and prevent access to cloud metadata or internal control-plane endpoints.
  • Store generated PDFs outside executable paths and return them with safe content-disposition headers.

WeasyPrint documents security concerns with untrusted sources, and wkhtmltopdf warns that unsafe HTML/JavaScript can lead to complete server takeover. Treat both HTML and CSS as potentially hostile input.

8. Validate the generated PDF

  1. Open the PDF in more than one viewer.
  2. Check page count, page size, margins, headers, footers, and intentional blank pages.
  3. Inspect long tables for split rows and repeated headings.
  4. Verify hyperlinks, bookmarks, attachments, form fields, and metadata when required.
  5. Check fonts, non-Latin scripts, ligatures, images, transparency, and color output.
  6. Run accessibility and archival validation when targeting PDF/UA or PDF/A.
  7. Compare representative documents in CI so renderer upgrades do not silently change pagination.

9. Troubleshooting

Symptom Likely cause Fix
Images or CSS are missing in WeasyPrint Relative URLs have no usable base URL, or files are inaccessible. Pass base_url, use absolute URLs, and verify permissions and paths.
PDF contains only a loading shell JavaScript data had not rendered. Use Puppeteer, wait for a readiness selector, and wait for fonts and required assets.
Fonts fall back or characters disappear Font files are unavailable, blocked, or not loaded before capture. Make fonts reachable, declare them in CSS, and await document.fonts.ready in Puppeteer.
Screen layout differs from PDF Print media rules are active. Keep print CSS intentional, or call emulateMediaType('screen') when screen styles are required.
Rows split badly across pages Page-break rules are missing or unsupported for the chosen layout. Add break-inside: avoid, repeat table headers, and test with long data.
Puppeteer times out Slow dependency, blocked request, infinite polling, or an overly short timeout. Log failed requests, wait for a specific readiness condition, and set a bounded but realistic timeout.
Chromium will not launch in a container Missing system libraries, executable path, or sandbox configuration. Use a compatible Puppeteer image or system Chromium setup and follow its container requirements.
wkhtmltopdf output is wrong for modern CSS Its Qt WebKit engine is old. Move to WeasyPrint for document CSS or Puppeteer for browser rendering.
Conversion exposes internal resources Untrusted HTML can request arbitrary URLs or files. Sanitize input, restrict network and filesystem access, and isolate the renderer.

10. Performance, reliability, and cost

Rendering cost depends on page complexity, asset size, JavaScript execution, fonts, and the renderer startup model. Reuse a controlled browser process for batches when using Puppeteer, but recycle it on a schedule and enforce per-job limits. Cache immutable assets and avoid loading analytics, advertisements, or unnecessary third-party resources. For WeasyPrint, keep templates and assets local or on reliable internal storage and avoid repeatedly downloading the same resources.

Reliability comes from deterministic readiness checks, bounded timeouts, structured logs, retries only for transient failures, and validation of the resulting PDF. Record the input version, renderer version, options, and output hash so a changed document can be diagnosed.

Open-source software has no per-request license charge, but you still pay for compute, memory, storage, browser images, native dependencies, and engineering time. Size workers from observed document complexity rather than assuming that one HTML page equals one unit of work.

11. Or skip the browser setup

ScreenshotNeo is a website screenshot API that can also return PDFs. It handles the capture service for you while keeping the request small. See the ScreenshotNeo API documentation for the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For PDF output, request the PDF format and configure paper size, margins, landscape mode, or page ranges with the documented parameters. ScreenshotNeo can accept cookie and consent banners, remove more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and let you turn each cleanup step off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan.

Create a free ScreenshotNeo account to get 1,000 screenshots each month without a card.

12. FAQ

Can WeasyPrint run JavaScript?

Use Puppeteer for pages whose required content is produced by JavaScript. WeasyPrint is designed for HTML/CSS document rendering.

Should I use print or screen CSS?

Use print CSS for document output. Puppeteer uses print media by default; explicitly emulate screen media only when that is the layout you intend to publish.

Is wkhtmltopdf maintained like Chromium?

No. Its project lists the 0.12.6 stable series from June 11, 2020, so treat it as a legacy compatibility choice.

Why does a PDF look correct locally but fail in production?

The environments often differ in fonts, native libraries, network access, file paths, browser versions, or JavaScript timing. Pin the runtime, make assets explicit, and validate in the same image used for production.

Which option is simplest for a Python report service?

Start with WeasyPrint when the report is document-shaped and its required CSS is supported. Move to Puppeteer when browser behavior or JavaScript is essential.