ScreenshotNeo

BlogComparisons

The Best HTML-to-PDF Conversion Tools for Developers

Compare Playwright, Puppeteer, WeasyPrint, managed APIs, and legacy tools for reliable HTML-to-PDF conversion.

By the ScreenshotNeo team1 October 20269 min read

Short answer: Use Playwright or Puppeteer when your HTML depends on JavaScript, charts, web fonts, or modern CSS. Use WeasyPrint, Paged.js, Vivliostyle, or OpenHTMLtoPDF for print-first documents with controlled templates and little JavaScript. Use a managed browser or PDF API when maintaining Chromium, scaling workers, patching security issues, and handling retries is not a good use of your team’s time.

This guide compares the main approaches, includes runnable Node.js, Python, and cURL examples, and covers pagination, fonts, dynamic content, security, performance, reliability, and migration from wkhtmltopdf.

Which HTML-to-PDF tool should you choose?

Use case Best starting point Why
JavaScript dashboards, SPAs, charts, or interactive pages Playwright or Puppeteer A real browser executes JavaScript and applies modern CSS.
Node-only service with a simple API Puppeteer page.pdf() provides a direct browser-to-PDF workflow.
Polyglot team or multiple browser engines Playwright It offers bindings beyond Node and browser selection.
Print-first reports with little JavaScript WeasyPrint, Paged.js, Vivliostyle, or OpenHTMLtoPDF Dedicated paged-media engines can provide deterministic pagination without a browser binary.
Burst traffic or no browser operations team A managed API such as Browserless or Doppio The provider operates browser infrastructure, scaling, and much of the patching work.
Existing wkhtmltopdf installation Plan a migration Comparison sources describe wkhtmltopdf as archived or unmaintained, with outdated JavaScript and CSS support.

How browser-based HTML-to-PDF conversion works

  1. Launch a browser process or connect to a managed browser.
  2. Navigate to a URL or load an HTML document.
  3. Wait for the document and its assets to reach a known state.
  4. Choose print or screen media, page size, margins, headers, and footers.
  5. Wait for fonts, charts, images, and application data.
  6. Generate the PDF and close or recycle the browser context.

Browser rendering is the most faithful option for client-rendered pages because the same layout engine that users see performs the conversion. It also means you must control navigation timeouts, untrusted content, network access, browser versions, and resource usage.

Playwright: the best general-purpose choice for modern pages

Playwright is a practical default when you need JavaScript execution, modern CSS, web fonts, charts, and language bindings beyond Node. The comparison evidence does not establish a universal performance or cost advantage over Puppeteer, so select based on your runtime, browser coverage, and operational preferences.

Node.js example

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
  viewport: { width: 1440, height: 900 },
  deviceScaleFactor: 1
});

await page.goto('https://example.com/report', {
  waitUntil: 'networkidle',
  timeout: 90000
});
await page.emulateMedia({ media: 'print' });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
  path: 'report.pdf',
  format: 'A4',
  printBackground: true,
  preferCSSPageSize: true,
  margin: { top: '18mm', right: '14mm', bottom: '18mm', left: '14mm' }
});
await browser.close();

For pages whose visual design is defined by screen media queries, use page.emulateMedia({ media: 'screen' }) before generating the PDF. Use a deterministic readiness condition, such as a report-complete selector, when networkidle is unreliable because of analytics, WebSockets, or polling.

Python example

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1440, "height": 900})
    page.goto("https://example.com/report", wait_until="networkidle", timeout=90000)
    page.emulate_media(media="print")
    page.evaluate("document.fonts.ready")
    page.pdf(
        path="report.pdf",
        format="A4",
        print_background=True,
        prefer_css_page_size=True,
        margin={"top": "18mm", "right": "14mm", "bottom": "18mm", "left": "14mm"},
    )
    browser.close()

Puppeteer: a direct Node.js workflow

Puppeteer’s official guide uses page.goto(), a network-idle wait, page.pdf(), and browser shutdown. Its API reference states that Page.pdf() generates a PDF with the print CSS media type and waits for fonts by default. If you need screen styles, call page.emulateMediaType('screen'). Print colors may be modified unless your CSS uses -webkit-print-color-adjust: exact.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.goto('https://example.com/report', {
  waitUntil: 'networkidle2',
  timeout: 90000
});
await page.emulateMediaType('print');
await page.evaluate(() => document.fonts.ready);
await page.pdf({
  path: 'report.pdf',
  format: 'A4',
  printBackground: true,
  preferCSSPageSize: true,
  margin: { top: '18mm', right: '14mm', bottom: '18mm', left: '14mm' }
});
await browser.close();

Use networkidle2 only when the page eventually becomes quiet. For applications with persistent connections, wait for an application-owned selector instead.

@page {
  size: A4;
  margin: 18mm 14mm;
}

@media print {
  .screen-only,
  .cookie-banner,
  .chat-widget,
  .actions {
    display: none !important;
  }

  * {
    -webkit-print-color-adjust: exact;
    print-color-adjust: exact;
  }

  h1, h2, h3 {
    break-after: avoid;
  }

  table, figure, pre {
    break-inside: avoid;
  }

  .page-break {
    break-before: page;
  }
}

@font-face {
  font-family: 'Report Sans';
  src: url('/fonts/report-sans.woff2') format('woff2');
  font-display: block;
}
  • Set @page size and margins explicitly.
  • Hide navigation, controls, cookie banners, and chat widgets in print CSS.
  • Use break rules around headings, tables, figures, and code blocks.
  • Load a font that the renderer can reach; do not rely on a developer workstation’s installed fonts.
  • Use preferCSSPageSize when CSS should control the paper size.

WeasyPrint and dedicated paged-media engines

WeasyPrint is a strong option for print-first documents with little or no client-side JavaScript. Paged.js, Vivliostyle CLI, and OpenHTMLtoPDF are other dedicated paged-media choices. These engines avoid a full browser in deployments where deterministic pagination matters more than rendering a JavaScript application.

from weasyprint import HTML

HTML("https://example.com/invoice").write_pdf("invoice.pdf")

For an HTML string, pass string=... and provide a base_url so relative images, stylesheets, and fonts resolve correctly:

from weasyprint import HTML

html = """<html><body><h1>Invoice</h1><p>Paid</p></body></html>"""
HTML(string=html, base_url="/srv/templates").write_pdf("invoice.pdf")

Do not choose WeasyPrint expecting browser-level execution of a React application or client-rendered chart. Render the data into HTML first, or use a browser engine.

Managed browser and PDF APIs

Browserless documents a REST /pdf endpoint for exporting a webpage or HTML and also supports browser connections for waiting on dynamic content. Its examples include A4 format, print backgrounds, headers, and footers. Managed APIs can be useful when your team wants REST or asynchronous jobs without operating Chromium workers.

Doppio’s options guide describes three approaches: legacy wkhtmltopdf, self-hosted Puppeteer or Playwright, and managed HTML-to-PDF APIs. Its claims about operational savings are vendor opinions; evaluate data residency, quotas, SLA, retry behavior, pricing, and delivery options before committing.

For any managed provider, verify:

  • Whether JavaScript execution and custom fonts are supported.
  • Navigation and rendering timeout limits.
  • Paper sizes, margins, headers, footers, page ranges, and background colors.
  • Webhook signing, idempotency, retries, and downloadable artifact retention.
  • Network egress rules, authentication, cookies, and private URLs.
  • Data retention and regional processing requirements.
  • Concurrency limits and burst behavior.

Why wkhtmltopdf is a migration target

Research sources describe wkhtmltopdf as archived or unmaintained and warn about unpatched vulnerabilities and outdated JavaScript and CSS support. Existing static templates may continue to render, but new systems should plan a controlled migration.

  1. Inventory representative templates, including charts, tables, fonts, RTL text, and long pages.
  2. Capture baseline PDFs from the current system.
  3. Rebuild templates with print CSS and a maintained renderer.
  4. Compare page count, line wrapping, images, colors, links, and metadata.
  5. Run the new renderer in parallel before switching production traffic.
  6. Keep a rollback path while fixing template-specific differences.

Options that matter in production

Option Recommendation
Wait strategy Prefer a known selector or application-ready flag. Use network idle only when the page has no long-lived requests.
Media type Use print for document CSS; use screen when the page’s layout is screen-specific.
Backgrounds Enable print backgrounds and set print color adjustment when brand colors are required.
Fonts Wait for document.fonts.ready and ensure every font URL is reachable from the renderer.
Page size Choose either CSS @page sizing or an explicit API size and keep it consistent.
Headers and footers Use renderer templates or CSS margin boxes where supported; test page numbers and long titles.
Page ranges Generate all pages first when content height is dynamic, then select ranges if the API supports it.
Assets Use absolute or correctly based URLs for images, stylesheets, and fonts.
Authentication Pass cookies, headers, or an authenticated session without exposing credentials in URLs.
Security Restrict outbound requests, isolate browser workers, limit PDF size, and never render untrusted HTML with powerful local access.

Common errors and fixes

Symptom Likely cause Fix
PDF contains a blank shell SPA JavaScript has not finished rendering. Wait for an application-owned selector or readiness flag instead of navigating and printing immediately.
Charts are missing Canvas or chart data is rendered after the capture. Wait for the chart container, animation completion, and required network requests.
Fonts fall back Font files are blocked, unresolved, or still loading. Use a correct base URL, allow the font origin, and await document.fonts.ready.
Colors look washed out Print color adjustment changed the output. Enable print backgrounds and use -webkit-print-color-adjust: exact where appropriate.
Content is cut off Fixed-height containers or overflow rules were designed for screens. Remove fixed heights in print CSS and inspect overflow, flex, and grid rules.
Unexpected extra pages Margins, large unbreakable elements, or forced breaks. Reduce margins, allow table and paragraph breaks, and use break-inside selectively.
Network-idle timeout Analytics, polling, WebSockets, or ads keep requests open. Block unnecessary resources or use a selector-based readiness condition.
Images are absent Relative URLs have no base, authentication is missing, or the image is lazy-loaded. Use absolute URLs or a base URL, pass credentials safely, and scroll or trigger lazy loading before printing.
Browser crashes under load Too many concurrent pages, large documents, or leaked browser processes. Bound concurrency, recycle contexts, cap document size, and monitor memory.
Different output after an upgrade Browser version or font changes altered layout. Pin browser and font versions, keep golden PDFs, and review diffs during upgrades.

Performance, reliability, and cost

Performance

Browser startup is usually more expensive than creating a page in an already-warm browser. Reuse a controlled browser process, create isolated contexts per job, bound concurrency, and avoid loading ads, trackers, and unrelated assets. A PDF4.dev 2026 benchmark reports 13 ms for Playwright, 58 ms for Puppeteer, and 629 ms for WeasyPrint on one warm, complex-document workload. Treat those figures as workload-specific rather than a universal ranking.

Reliability

  • Set separate navigation and PDF-generation timeouts.
  • Retry transient navigation failures with a limit and backoff.
  • Make jobs idempotent so a retry cannot create duplicate records.
  • Record browser version, URL, readiness condition, page count, and failure reason.
  • Keep a small set of representative PDFs for visual regression checks.
  • Fail clearly when authentication expires or required assets cannot load.

Cost

Self-hosting moves the bill into browser CPU, memory, storage, engineering time, patching, and operations. Managed APIs trade that work for per-request pricing and provider limits. Compare total cost at your actual document size, concurrency, retry rate, and peak traffic rather than comparing only an advertised request price.

Or skip the browser setup

ScreenshotNeo provides a website capture API and MCP server. Its PDF and capture workflow handles the browser layer through one request, while also removing cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for request options. The same endpoint also supports full-page capture, element selection, custom CSS and JavaScript, waits, blocking rules, cookies, headers, user agents, timezone, geolocation, caching, signed links, asynchronous jobs, bulk capture, and PDF settings.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes 1,000 shots per month free with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Can Playwright or Puppeteer render JavaScript charts?

Yes. They drive a real browser, but you must wait for the chart data and rendering to finish before printing.

Should I use print or screen media?

Use print for document-specific styles. Use screen when the page’s existing layout is intended for the screen and should remain unchanged.

Is WeasyPrint faster than a browser?

It can be efficient for simple, print-first documents, but performance depends on the document and workload. The cited benchmark is not a universal ranking.

Is wkhtmltopdf safe for a new service?

It is a poor starting point for new systems because research sources describe it as archived or unmaintained. Plan a migration for existing deployments.

When should I use a managed API?

Use one when browser patching, scaling, retries, and burst capacity would otherwise require significant operational work. Confirm its limits, data handling, and pricing first.