ScreenshotNeo

BlogHTML to image & PDF

HTML to PDF Conversion: Complete Developer Guide

Convert HTML to reliable PDFs with browser APIs, CLI tools, print CSS, troubleshooting, and a practical ScreenshotNeo option.

By the ScreenshotNeo team29 September 20269 min read

HTML to PDF Conversion: Complete Developer Guide

HTML to PDF conversion means rendering an HTML document with CSS, fonts, images, and (sometimes) JavaScript into a paginated PDF file. The most dependable choice for modern, dynamic pages is a headless browser such as Puppeteer or Playwright. Browser APIs execute page scripts and expose print controls, but they use print CSS by default, so you must design and test the print output separately from the screen layout.

For static templates, a command-line renderer such as wkhtmltopdf or a Python library such as WeasyPrint may be simpler. For publishing workflows that need advanced paged-media features, Prince documents generated content, page regions, footnotes, and server-side integration in its user guide.

Choose a conversion approach

Approach Best fit Key considerations
Puppeteer or Playwright Apps with JavaScript, charts, authenticated pages, and modern CSS Runs a browser; PDF uses print media by default; wait for assets before export
wkhtmltopdf Simple server or command-line jobs Headless Qt WebKit renderer; check compatibility with your CSS and scripts
WeasyPrint Python publishing and document templates Free, open-source HTML-to-PDF project; validate CSS support for your design
Prince Books, reports, invoices, and complex paged layouts Focused on paged media, generated content, and publishing controls

There is no source-backed universal winner. Compare rendering basis, JavaScript requirements, layout controls, accessibility workflow, deployment needs, licensing, and support. Test your own pages because output depends on fonts, assets, scripts, and CSS.

HTML, CSS, fonts, and scripts pass through a renderer before becoming paginated PDF output.
HTML, CSS, fonts, and scripts pass through a renderer before becoming paginated PDF output.

Browser conversion with Puppeteer

Puppeteer’s page.pdf() API generates a PDF using the print CSS media type by default (documentation). Install it with:

npm install puppeteer

Complete Node.js example:

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({
    headless: true,
    args: ['--no-sandbox', '--disable-setuid-sandbox']
  });
  const page = await browser.newPage();
  await page.setViewport({ width: 1280, height: 900, deviceScaleFactor: 1 });
  await page.goto('https://example.com/report', {
    waitUntil: 'networkidle0',
    timeout: 90000
  });
  await page.emulateMediaType('print');
  await page.addStyleTag({ content: `
    @page { size: A4; margin: 18mm 14mm 20mm; }
    html { -webkit-print-color-adjust: exact; print-color-adjust: exact; }
    .page-break { break-before: page; }
  ` });
  await page.pdf({
    path: 'report.pdf',
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true,
    displayHeaderFooter: false,
    margin: { top: '18mm', right: '14mm', bottom: '20mm', left: '14mm' }
  });
  await browser.close();
})();

Important Puppeteer options

  • format selects a paper preset such as A4 or Letter. Use width and height for custom dimensions.
  • margin accepts CSS lengths for each edge.
  • printBackground: true preserves background colors and images.
  • preferCSSPageSize: true lets an @page rule control size instead of scaling to the API format.
  • displayHeaderFooter, headerTemplate, and footerTemplate add running page elements. Template CSS is limited; keep it self-contained.
  • pageRanges exports selected pages, for example 1-3 or 2,5.
  • landscape: true rotates the selected paper format.
  • omitBackground: true creates a transparent page background where supported by the renderer.

Browser conversion with Playwright

Playwright exposes a similar page.pdf() API. Its documentation covers paper formats, margins, backgrounds, outlines, and tagged output (Page API).

npm install -D playwright
npx playwright install chromium
import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
await page.goto('https://example.com/report', {
  waitUntil: 'networkidle',
  timeout: 90000
});
await page.emulateMedia({ media: 'print' });
await page.pdf({
  path: 'report.pdf',
  format: 'A4',
  printBackground: true,
  preferCSSPageSize: true,
  margin: { top: '18mm', right: '14mm', bottom: '20mm', left: '14mm' },
  tagged: true,
  outline: true
});
await browser.close();

The tagged option defaults to false. Enabling it does not prove that the result meets an accessibility standard; review the generated artifact with an accessibility process appropriate to your project.

Keep print rules in a dedicated stylesheet or an @media print block. Browser PDF generation is print-oriented, so screen-only assumptions often produce surprises.

@page {
  size: A4;
  margin: 16mm 14mm 20mm;
}

@media print {
  nav, .cookie-banner, .chat-widget, .interactive-controls {
    display: none !important;
  }

  a { color: inherit; text-decoration: none; }
  h1, h2, h3 { break-after: avoid; }
  table, figure, pre { break-inside: avoid; }
  .page-break { break-before: page; }
}

html {
  -webkit-print-color-adjust: exact;
  print-color-adjust: exact;
}

Puppeteer documents that PDF generation modifies colors for printing by default and points to -webkit-print-color-adjust when exact colors are needed. Even with that rule, inspect the output because printers, PDF viewers, and color profiles can differ.

Wait for dynamic content, fonts, and images

A navigation event only tells you that the page reached a lifecycle point. Data fetched after navigation, web fonts, lazy images, and chart libraries can still be unfinished. Use an application-specific readiness signal where possible.

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 90000 });
await page.waitForSelector('#report-ready', { timeout: 30000 });
await page.evaluate(() => document.fonts.ready);
await page.evaluate(async () => {
  const images = [...document.images];
  await Promise.all(images.map(img => img.complete
    ? Promise.resolve()
    : new Promise(resolve => {
        img.addEventListener('load', resolve, { once: true });
        img.addEventListener('error', resolve, { once: true });
      })));
});
await page.pdf({ path: 'ready.pdf', printBackground: true });

If your page has infinite scrolling, explicitly load the content needed for the document before export. For animations, disable transitions in print CSS or pause them in JavaScript so a capture does not land between states.

Command-line conversion with wkhtmltopdf

wkhtmltopdf describes a headless command-line HTML-to-PDF tool based on Qt WebKit. It runs without a display service.

wkhtmltopdf \
  --print-media-type \
  --enable-local-file-access \
  --margin-top 18mm \
  --margin-right 14mm \
  --margin-bottom 20mm \
  --margin-left 14mm \
  --javascript-delay 1500 \
  https://example.com/report report.pdf

Use --enable-local-file-access only when the input requires local assets, and restrict the files exposed to the process. The project page identifies the license as LGPLv3. Verify that your page’s JavaScript and CSS are compatible with its rendering engine before committing to it.

Python conversion with WeasyPrint

WeasyPrint is a free, open-source project for producing PDF documents from HTML. It is useful for server-rendered templates where you control the HTML and CSS.

pip install weasyprint
from weasyprint import HTML, CSS

HTML(url='https://example.com/report').write_pdf(
    'report.pdf',
    stylesheets=[CSS(string='''
        @page { size: A4; margin: 18mm 14mm 20mm; }
        @media print { .screen-only { display: none } }
    ''')]
)

For local templates, use an explicit base URL so relative stylesheets, images, and fonts resolve correctly:

HTML(string=html, base_url='/srv/app/templates').write_pdf('/tmp/report.pdf')

Publishing layouts with Prince

Prince’s user guide covers HTML, Markdown, and XML conversion, CSS styling, JavaScript, paged media, generated content, page numbering, headers, footers, list markers, and footnotes. A minimal command is:

prince invoice.html -o invoice.pdf

Prince is worth evaluating when the document itself is the product: books, reports, long invoices, or archival publications with strict page composition. Treat vendor documentation as feature documentation, not as an independent benchmark.

Handling assets, security, and private pages

  1. Use absolute HTTPS URLs or a controlled base URL for stylesheets, fonts, and images.
  2. Allow the renderer to reach only required hosts. Do not let untrusted HTML access internal metadata services or private network addresses.
  3. Pass authentication through a controlled browser context, cookie jar, or server-side template rather than embedding long-lived secrets in page markup.
  4. For deterministic output, pin font files and versions. A missing fallback font changes line wrapping and page count.
  5. Set navigation, script, and PDF generation timeouts. Return a useful error when a page never reaches its readiness condition.

Performance, reliability, and cost

Browser startup is usually the most expensive part of a small job. Reuse a browser process while isolating each request in a fresh context, cap concurrency, and close pages in a finally block. Cache immutable assets and avoid waiting for a global network-idle event when third-party analytics keep connections open; an explicit readiness element is more predictable.

A capture service can remove consent banners and overlays before generating the document.
A capture service can remove consent banners and overlays before generating the document.

Measure the output you actually need: conversion duration, PDF size, page count, memory, and failure rate by URL type. Keep a sample set containing long tables, web fonts, charts, right-to-left text, large images, and authenticated pages. Generate PDFs in a queue when a request can exceed normal HTTP timeouts.

Costs depend on your runtime, browser hosting, commercial licenses, storage, and traffic. The cited sources do not provide a controlled speed or total-cost comparison, so avoid assuming one engine is cheaper or faster without measuring your workload.

Quality checklist before shipping a PDF

  • Confirm paper size, orientation, margins, and intentional page breaks.
  • Check for clipped text, orphaned headings, split rows, and blank pages.
  • Verify fonts, images, SVGs, backgrounds, and chart labels.
  • Open links and inspect bookmarks or outlines where required.
  • Check reading order, selectable text, contrast, and tagged output if accessibility matters.
  • Compare a representative PDF against the source at multiple viewport widths.
  • Inspect the file in more than one PDF viewer before publishing.

Troubleshooting common errors

PDF is blank or missing data

Cause: the export ran before client-side rendering finished. Fix: wait for a page-specific selector, API completion flag, fonts, and images. Avoid relying only on a short sleep.

Colors or backgrounds disappeared

Cause: print color adjustment or background printing is disabled. Fix: set printBackground: true (or the equivalent), add print-color-adjust: exact, and verify the viewer.

Content is cut off at the page edge

Cause: fixed-width elements exceed the printable area. Fix: use responsive print widths, define margins with @page, and test wide tables separately.

Fonts wrap differently

Cause: the font was not loaded or is unavailable in the renderer. Fix: self-host or preload the font, await document.fonts.ready, and confirm the font file is reachable.

Images are absent

Cause: blocked URLs, lazy loading, CORS, or a premature capture. Fix: use reachable URLs, trigger lazy loading, wait for image completion, and log failed requests.

Cause: a stalled third-party request or page script. Fix: set a bounded timeout, block unnecessary resources, and use an application readiness condition instead of waiting forever for network idle.

Local files are rejected by wkhtmltopdf

Cause: local file access is disabled. Fix: enable it only for a controlled directory with --enable-local-file-access.

PDF accessibility is incomplete

Cause: a tagged-output option alone does not establish conformance. Fix: run an accessibility review of structure, reading order, headings, links, language, and contrast.

Or skip the browser setup

ScreenshotNeo can return a PDF from one API request, with the browser capture and cleanup handled for you. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Set the PDF output option for a PDF response. ScreenshotNeo can also handle full-page captures with lazy images loaded, custom CSS and JavaScript, waiting for selectors or network idle, cookies and headers, viewport and device presets, page ranges, paper size, margins, landscape mode, and other capture controls. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Should I use screen CSS or print CSS?

Use print CSS when the PDF is a document. Browser APIs default to print media; emulate screen media only when you intentionally want screen styling.

Can HTML-to-PDF conversion run without a display server?

Yes. Headless browser APIs and wkhtmltopdf are designed for server environments without a visible desktop.

Why does the PDF have more pages than the browser view?

PDF pagination applies paper dimensions, margins, print rules, font metrics, and break rules. A continuous screen layout has no equivalent page boundary.

Is a tagged PDF automatically accessible?

No. Tagged output can help, but accessibility still requires reviewing structure, reading order, labels, links, and other requirements.

Which tool should I choose for a JavaScript-heavy page?

Start with Puppeteer or Playwright because they automate a real browser. Confirm the page’s readiness state and inspect the resulting PDF.