ScreenshotNeo

BlogHTML to image & PDF

HTML to PDF Generator Tool: A Complete Developer Guide

Compare browser printing, WeasyPrint, wkhtmltopdf and hosted APIs, with runnable code, options, troubleshooting and production guidance.

By the ScreenshotNeo team1 October 20263 min read

An HTML to PDF generator converts HTML and CSS into a PDF file. Choose a browser-based renderer when your page depends on JavaScript or browser layout; choose a document engine such as WeasyPrint when your HTML can render without JavaScript; consider a command-line converter or hosted API when deployment simplicity matters.

For most web applications, start by rendering a representative document with Puppeteer, compare it against your required paper size and print styles, then add visual regression checks before upgrading the browser. The examples below cover browser printing, Python document rendering, command-line conversion and a managed screenshot-to-PDF option.

Choose an HTML to PDF approach

Approach Best fit Important limitation
Puppeteer Pages that need JavaScript, modern browser CSS, web fonts or application state You must operate a browser process and check output after browser upgrades
WeasyPrint Server-rendered HTML/CSS/SVG documents that do not require JavaScript Its product documentation says JavaScript is not executed
wkhtmltopdf Existing command-line integrations using its Qt WebKit renderer The project page is old; verify maintenance and compatibility before adopting it
Hosted HTML-to-PDF API Teams that want to avoid browser installation and operations Verify current pricing, data handling, regions, limits and migration options with each provider

Compare candidate tools on layout fidelity, print CSS, JavaScript execution, font and asset loading, paper controls, headers and footers, pagination, PDF metadata, deployment complexity and how you will detect rendering changes.

Generate a PDF with Puppeteer (Node.js)

Puppeteer’s page.pdf() uses print CSS media by default. If your design only has screen styles, call page.emulateMediaType('screen') before generating the file. The official Page.pdf documentation and PDFOptions reference list the available controls.

Install

npm install puppeteer

Complete runnable example

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({
    headless: true,
    // In a container, you may need these flags if the image has no sandbox setup.
    args: ['--no-sandbox', '--disable-setuid-sandbox']
  });

  try {
    const page = await browser.newPage();
    await page.setViewport({ width: 1280, height: 900, deviceScaleFactor: 1 });
    await page.goto('https://example.com/invoice/123', {
      waitUntil: 'networkidle0',
      timeout: 60000
    });

    // Use this only when the page is designed for screen media.
    // await page.emulateMediaType('screen');

    await page.pdf({
      path: 'invoice.pdf',
      format: 'A4',
      printBackground: true,
      preferCSSPageSize: true,
      margin: {
        top: '16mm',
        right: '14mm',
        bottom: '16mm',
        left: '14mm'
      },
      displayHeaderFooter: false
    });
  } finally {
    await browser.close();
  }
})();
@page {
  size: A4;
  margin: 16mm 14mm;
}

@media print {
  .screen-only { display: none !important; }
  .page-break { break-before: page; }
  .avoid-break { break-inside: avoid; }
  a { color: inherit; text-decoration: none; }
}

Useful Puppeteer PDF options

Option Purpose
format Preset such as A4 or Letter. The documented default is Letter.
width, height Custom paper dimensions when a preset is insufficient.
landscape Switches the page orientation.
margin Set top, right, bottom and left margins with CSS units.
printBackground Print background colors and images. The documented default is off.
preferCSSPageSize Use the CSS @page size instead of scaling the content to the requested format.
pageRanges Emit selected pages, for example 1-5, 8, 11-13.
scale Scale the rendered page when fitting content to paper.
displayHeaderFooter Enable Chromium header and footer templates.
headerTemplate, footerTemplate Insert HTML templates with supported page-number and date classes.
timeout Limit PDF generation time.

Authenticated pages and dynamic content

await page.setExtraHTTPHeaders({ Authorization: `Bearer ${process.env.API_TOKEN}` });
await page.setCookie({
  name: 'session',
  value: process.env.SESSION_COOKIE,
  domain: 'example.com',
  path: '/',
  httpOnly: true,
  secure: true
});
await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#report-ready', { timeout: 30000 });
await page.evaluate(() => document.fonts.ready);
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

Do not rely on a fixed delay alone for application state. Wait for a selector that means the report is complete, wait for fonts, and use a bounded timeout. If a chart is painted on a canvas, wait for the chart’s own ready signal before printing.

Generate a PDF with WeasyPrint (Python)

WeasyPrint converts HTML, CSS and SVG to PDF in Python and documents links, bookmarks and attachments. Its product site states that it does not execute JavaScript, so use it for server-rendered documents or precompute dynamic values before conversion.

Install and run

python -m pip install weasyprint
from weasyprint import HTML

HTML(string='''



  <meta charset="utf-8">
  <style>
    @page { size: A4; margin: 16mm; }
    body { font-family: sans-serif; }
    h1 { break-after: avoid; }
    .avoid-break { break-inside: avoid; }
  </style>

Use base_url when the document references relative images, stylesheets or fonts. For a file, pass HTML(filename='report.html').write_pdf('report.pdf'). Keep your WeasyPrint version pinned and compare representative PDFs after upgrades because its API reference warns that rendering can change between versions.

Generate a PDF with wkhtmltopdf

wkhtmltopdf describes a headless command-line HTML-to-PDF and image converter based on Qt WebKit. Existing deployments may still depend on it, but the surfaced project page is old, so verify supported operating systems, security posture and rendering compatibility before choosing it for a new system.

wkhtmltopdf --page-size A4 --margin-top 16mm --margin-right 14mm --margin-bottom 16mm --margin-left 14mm https://example.com report.pdf

HTML and CSS details that affect PDF output

Pagination

  • Define @page size and margins explicitly.
  • Use break-before, break-after and break-inside: avoid for sections, tables and cards.
  • Keep headings with the following content using break-after: avoid.
  • Test long words, overflowing code blocks and tables with many columns.

Fonts and assets

Make fonts and images reachable from the rendering process. Use absolute URLs or a correct base URL, wait for fonts in browser workflows, and check that private assets receive the required authentication headers. Missing fonts can change line wrapping and therefore every page break.

Colors and backgrounds

Browser printing commonly omits backgrounds unless you enable them. In Puppeteer, set printBackground: true. Also check whether your print stylesheet changes colors, shadows or visibility.

WeasyPrint documents clickable links, bookmarks and embedded attachments. Browser-generated PDFs may need additional post-processing if you require a specific outline structure, metadata policy or attachment behavior.

Or skip the browser setup

ScreenshotNeo provides an HTML-to-PDF capture path through its API. See the ScreenshotNeo documentation for the current parameters. This one-call example captures a URL as a PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o document.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("document.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('document.pdf', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo removes cookie and consent banners, newsletter popups and chat widgets before the capture. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and whether it was billed. It also offers PDF paper size, margins, landscape mode and page ranges, plus custom CSS and JavaScript, waits, headers, cookies, user agents, authorization, timezone and geolocation. Its MCP server lets AI agents call take_screenshot, get_page_info and capture_pdf. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Production checklist

  1. Pin the renderer and browser versions.
  2. Store a small fixture set containing long text, tables, images, non-Latin characters and authenticated content.
  3. Compare rendered PDFs after dependency upgrades.
  4. Wait for a real readiness selector, fonts and critical images.
  5. Set navigation and conversion timeouts and cancel orphaned browser jobs.
  6. Limit URL access to prevent server-side request forgery when users supply URLs.
  7. Keep credentials out of generated PDFs, logs and query strings where possible.
  8. Define retention and deletion rules for source HTML and generated files.
  9. Measure queue time, render time, output size and failure reason by document type.

Performance, reliability and cost

Browser startup is expensive compared with reusing a controlled browser process, but long-lived processes must be isolated and recycled to avoid leaks. Reuse pages carefully, clear cookies and storage between jobs, and cap concurrency so CPU and memory pressure do not create timeouts.

Reduce output size by using appropriately sized images, avoiding unnecessary background assets and selecting the required page range. Cache deterministic documents when their inputs have not changed. For a hosted API, include network latency, transfer size, retry behavior, retention, geographic processing and per-document pricing in your cost model. ScreenshotNeo bills only clean shots; cache hits, bot checks, blank pages, timeouts and failed loads are not billed, and each response includes X-Page-Verdict and X-Billed headers.

Troubleshooting

Symptom Likely cause Fix
PDF uses the wrong layout Print media rules are active Review @media print; call emulateMediaType('screen') only when screen CSS is intended.
Background colors are missing Background printing is disabled Set Puppeteer’s printBackground: true and verify print CSS.
Images or CSS are missing Relative URLs or private assets cannot be resolved Set a base URL, use absolute URLs and provide authentication headers or cookies.
Fonts change page breaks Fonts were not loaded before printing Wait for document.fonts.ready; verify font responses and family names.
Charts are blank Canvas or application data was not ready Wait for an application readiness selector or chart completion event.
Content is cut off Fixed heights, overflow rules or unsuitable margins Remove restrictive heights, inspect overflow and tune @page margins and breaks.
WeasyPrint output lacks dynamic content JavaScript is not executed Render the data into HTML first or use a browser-based renderer.
Navigation times out Slow dependency, blocked request or page waiting forever Set a bounded timeout, wait for a specific selector and inspect failed network requests.
Jobs fail only in containers Missing browser libraries or sandbox permissions Use a compatible image, install required dependencies and configure sandboxing according to your environment.

FAQ

Which tool should I use for a React or Vue page?

Use a browser renderer such as Puppeteer when the final content is assembled by JavaScript. Wait for a deterministic ready signal before calling page.pdf().

Can WeasyPrint render JavaScript?

No. Its product documentation says it does not execute JavaScript.

How do I select only some pages?

Puppeteer’s pageRanges accepts ranges such as 1-5, 8, 11-13.

Why did an upgrade change the PDF?

Rendering engines, browser versions, fonts and pagination behavior can change. Pin versions and run visual comparisons on representative documents.

Can I generate a PDF without installing Chromium?

Yes. Use a document engine such as WeasyPrint for non-JavaScript HTML, a command-line converter already approved in your environment, or a hosted service such as ScreenshotNeo.