ScreenshotNeo

BlogHTML to image & PDF

HTML-to-PDF Examples for Developers

Runnable Puppeteer, Playwright, Python, cURL and Prince examples for reliable HTML-to-PDF generation, print CSS, page breaks and troubleshooting.

By the ScreenshotNeo team1 October 20268 min read

To generate a PDF from HTML, render the document in a browser or print-focused engine, wait for its content and fonts to load, then export it. Puppeteer and Playwright are the simplest choices for JavaScript applications; Prince is designed for documents that need advanced CSS paged-media features such as generated page numbers, headers and footers.

This guide starts with runnable examples, then covers print CSS, page breaks, media selection, assets, security, reliability, performance, costs and common failures.

1. Choose an HTML-to-PDF approach

Tool Rendering model Output Best fit Important detail
Puppeteer Chromium browser automation Writes a PDF file A straightforward Node.js pipeline Uses print CSS by default
Playwright Chromium browser automation Returns a PDF buffer Projects already using Playwright contexts and pages PDF export is Chromium-only
Prince XML Dedicated HTML/XML-to-PDF engine Engine-generated PDF Print-heavy reports, books and composed documents Investigate commercial licensing separately

Use Puppeteer when you need the smallest Node.js implementation. Use Playwright when the rest of your application already uses its browser, context and page APIs. Investigate Prince when CSS paged-media controls and document composition matter more than browser automation.

2. Generate a PDF with Puppeteer

Puppeteer’s basic sequence is: launch Chromium, create a page, navigate to the URL, call page.pdf(), and close the browser. Its PDF operation waits for fonts to load by default.

Install

npm install puppeteer

URL to PDF

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle2' });
  await page.pdf({
    path: 'example.pdf',
    format: 'A4',
    printBackground: true,
  });
} finally {
  await browser.close();
}

waitUntil: 'networkidle2' waits for a quiet network, but it is not a guarantee that application data has finished rendering. For dashboards and other client-rendered pages, wait for a selector that proves the content is ready.

await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-report-ready]', { timeout: 30000 });
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

Use screen CSS instead of print CSS

page.pdf() generates the page with the print CSS media type. If the page is designed for screens, switch media before exporting:

await page.emulateMediaType('screen');
await page.pdf({
  path: 'screen-styled.pdf',
  printBackground: true,
});

For exact colors, use -webkit-print-color-adjust: exact in the print stylesheet. This affects color adjustment; it does not replace printBackground: true.

Control paper, margins and orientation

await page.pdf({
  path: 'invoice.pdf',
  format: 'A4',
  landscape: false,
  printBackground: true,
  margin: {
    top: '18mm',
    right: '14mm',
    bottom: '18mm',
    left: '14mm',
  },
  displayHeaderFooter: false,
});

Use format for standard paper sizes or explicit width and height values for custom pages. Keep units explicit, such as mm, in or px.

3. Generate a PDF with Playwright

Playwright’s page.pdf() returns a PDF buffer and also uses print CSS by default. PDF generation is available through Chromium.

Install and export a buffer

npm install playwright
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';

const browser = await chromium.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle' });
  const pdf = await page.pdf({
    format: 'A4',
    printBackground: true,
  });
  await writeFile('example.pdf', pdf);
} finally {
  await browser.close();
}

Use screen styling and wait for application state

await page.emulateMedia({ media: 'screen' });
await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.locator('[data-report-ready]').waitFor({ state: 'visible' });
const pdf = await page.pdf({
  format: 'A4',
  printBackground: true,
  preferCSSPageSize: true,
});

preferCSSPageSize: true lets a document’s @page rule control the page size when that is part of your layout design.

4. Print CSS, page breaks and document layout

Put PDF-specific rules in a print stylesheet or an @media print block. A reliable baseline is:

@page {
  size: A4;
  margin: 18mm 14mm;
}

@media print {
  * {
    -webkit-print-color-adjust: exact;
    print-color-adjust: exact;
  }

  .page-break {
    break-before: page;
  }

  .avoid-break {
    break-inside: avoid;
  }

  thead {
    display: table-header-group;
  }

  a {
    color: inherit;
    text-decoration: none;
  }
}

Page breaks

  • break-before: page starts a new page before an element.
  • break-after: page starts a new page after an element.
  • break-inside: avoid helps keep cards, figures and table rows together.
  • Legacy page-break-before, page-break-after and page-break-inside remain useful for older stylesheets.

Browsers may still move content when an unbreakable block is taller than a page. Split very large code blocks, images or tables into intentional sections.

Headers, footers and page numbers

Browser PDF headers and footers are controlled by the engine’s header/footer options and have limited layout capabilities. For generated page numbers, running headers and complex page furniture, Prince’s CSS paged-media workflow is the option to investigate.

5. Convert HTML with Prince XML

Prince converts HTML and XML documents into PDF by applying CSS. Its documentation covers CSS-generated content for page numbers, headers and footers, plus networking and server integration.

prince invoice.html -o invoice.pdf

A typical paged document uses @page, generated content and named strings. Check the Prince documentation for the exact syntax and deployment model that match your version.

6. HTML-to-PDF options you will use often

Requirement Typical setting or technique
Background colors and images Enable printBackground (Puppeteer/Playwright) and use print color adjustment CSS.
Screen-designed page Call emulateMediaType('screen') or emulateMedia({ media: 'screen' }).
Print-designed page Use the default print media and define @page.
Custom paper Set explicit width and height or a CSS @page size.
Landscape report Set landscape: true or an appropriate CSS page size.
Dynamic data Navigate, wait for a readiness selector, then export.
Fonts Load web fonts before export; PDF generation waits for fonts in Puppeteer.
Images Use absolute URLs or data URLs, wait for lazy content, and verify access from the rendering environment.
Links Keep URLs absolute when the PDF will be opened away from the source site.
Large tables Repeat table headers and allow rows to break only when necessary.

7. A production-ready capture checklist

  1. Pin the browser or rendering-engine version used in production.
  2. Set an explicit navigation timeout and an export timeout.
  3. Wait for a semantic readiness selector, not only network idleness.
  4. Set the media type deliberately.
  5. Define paper size and margins in one place.
  6. Enable backgrounds when the design depends on them.
  7. Load fonts and images from reachable, authenticated locations.
  8. Use print CSS to control page breaks and avoid splitting important blocks.
  9. Close pages and browsers in a finally block.
  10. Store the source URL, rendering version and relevant options with the generated PDF for debugging.

8. Troubleshooting common failures

Symptom Likely cause Fix
PDF is blank Navigation failed, content is client-rendered, or the page is blocked. Capture navigation errors, wait for a readiness selector, and verify the URL from the worker environment.
Styles look wrong The page uses screen CSS while the exporter uses print CSS. Switch to screen media or add an intentional print stylesheet.
Colors are missing Background printing is disabled or the browser adjusts colors. Enable printBackground and use print-color-adjust.
Web fonts fall back Fonts have not loaded or are blocked by CORS/authentication. Wait for document.fonts.ready, fix access headers, and keep font URLs reachable.
Images are absent Lazy loading has not been triggered, or URLs cannot be fetched. Scroll or trigger loading before export and inspect image requests.
Content is cut off Fixed heights, overflow rules or an oversized unbreakable element. Remove restrictive heights in print CSS and split large blocks.
Unexpected blank pages Conflicting page-break rules, margins or an empty break element. Inspect computed print styles and remove duplicate break rules.
PDF export hangs Long-polling, analytics or never-ending requests prevent an idle condition. Wait for a selector instead of network idle and block nonessential requests.
Chromium will not start Missing browser binary or OS libraries in the deployment image. Install the required browser during the build and verify the container dependencies.
Protected page returns a challenge Bot protection or authentication prevents normal rendering. Use an authorized session, custom headers/cookies, or a service designed to report failed captures clearly.

9. Performance, reliability and cost

Performance

Browser startup is expensive relative to reusing a process. Keep one browser process per worker and create isolated pages or contexts per job. Reuse pages only when you can reliably clear cookies, storage and injected state. Block analytics, advertising and other resources that are irrelevant to the document, but do not block fonts, stylesheets or images required for the output.

Reliability

Use bounded timeouts, retries for transient network failures, and idempotent job identifiers. Record whether failure occurred during navigation, readiness waiting or PDF encoding. Validate the resulting file before publishing it. For repeatable output, pin dependencies and avoid time-dependent content where possible.

Cost

Self-hosted browser rendering consumes compute, memory and operational time. A hosted API trades browser deployment for per-capture pricing. Estimate the number of successful documents, retries and long-running pages; cache identical inputs when the source and rendering options have not changed.

10. Or skip the browser setup

ScreenshotNeo can return a PDF from one GET request. See the ScreenshotNeo API documentation for the available options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o page.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const pdf = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('page.pdf', pdf));

ScreenshotNeo accepts the cookie or consent banner before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and start with 1,000 shots per month at no charge.

11. FAQ

Does PDF generation use print or screen CSS?

Puppeteer and Playwright use print CSS by default. Explicitly emulate screen media when the page’s layout is intended for screens.

Which tool should I use for a Node.js service?

Start with Puppeteer for a minimal pipeline. Choose Playwright when your service already uses Playwright’s browser and context APIs.

Can a browser PDF contain page numbers?

Basic browser header/footer support is limited. Prince is the option to investigate when generated page numbers, running headers and other paged-media features are central requirements.

Why does network idle not guarantee complete output?

Applications can render after the network becomes quiet, and some connections never become idle. Wait for a selector or application-specific readiness signal.

Can I generate PDFs without installing Chromium?

Yes. A hosted service such as ScreenshotNeo handles the capture infrastructure and returns the PDF response from its API.