ScreenshotNeo

BlogHTML to image & PDF

How to Generate Editable PDFs from HTML with Puppeteer

Generate print-ready PDFs from HTML with Puppeteer, control layout and assets, and understand when you need real fillable form fields.

By the ScreenshotNeo team1 October 202611 min read

Direct answer: use Puppeteer’s page.setContent() (or page.goto()) to load HTML, wait for its assets to finish loading, then call page.pdf(). Puppeteer renders with print CSS by default and returns PDF bytes. This produces a print-ready, selectable-text PDF. It does not, according to the documented API, turn HTML form controls into interactive AcroForm fields automatically.

The word editable can describe three different outcomes:

  • Editable source: you can change the HTML/CSS and regenerate the PDF.
  • Selectable/searchable text: the PDF contains text rather than a screenshot.
  • Fillable PDF: recipients can type into actual PDF form fields, check boxes, or choose options in a viewer.

Puppeteer handles the first two well. Treat the third as a separate PDF form-authoring or post-processing step and test the resulting file in the viewers your recipients use.

1. Install Puppeteer and create a PDF

Start a Node.js project and install Puppeteer. The package downloads a compatible browser during installation unless your setup is configured to use an existing browser.

mkdir html-pdf
cd html-pdf
npm init -y
npm install puppeteer

Create generate-pdf.mjs:

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const html = `<!doctype html>
<html>
<head>
  <meta charset="utf-8">
  <title>Invoice</title>
  <style>
    @page { size: A4; margin: 18mm 16mm 20mm; }
    * { box-sizing: border-box; }
    body {
      margin: 0;
      color: #202124;
      font-family: Arial, sans-serif;
      font-size: 11pt;
      line-height: 1.45;
    }
    h1 { margin: 0 0 8mm; font-size: 24pt; }
    .meta { color: #5f6368; margin-bottom: 12mm; }
    table { width: 100%; border-collapse: collapse; }
    th, td { padding: 3mm; border-bottom: 1px solid #d9d9d9; text-align: left; }
    th { background: #f1f3f4; }
    .total { margin-top: 10mm; text-align: right; font-size: 14pt; font-weight: 700; }
    .avoid-break { break-inside: avoid; }
  </style>
</head>
<body>
  <h1>Invoice 1042</h1>
  <p class="meta">Issued 2026-10-01 · Acme Example Ltd.</p>
  <table>
    <thead><tr><th>Description</th><th>Amount</th></tr></thead>
    <tbody>
      <tr><td>Design and implementation</td><td>$1,200.00</td></tr>
      <tr><td>Support</td><td>$300.00</td></tr>
    </tbody>
  </table>
  <p class="total">Total: $1,500.00</p>
</body>
</html>`;

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setContent(html, { waitUntil: 'networkidle0' });
  const pdf = await page.pdf({
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true,
  });
  await writeFile('invoice.pdf', pdf);
} finally {
  await browser.close();
}

Run it with:

node generate-pdf.mjs

This follows Puppeteer’s documented PDF path: PDF generation guide, Page.pdf(), and Page.setContent().

2. Load a URL instead of an HTML string

For an existing page, use page.goto(). Choose a wait condition that matches the page rather than assuming that the initial HTML response means the page is ready.

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com/report', {
    waitUntil: 'networkidle0',
    timeout: 45_000,
  });
  await page.pdf({
    path: 'report.pdf',
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true,
  });
} finally {
  await browser.close();
}

networkidle0 waits for there to be no active network connections. Pages with analytics, polling, advertisements, or web sockets may never reach that state. In those cases, use domcontentloaded or load, then wait for a specific selector or application-ready signal:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector('[data-report-ready]', { timeout: 20_000 });

3. Control page size, margins, orientation, and ranges

The PDFOptions interface contains the layout controls you will use most often.

Option Purpose Notes
format Named paper size such as A4 or Letter The documented default is Letter.
width, height Custom paper dimensions Use CSS units such as mm, in, or px.
landscape Rotates the page orientation Set true for wide tables or slides.
margin Top, right, bottom, and left margins Use strings with units or numeric values.
printBackground Print background colors and images Defaults to false; enable it for designed documents.
preferCSSPageSize Lets CSS @page dimensions win Useful when the document owns its print specification.
pageRanges Prints selected pages Use ranges such as 1-3 or 2,4.
scale Scales the rendered page Valid range is 0.1–2; default is 1.

Use either CSS or PDFOptions as the source of truth for dimensions. If you set preferCSSPageSize: true, define the intended paper size in CSS:

@page {
  size: A4 portrait;
  margin: 12mm 15mm;
}

@media print {
  .screen-only { display: none !important; }
  .avoid-break { break-inside: avoid; }
  .page-break { break-before: page; }
}

For a landscape report with a selected range:

await page.pdf({
  path: 'wide-report.pdf',
  format: 'A4',
  landscape: true,
  pageRanges: '1-3,5',
  margin: { top: '12mm', right: '10mm', bottom: '14mm', left: '10mm' },
  printBackground: true,
});

4. Print CSS versus screen CSS

Puppeteer uses print media by default. This means rules inside @media print apply and screen-only styles may not. If the PDF should match the screen layout, switch media before generating it:

await page.emulateMediaType('screen');
const pdf = await page.pdf({
  format: 'A4',
  printBackground: true,
});

For a print-specific document, leave the default print media in place and define page breaks explicitly. Long cards, table rows, and headings often need break-inside: avoid, break-before, or break-after. Always inspect multi-page output because a rule that looks correct on one page can create a large blank area when the remaining space is too small.

Chromium adjusts colors for printing. When exact brand colors matter, add -webkit-print-color-adjust: exact to the relevant element and still verify the output in your target PDF viewer and printer.

5. Make fonts, images, and scripts ready before printing

HTML-to-PDF failures are frequently readiness failures. A page can have its DOM but still be waiting for web fonts, lazy images, client-side data, or chart rendering.

Wait for application state

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#report-complete', { timeout: 30_000 });
await page.evaluate(() => document.fonts.ready);
await page.waitForFunction(() => {
  return [...document.images].every((image) => image.complete);
}, { timeout: 30_000 });

The documented PDF options set waitForFonts to true by default. That helps with font readiness, but it does not guarantee that your application’s data fetches, charts, or lazy-loading code has finished. Give every external dependency a bounded timeout and provide a deterministic ready marker.

Make assets resolvable

  • Use absolute HTTPS URLs for fonts, images, and stylesheets when rendering supplied HTML.
  • Ensure the browser process can reach private assets, or inline them as data URLs.
  • Check certificate, DNS, and authentication requirements in the environment where Chromium runs.
  • Use page.setExtraHTTPHeaders() or cookies when the page requires authentication.

Wait for charts and client-side rendering

await page.waitForFunction(() => {
  const chart = document.querySelector('#revenue-chart');
  return chart?.dataset.rendered === 'true';
}, { timeout: 20_000 });

A fixed delay such as await new Promise(r => setTimeout(r, 1000)) can be a fallback, but a selector or application signal is more reliable because rendering time changes with load and data size.

6. Add headers, footers, and page numbers

Set displayHeaderFooter: true and provide HTML templates. Puppeteer replaces supported classes such as pageNumber and totalPages when it renders the document.

await page.pdf({
  path: 'manual.pdf',
  format: 'A4',
  printBackground: true,
  displayHeaderFooter: true,
  headerTemplate: '<div style="font-size:9px;width:100%;text-align:center;color:#666">Acme Manual</div>',
  footerTemplate: '<div style="font-size:9px;width:100%;text-align:center;color:#666">Page <span class="pageNumber"></span> of <span class="totalPages"></span></div>',
  margin: { top: '22mm', bottom: '20mm' },
});

Header and footer templates have limited layout context. Keep them self-contained, use inline styles, and reserve enough top and bottom margin so body content does not overlap them.

7. What “editable PDF” means in practice

Source-editable workflow

Keep your HTML, CSS, and data separate from the PDF output. Regenerate whenever the source changes. This is the simplest interpretation of editable and gives you full control over typography, page breaks, and content.

Selectable and searchable text

Puppeteer renders real text through Chromium, so normal text can usually be selected and searched in a PDF viewer. Avoid converting all content to a canvas or a single raster image if text accessibility matters. Validate characters, ligatures, right-to-left scripts, and fallback fonts with representative documents.

Interactive fillable fields

HTML elements such as <input>, <select>, and <textarea> are rendered as their visual appearance in the print output. The reviewed Puppeteer Page.pdf() and PDFOptions documentation describes printing and layout; it does not promise conversion of those controls into AcroForm widgets.

If recipients must type into the delivered PDF, use a dedicated PDF form-field generation or post-processing step after rendering. Then test:

  • field names and tab order;
  • required fields and validation behavior;
  • font embedding and non-Latin input;
  • desktop viewers, browser viewers, and mobile viewers used by recipients;
  • printing and flattening behavior.

8. A reusable production function

import puppeteer from 'puppeteer';

export async function htmlToPdf({ html, outputPath, landscape = false }) {
  const browser = await puppeteer.launch({
    // Add launch arguments only when required by your container or host.
  });

  try {
    const page = await browser.newPage();
    await page.setContent(html, {
      waitUntil: 'networkidle0',
      timeout: 45_000,
    });
    await page.evaluate(() => document.fonts.ready);

    await page.pdf({
      path: outputPath,
      format: 'A4',
      landscape,
      printBackground: true,
      preferCSSPageSize: true,
      waitForFonts: true,
      margin: {
        top: '16mm',
        right: '14mm',
        bottom: '18mm',
        left: '14mm',
      },
    });
  } finally {
    await browser.close();
  }
}

await htmlToPdf({
  html: '<html><body><h1>Report</h1></body></html>',
  outputPath: 'report.pdf',
});

In a server, reuse a browser process when appropriate and create a fresh page per job. Always close pages and browsers in finally blocks so failed jobs do not accumulate Chromium processes.

9. Troubleshooting

Symptom Likely cause Fix
Blank or incomplete PDF Client-side content had not rendered Wait for a ready selector, application signal, fonts, and images before page.pdf().
Missing background colors printBackground is false Set printBackground: true and check print color adjustment.
Wrong paper size CSS and PDFOptions disagree Choose one source of truth; use preferCSSPageSize: true when CSS owns the size.
Text uses a fallback font Font request failed or was still loading Use reachable URLs or inline fonts, wait for document.fonts.ready, and inspect network errors.
Images are missing Lazy loading or inaccessible asset URL Scroll or trigger loading, wait for every image to complete, and verify the browser can access the asset.
networkidle0 never resolves Analytics, polling, sockets, or ads keep connections open Use domcontentloaded or load plus a specific readiness condition.
Content overlaps the footer Footer enabled without enough bottom margin Increase the PDF bottom margin and simplify the footer template.
Unexpected page breaks Large blocks cannot fit in remaining space Use break-inside: avoid selectively and inspect the resulting whitespace.
Timeout during PDF generation Slow resources or an overloaded browser Set a suitable timeout, reduce external dependencies, and limit concurrent jobs.
Chromium fails to launch in a container Missing system libraries or incompatible sandbox settings Use Puppeteer’s supported browser installation and follow your container’s Chromium requirements; do not assume an arbitrary system browser is interchangeable.
HTML inputs are not fillable Print rendering is not form-field authoring Add a dedicated AcroForm creation or post-processing stage.

10. Compatibility, performance, and reliability

Pin compatible browser versions

Puppeteer’s browser mapping changes. Its support table currently documents Puppeteer 25.12.0 with Chrome for Testing 154.0.8037.57 and Firefox 156.0.1, and notes that Chrome for Testing has been downloaded and used since Puppeteer v20. Check the official support table for the version you install instead of assuming that any system Chrome build is equivalent.

Keep jobs bounded

  • Set navigation, selector, and PDF timeouts.
  • Reject unbounded pages with endless polling or oversized data.
  • Cap concurrency so each Chromium process has enough memory.
  • Close pages after every job and browsers during graceful shutdown.
  • Record the URL, Puppeteer version, browser version, options, and failure stage for diagnosis.

Reduce rendering cost

Reuse a browser where safe, but isolate each document in a new page. Inline small critical CSS, avoid unnecessary third-party resources, and wait for a deterministic application-ready signal instead of a long fixed sleep. Large images, complex SVG, very long tables, and high concurrency increase memory and latency.

Validate output

For important documents, inspect page count, file size, text extraction, expected headings, and representative visual pages. Test portrait and landscape variants, long and short data sets, missing optional fields, and the fonts and locales your users actually receive.

11. Or skip the browser setup

If you need a clean capture of a web page rather than maintaining Chromium yourself, ScreenshotNeo provides a single HTTP endpoint. Its cleanup steps accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents.

See the ScreenshotNeo API documentation for request options and PDF capture details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every feature is available on every plan. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

12. Cost and service choices

Self-hosting Puppeteer means paying for your own compute, browser maintenance, storage, and operational work. It is a good fit when you need complete control over HTML execution, private network access, or a separate fillable-form pipeline. A hosted capture API can be simpler when the input is a public URL and you want request-level handling for popups, failed loads, caching, and asynchronous jobs.

For ScreenshotNeo, only clean shots are billed. Plans are Free (1,000 shots/month), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free.

13. FAQ

Does Puppeteer create a PDF from a local HTML file?

Yes. Read the file, call page.setContent() with its contents, wait for any referenced assets, and call page.pdf(). For local relative assets, ensure their URLs resolve in the browser context.

Can I generate only selected pages?

Yes. Pass a pageRanges value such as 1-3 or 2,4 in the PDF options.

Should I use networkidle0 for every page?

No. It is useful for pages that become quiet, but polling and analytics can prevent it from completing. A page-specific ready signal is often more reliable.

How do I make a PDF match screen styling?

Call page.emulateMediaType('screen') before page.pdf(), then enable printBackground if the design depends on backgrounds.

Can I use Puppeteer’s PDF as a fillable government or tax form?

Not without an additional form-field authoring step. Render the visual document with Puppeteer, create the required PDF widgets separately, and validate the final file in the official viewer requirements for that form.

Where should I check browser compatibility?

Use Puppeteer’s supported browsers table for the installed Puppeteer version and its corresponding browser builds.