ScreenshotNeo

BlogHTML to image & PDF

How to Convert HTML to PDF with an npm Library

Convert HTML to PDF in Node.js with Puppeteer, configure print CSS, margins and page sizes, troubleshoot rendering, and compare managed alternatives.

By the ScreenshotNeo team1 October 20268 min read

Short answer: For browser-accurate HTML and CSS, install Puppeteer, launch a browser, load the HTML, call page.pdf(), and close the browser. Puppeteer’s PDF API supports paper formats, custom dimensions, margins, print or screen media, backgrounds, CSS page sizing, page ranges and more.

This guide shows a complete Node.js implementation, explains the options that affect layout, and covers the failure modes that commonly produce blank, clipped or incorrectly paginated PDFs.

1. Install an npm library

Puppeteer’s PDF guide documents the browser-rendering workflow used here. Install it in a new project:

mkdir html-to-pdf
cd html-to-pdf
npm init -y
npm install puppeteer

Puppeteer downloads a compatible browser during installation. Your deployment must be able to start that browser and provide the libraries it needs.

2. Convert an HTML string to PDF

This runnable ES module loads an HTML string, waits for the page to become ready, and returns PDF bytes. It writes those bytes to output.pdf.

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const html = `
<!doctype html>
<html>
  <head>
    <meta charset="utf-8">
    <title>Invoice</title>
    <style>
      @page { size: A4; margin: 18mm 16mm; }
      * { box-sizing: border-box; }
      body {
        margin: 0;
        color: #172033;
        font: 14px/1.5 Arial, sans-serif;
      }
      h1 { margin: 0 0 12px; }
      .avoid-break { break-inside: avoid; }
      .page-break { break-before: page; }
    </style>
  </head>
  <body>
    <h1>Invoice 1001</h1>
    <div class="avoid-break">Rendered from HTML with Puppeteer.</div>
  </body>
</html>`;

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setContent(html, { waitUntil: 'networkidle0' });
  const pdf = await page.pdf({
    path: 'output.pdf',
    format: 'A4',
    printBackground: true
  });
  console.log(`Wrote ${pdf.length} bytes`);
} finally {
  await browser.close();
}

The API returns PDF bytes (a Uint8Array in current documentation), so you can upload the result to object storage or return it from an HTTP response instead of writing a file.

3. Convert a URL to PDF

For a page hosted by your application or another site, navigate to its URL before calling page.pdf():

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com/invoice/1001', {
    waitUntil: 'networkidle2',
    timeout: 30_000
  });
  await page.pdf({
    path: 'invoice-1001.pdf',
    format: 'A4',
    printBackground: true
  });
} finally {
  await browser.close();
}

The official guide uses networkidle2 in its URL example. Treat that as an example choice, not a universal readiness rule. Dynamic applications may need a selector, an explicit delay, or an application-specific promise.

4. Choose the rendering mode and page layout

page.pdf() uses the print CSS media type by default. To render the styles intended for a screen, select it before generating the PDF:

await page.emulateMediaType('screen');
await page.pdf({
  path: 'screen-styled.pdf',
  printBackground: true
});

Use print styles for documents and reports. Use screen media when the page’s screen layout is the one you need to preserve.

Backgrounds and exact colors

printBackground defaults to false. Set it to true when colored panels, gradients or background images belong in the output. Chromium can adjust colors for printing; CSS -webkit-print-color-adjust controls whether exact colors are forced:

html {
  -webkit-print-color-adjust: exact;
  print-color-adjust: exact;
}

Paper size, orientation and margins

await page.pdf({
  path: 'landscape-letter.pdf',
  format: 'Letter',
  landscape: true,
  margin: {
    top: '12mm',
    right: '12mm',
    bottom: '14mm',
    left: '12mm'
  },
  printBackground: true
});

The documented options include format, width, height, landscape, margin, scale, pageRanges, timeout and preferCSSPageSize. If format is present, it takes priority over width and height.

When your stylesheet defines @page { size: ... }, set preferCSSPageSize: true to give that CSS size priority:

await page.pdf({
  path: 'css-sized.pdf',
  preferCSSPageSize: true,
  printBackground: true
});

Page ranges and scaling

await page.pdf({
  path: 'pages-2-and-4.pdf',
  format: 'A4',
  pageRanges: '2,4',
  scale: 0.95,
  printBackground: true
});

Scaling changes the rendered content size and can alter pagination. Fix widths, margins and page breaks before using scale to compensate for an overflow problem.

5. Wait for fonts, images and application data

Puppeteer’s current guide and API documentation say PDF generation waits for fonts by default. External images, web fonts and client-rendered data still need a readiness strategy that matches your page.

await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#report-ready', { timeout: 30_000 });
await page.evaluate(() => document.fonts.ready);
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

Have the application add #report-ready only after its data and critical images are rendered. This is more deterministic than assuming that all network activity has stopped.

6. Avoid broken pagination

PDF pages are created from the final layout. These CSS rules help keep cards, table rows and headings together:

.card,
.table-row,
.avoid-break {
  break-inside: avoid;
}

h2, h3 {
  break-after: avoid;
}

.page-break {
  break-before: page;
}

@page {
  size: A4;
  margin: 16mm;
}

Very tall elements cannot fit on one page even with break-inside: avoid. Reduce their height, allow a break, or split the content deliberately.

7. Return the PDF from an HTTP endpoint

Because page.pdf() returns bytes, an Express route can send them directly:

import express from 'express';
import puppeteer from 'puppeteer';

const app = express();
const browser = await puppeteer.launch();

app.get('/invoices/:id.pdf', async (req, res) => {
  const page = await browser.newPage();
  try {
    await page.goto(`https://app.example.com/invoices/${req.params.id}`, {
      waitUntil: 'domcontentloaded',
      timeout: 30_000
    });
    await page.waitForSelector('#report-ready');
    const pdf = await page.pdf({
      format: 'A4',
      printBackground: true,
      preferCSSPageSize: true
    });
    res.type('application/pdf').send(Buffer.from(pdf));
  } catch (error) {
    res.status(500).json({ error: 'PDF generation failed' });
  } finally {
    await page.close();
  }
});

app.listen(3000);

Reuse a browser process for multiple requests, but create and close a page per job. Put a queue and concurrency limit in front of generation when traffic can spike.

8. Puppeteer versus other npm packages

Route How it works Choose it when
Puppeteer Runs a browser and converts the rendered page with page.pdf(). You need HTML, CSS, JavaScript, web fonts and browser layout fidelity.
puppeteer-html-pdf A wrapper around browser-based PDF generation. You have verified its current Node compatibility, dependencies and maintenance for your project.
html-pdf-node A package listing that accepts a URL or HTML content. You have checked its current behavior and operational requirements.
PDFKit Programmatically constructs PDF documents. You want to build a document through drawing and document APIs rather than render arbitrary HTML/CSS.

A wrapper’s existence does not establish better performance, security or maintenance. Compare browser requirements, HTML/CSS fidelity, print-media control, page sizing and operational support before adopting one.

9. Troubleshooting

Blank or partially rendered PDF

  • Cause: The page was captured before client-side data or assets finished loading.
  • Fix: Wait for an application-specific selector, then wait for document.fonts.ready and any required image state.

Background colors are missing

  • Cause: printBackground defaults to false.
  • Fix: Set printBackground: true and use print color adjustment CSS when exact colors matter.

The PDF uses the wrong layout

  • Cause: Print media is the default, while the page’s important rules target screen.
  • Fix: Call page.emulateMediaType('screen'), or add the required print styles.

Content is clipped or unexpectedly split

  • Cause: Fixed heights, oversized elements, margins or conflicting @page rules.
  • Fix: Inspect the print layout, remove rigid heights, adjust margins, use break rules and decide whether CSS or the PDF format controls page size.
  • Cause: A page keeps long-lived connections open, or the chosen readiness condition never occurs.
  • Fix: Use a suitable waitUntil value, then wait for a concrete selector. Increase the documented 30-second PDF timeout only when the workload requires it.

Browser fails to start in production

  • Cause: The runtime lacks Puppeteer’s browser binary or required system libraries.
  • Fix: Install the browser and OS dependencies in the image, verify the executable path, and run a small startup check during deployment.

10. Performance, reliability and cost

  • Keep one browser process alive and create isolated pages for jobs; launching a new browser for every request adds startup work.
  • Limit concurrent pages so memory use remains predictable. Queue excess work and apply a job timeout.
  • Reuse templates and avoid loading resources that do not belong in the document. Do not block assets required for the final layout.
  • Use deterministic readiness selectors instead of waiting indefinitely for network idle on applications with analytics, polling or WebSockets.
  • Store the returned bytes directly when possible; avoid unnecessary temporary files.
  • Measure your own page mix. Rendering cost depends on HTML complexity, JavaScript, fonts, images and concurrency; the reviewed documentation does not establish a universal benchmark.

11. Or skip the browser setup

ScreenshotNeo provides a website capture API and PDF support through one GET request. See the ScreenshotNeo documentation for PDF output options and the full set of capture controls.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o document.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("document.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const pdf = await res.arrayBuffer();

ScreenshotNeo removes cookie banners, newsletter popups and chat widgets before capture. Bot checks, blank pages and failed loads are not billed, and the response identifies the page verdict and billing status. Its MCP server lets AI agents call take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to start with 1,000 screenshots per month and no card.

12. FAQ

How do I convert HTML to PDF in Node.js?

Use Puppeteer: launch a browser, load HTML with setContent or navigate with goto, call page.pdf(), then close the browser.

What npm package converts HTML to PDF?

Puppeteer is the clearest documented browser-rendering route. Wrappers exist, but verify their current compatibility and maintenance before choosing one.

Can Puppeteer render JavaScript before creating the PDF?

Yes. It renders the page in a browser; your code must still wait for the application’s data and assets to be ready.

Can I use CSS @page rules?

Yes. Use preferCSSPageSize: true when the CSS page size should take priority over the PDF format option.

Should I use PDFKit for HTML conversion?

PDFKit is a programmatic PDF document-generation library. It is not presented by its npm description as a drop-in arbitrary HTML/CSS renderer.