Best Node.js Libraries for Converting HTML to PDF
Compare Puppeteer, Playwright, PDFKit, and hosted APIs for Node.js HTML-to-PDF conversion, with runnable code and troubleshooting guidance.
The best Node.js HTML-to-PDF library depends on what you are converting. Use Puppeteer or Playwright when a browser must render an existing HTML page and its CSS. Use PDFKit when your application can create the PDF layout directly. Use a hosted conversion API when you want to avoid operating a local browser process.
| Approach | Best fit | Key decision | Limit to understand |
|---|---|---|---|
| Puppeteer | Print a page with a Chromium browser | Print CSS, page format, headers and footers | Requires browser automation setup |
| Playwright | Print a page in a browser automation project | Reuse the browser tooling already in your stack | The supplied research does not establish a speed or output-quality winner over Puppeteer |
| PDFKit | Programmatically draw a PDF | Direct control of content, layout and streams | Its documentation does not establish arbitrary HTML rendering |
| Hosted API | Managed conversion without a local browser | Deployment, privacy, reliability and service limits | Provider claims and terms must be checked for your workload |
1. Puppeteer: browser-based HTML-to-PDF conversion
Puppeteer’s documented workflow is to navigate to a page and call page.pdf(). PDF generation uses print CSS. If the page is designed for the screen, call page.emulateMediaType('screen') before generating the file. The method waits for fonts by default. See the Puppeteer PDF guide, Page.pdf API, and PDF options reference.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com/invoice/123', {
waitUntil: 'networkidle0'
});
// Use this only when the page's screen CSS should be printed.
await page.emulateMediaType('screen');
await page.pdf({
path: 'invoice.pdf',
format: 'A4',
printBackground: true,
margin: { top: '16mm', right: '14mm', bottom: '16mm', left: '14mm' },
displayHeaderFooter: true,
headerTemplate: '<span></span>',
footerTemplate: '<div style="font-size:9px;width:100%;text-align:center">Page <span class="pageNumber"></span> of <span class="totalPages"></span></div>'
});
} finally {
await browser.close();
}
Print colors can differ from the browser view. Puppeteer documents that printing adjusts colors by default; add -webkit-print-color-adjust: exact in your print stylesheet when exact colors are required, then verify the result in your own documents.
@media print {
-webkit-print-color-adjust: exact;
print-color-adjust: exact;
}
Useful Puppeteer options
format, or explicitwidthandheight, controls paper size.marginreserves printable space.printBackgroundincludes CSS backgrounds.displayHeaderFooter,headerTemplate, andfooterTemplateadd repeating page content. Page number and total-page placeholders are available in templates.pageRangeslimits output to selected pages when your use case needs an excerpt.- Wait for your own readiness signal when the document depends on client-side rendering:
await page.waitForSelector('#report-ready')or an application-specific delay.
2. Playwright: PDF output with the same browser-print model
Playwright’s page.pdf() returns a PDF buffer and renders with print CSS. Its API also documents emulating screen media before PDF generation when screen styling is required. Choose it when Playwright is already the browser automation environment used by your project; the supplied research does not establish that it is faster or produces higher-quality PDFs than Puppeteer.
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';
const browser = await chromium.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com/report', { waitUntil: 'networkidle' });
await page.emulateMedia({ media: 'screen' });
await page.waitForSelector('#report-ready');
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
margin: { top: '16mm', right: '14mm', bottom: '16mm', left: '14mm' }
});
await writeFile('report.pdf', pdf);
} finally {
await browser.close();
}
Keep the same checks as with Puppeteer: choose print or screen media deliberately, wait for data and fonts, include backgrounds when needed, and inspect page breaks. The Playwright API is the authoritative place to check the options supported by the version installed in your project: Page API.
3. PDFKit: direct PDF generation in Node.js
PDFKit is a JavaScript library for creating PDF documents directly. Its documentation describes PDFDocument instances as readable Node streams and shows piping one to a file or HTTP response before calling end(). This is a good fit when the application owns the document layout and can place text, images and drawing primitives itself. Do not treat the cited documentation as evidence that PDFKit is a drop-in renderer for arbitrary HTML and CSS.
import PDFDocument from 'pdfkit';
import { createWriteStream } from 'node:fs';
const doc = new PDFDocument({ size: 'A4', margin: 50 });
doc.pipe(createWriteStream('summary.pdf'));
doc.fontSize(22).text('Monthly summary');
doc.moveDown();
doc.fontSize(11).text('Generated directly with PDFKit.');
doc.moveDown();
doc.text('Add your own layout, tables, images and page management here.');
doc.end();
For an HTTP response, pipe the document to the response and set the content type before writing:
app.get('/summary.pdf', (req, res) => {
res.type('application/pdf');
const doc = new PDFDocument();
doc.pipe(res);
doc.fontSize(18).text('Summary');
doc.end();
});
Choose PDFKit when deterministic, application-owned drawing is more important than reproducing an existing web page. If your source is already HTML with complex CSS, a browser renderer usually requires less rewriting.
4. Hosted HTML-to-PDF APIs
A hosted API receives HTML or a URL and returns PDF bytes. This model can remove browser installation and process management from your deployment. The research dossier includes one provider-authored example, pdfkitt’s Node.js HTML-to-PDF page; its service claims, security model, retention, limits, pricing and reliability should be verified directly before adoption. Compare providers on data handling, authentication, page controls, regional availability, failure reporting and total cost for your traffic.
5. How to choose
- Existing web page or template: start with Puppeteer or Playwright.
- Need screen styling: emulate screen media explicitly and test colors and page breaks.
- Application creates every layout element: evaluate PDFKit.
- Do not want a local browser: evaluate a hosted API and its data-handling terms.
- Need a reliable production path: define readiness checks, timeouts, retries, observability and a failure policy before shipping.
6. Production checklist
- Pin and review the Puppeteer, Playwright or PDFKit version used by your application.
- Wait for application data, images and fonts before capturing.
- Set an explicit paper size, margins and background policy.
- Include a print stylesheet for page breaks, hidden controls and color handling.
- Test long tables, very wide content, missing images, empty states and right-to-left or non-Latin text if applicable.
- Close browser pages and processes in success and error paths.
- Stream large PDFs where practical instead of buffering multiple copies.
- Record the input URL or document identifier, duration, output size and failure reason without logging secrets or sensitive HTML.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF looks different from the screen | PDF uses print CSS | Use emulateMediaType('screen') or emulateMedia({media:'screen'}), then maintain a print stylesheet deliberately. |
| Background colors are missing | Background printing is disabled | Set printBackground: true and check print color adjustment CSS. |
| Charts or data are absent | Capture happened before client rendering completed | Wait for a specific selector or readiness event after data loading. |
| Fonts change layout | Web fonts were not available when layout was captured | Wait for font loading and verify the font files are reachable from the browser context. |
| Content is clipped | Fixed dimensions, overflow rules or unsuitable margins | Inspect print CSS, remove unintended fixed heights, and set paper dimensions and margins explicitly. |
| Browser launch fails in deployment | Browser executable or runtime dependencies are unavailable | Install the supported browser/runtime for your chosen package or use a managed conversion service. |
| Output hangs | Page keeps connections open or waits for an event that never occurs | Use a bounded navigation and application timeout, define a readiness selector, and close the page in a finally block. |
| PDFKit output is not HTML-like | PDFKit is drawing a document rather than rendering a browser page | Use a browser renderer for existing HTML, or implement the required layout directly with PDFKit. |
8. Performance, reliability and cost considerations
The supplied sources do not provide a comparable benchmark, compatibility matrix or pricing study, so there is no evidence-based universal winner. Measure your own representative pages. Browser conversion cost is affected by page complexity, assets, fonts and concurrency; PDFKit avoids browser rendering when its direct drawing model fits. Hosted services trade local process management for provider pricing, network latency and service limits.
For reliability, make readiness and timeout behavior explicit, retry only failures that are safe to retry, and keep generated files identifiable so duplicate jobs can be detected. For cost, count browser resources and hosted API requests in your workload model, and test the largest and slowest documents rather than relying on a single small page.
9. Or skip the browser setup
ScreenshotNeo is a hosted website capture API that can return PNG, JPEG, WebP or PDF. Its PDF options include paper size, margins, landscape mode and page ranges. A single request can capture a URL without installing or managing Puppeteer or Playwright.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for request options. Cookie and consent banners, newsletter popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed. Responses identify the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also provides an MCP server so AI agents can take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
10. FAQ
Which Node.js library converts HTML to PDF?
For existing HTML pages, use Puppeteer or Playwright. For direct programmatic document creation, use PDFKit. A hosted API is an alternative when browser operations should be managed remotely.
Can PDFKit convert arbitrary HTML?
The cited PDFKit documentation covers direct PDF generation and Node.js streams. It does not establish arbitrary HTML and CSS rendering.
Should I use Puppeteer or Playwright?
Use the one that matches your existing browser automation stack and verify the output and deployment behavior for your pages. The supplied research does not establish a general performance or quality winner.
Why does my PDF not match the browser?
PDF generation uses print CSS by default. Emulate screen media when appropriate, define print styles, wait for content and fonts, and set background and color behavior explicitly.
Is a hosted API always cheaper?
No conclusion is supported without your page mix, volume, infrastructure and provider terms. Compare measured conversion cost, operations work, latency and data requirements.
