How to Build an HTML-to-PDF Converter App
Build a secure HTML-to-PDF service with Puppeteer, print CSS, validation, SSRF defenses, queues, troubleshooting, and a managed API option.
Direct answer: run a Chromium worker with Puppeteer or Playwright, load a controlled HTML document, wait for fonts and images, switch to print media, and call page.pdf(). Put validation, sanitization, network controls, timeouts, queueing, and browser isolation around that call before accepting production traffic.
This guide builds a Node.js converter, explains print CSS and engine choices, and covers the security and operations work that determines whether it remains reliable.
1. Choose the input contract
The safest API accepts a server-side template identifier plus validated data. If callers submit HTML, treat it as executable browser input: cap its size, sanitize it, allow only approved asset URLs, and disable navigation to arbitrary hosts.
| Contract | Use when | Controls |
|---|---|---|
| Template ID + data | Invoices, reports, certificates | Allowlisted templates, schema validation, escaped values |
| Sanitized HTML | Limited user authoring | HTML sanitizer, URL and CSS policy, size limits |
| URL | Controlled sites or migration | Host allowlist, DNS and redirect checks, egress filtering |
2. Minimal Puppeteer service
npm install express puppeteer
import express from 'express';
import puppeteer from 'puppeteer';
const app = express();
app.use(express.json({ limit: '1mb' }));
app.post('/convert', async (req, res, next) => {
if (typeof req.body?.html !== 'string' || !req.body.html.length) {
return res.status(400).json({ error: 'html is required' });
}
const browser = await puppeteer.launch({ headless: true, args: ['--disable-dev-shm-usage'] });
try {
const page = await browser.newPage();
await page.setContent(req.body.html, { waitUntil: 'networkidle0' });
await page.emulateMediaType('print');
const pdf = await page.pdf({ format: 'A4', printBackground: true, preferCSSPageSize: true, tagged: true, timeout: 30000 });
res.type('application/pdf').set('Content-Disposition', 'attachment; filename="document.pdf"').send(pdf);
} catch (err) {
next(err);
} finally {
await browser.close();
}
});
app.listen(3000, () => console.log('listening on 3000'));
Puppeteer documents the launch, navigation, page.pdf(), and close sequence. Its PDF guide says fonts are awaited by default: Puppeteer PDF generation.
node server.mjs
curl -X POST http://localhost:3000/convert \
-H 'content-type: application/json' \
--data '{"html":"<h1>Invoice</h1><p>Paid</p>"}' \
-o document.pdf
3. Print CSS and pagination
@page { size: A4; margin: 18mm 14mm 20mm; }
html { font-family: 'Inter', Arial, sans-serif; }
body { margin: 0; color: #111; }
h1, h2 { break-after: avoid; }
table, tr, .invoice { break-inside: avoid; }
.cover { break-after: page; }
@media print { .screen-only { display: none !important; } }
- Use
@pagefor paper size and margins, withpreferCSSPageSize: truewhen CSS owns geometry. - Use
break-before,break-after, andbreak-insidefor headings, rows, and invoice blocks. - Set
print-color-adjust: exactonly where backgrounds are essential. - Embed or preload exact fonts. Missing fonts change line wrapping and pagination.
- Choose explicitly whether external images, fonts, and JavaScript are allowed.
4. Wait for deterministic readiness
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-pdf-ready="true"]', { timeout: 15000 });
await page.evaluate(() => document.fonts.ready);
await page.evaluate(() => Promise.all([...document.images].map(img => img.complete ? null : new Promise(r => { img.onload = img.onerror = r; }))));
For controlled pages, set data-pdf-ready='true' after charts and asynchronous data have finished. A fixed short delay is less reliable than a readiness signal.
5. Engine choices
| Engine | Best fit | Tradeoffs |
|---|---|---|
| Puppeteer | Chromium fidelity and JavaScript-heavy pages | Browser process cost; sandbox and network hardening |
| Playwright | Chromium rendering with broader automation | Same worker and egress concerns |
| wkhtmltopdf | Simple CLI deployment and legacy layouts | Older WebKit engine |
| PDFKit | Structured programmatic documents | Coordinate-based; does not render HTML/CSS |
Playwright also uses print CSS by default. Use screen media only when required. PDFKit is a different, coordinate-driven model rather than a drop-in HTML renderer.
6. Security controls
- HTML/XSS: sanitize untrusted markup, remove event-handler attributes, and reject dangerous URL schemes such as
javascript:. - SSRF: prefer an identifier or allowlisted host over a complete user URL. Block loopback, link-local, private, and metadata ranges. Re-check redirects and protocol changes.
- Isolation: run workers as a low-privilege user in a separate container with a read-only filesystem, no cloud credentials, and restricted egress.
- Resource exhaustion: cap HTML bytes, image dimensions, page count, render time, memory, concurrency, and output bytes.
- Data handling: avoid logging raw HTML or PDFs, encrypt stored results, use short retention, and remove temporary files.
7. Production operations
Use a synchronous endpoint for small documents. For large files or bursts, enqueue a job and return 202 Accepted with a status identifier; store the result in object storage and provide a time-limited download URL.
- Reuse browsers for several jobs, then recycle them after a fixed count or memory threshold.
- Record validation, navigation, readiness, render duration, bytes, page count, and failure class.
- Return stable errors such as
invalid_html,blocked_url,timeout,renderer_crash, andoutput_too_large. - Attach a renderer version to every job.
- Keep fixtures for long tables, RTL text, Unicode, charts, headers, footers, and large documents. Compare page count, text, and rasterized snapshots during upgrades.
8. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Blank or partial PDF | Async content was not ready | Wait for a readiness selector, fonts, and images. |
| Wrong page breaks | Fallback fonts or conflicting margins | Embed fonts, define @page, and choose CSS sizing deliberately. |
| Missing backgrounds | Print backgrounds disabled | Set printBackground: true. |
| Missing images | Blocked origin, CORS, or lazy loading | Allow the asset host and wait for images. |
| Navigation timeout | Third-party request or never-ending page | Allowlist hosts, block unnecessary resources, and enforce a deadline. |
| Browser crashes | Too many concurrent pages or huge assets | Bound concurrency, cap input/output, monitor memory, and recycle workers. |
9. Performance and cost
Browser startup is expensive, so warm workers reduce latency. Queue limits protect tail latency. Cache immutable inputs with a key containing template, CSS, font, and renderer versions. Store large PDFs outside the API process and stream downloads.
Costs come from CPU and memory time, concurrency, asset transfer, and storage retention. Measure render duration and output bytes by tenant before setting quotas. Self-hosted assets are usually more reproducible than many third-party requests.
10. Test checklist
- Test malformed HTML, oversized payloads, deep nesting, huge images, and disallowed URLs.
- Verify paper sizes, landscape mode, margins, page ranges, headers, footers, and print colors.
- Test fonts, Unicode, RTL text, long strings, tables crossing pages, and charts.
- Kill workers during navigation and verify clean failure and temporary-file removal.
- Load-test queue backpressure so one tenant cannot consume every worker.
11. Or skip the browser setup
ScreenshotNeo is a website screenshot API that also returns PDFs. Use its managed endpoint when you do not want to operate browser workers:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o document.pdf
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("document.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const data = await res.arrayBuffer();
await import('node:fs/promises').then(fs => fs.writeFile('document.pdf', Buffer.from(data)));
See the ScreenshotNeo API documentation for PDF options. Cookie and consent banners are accepted, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the verdict and billing. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account.
FAQ
Should I accept arbitrary HTML?
Only with sanitization, strict limits, and an isolated worker. Templates plus validated data are safer for multi-tenant products.
Why does the PDF differ from the browser tab?
PDF uses print media by default. Check print CSS, fonts, viewport, margins, and readiness timing.
When should conversion become asynchronous?
Use a queue when documents are large, traffic is bursty, or callers can tolerate polling or a webhook.
Is PDFKit a drop-in HTML converter?
No. It is coordinate-driven; choose it when you control layout programmatically.


