Convert HTML to PDF with an API
Compare hosted HTML-to-PDF APIs with Puppeteer and Playwright, then use complete code for reliable CSS, fonts, JavaScript, and page breaks.
Use a managed HTML-to-PDF API when you want conversion without operating browsers. Use Puppeteer or Playwright when you need complete control over the browser and deployment. In either architecture, load the page, wait for fonts and dynamic content, choose print or screen CSS, then export with paper, margin, background, and page-break options.
This guide covers URL conversion, HTML strings, JavaScript-rendered pages, CSS and font loading, headers, cookies, page ranges, troubleshooting, and production trade-offs.
1. Choose an architecture
| Approach | Best for | Trade-offs |
|---|---|---|
| Hosted REST API | Teams that want browser operations, scaling, and patching handled by a provider | Provider-specific authentication, quotas, retention, regions, and pricing; verify current settings |
| Puppeteer | Node.js services requiring Chromium control and custom network access | You operate browser binaries, memory, concurrency, isolation, and updates |
| Playwright | Teams already using Playwright and Chromium automation | Documented PDF export is Chromium-only |
Adobe documents an HTML-to-PDF REST operation for static HTML, dynamic HTML, ZIP input, and URLs: Adobe HTML-to-PDF documentation. Authentication, upload flow, quotas, retention, data residency, and pricing are provider-specific and should be checked before production use.
2. Convert a URL with Puppeteer
Puppeteer’s page.pdf() generates a PDF using the print CSS media type by default. Use emulateMediaType('screen') when the PDF should use screen styling. The API supports paper format or dimensions, orientation, margins, page ranges, background graphics, scale, CSS page-size preference, and headers or footers. See the Puppeteer page.pdf() reference.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({
headless: true
});
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.goto('https://example.com', {
waitUntil: 'networkidle2',
timeout: 60_000
});
// Omit this line to use print CSS (the default).
await page.emulateMediaType('screen');
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'document.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: {
top: '18mm',
right: '16mm',
bottom: '18mm',
left: '16mm'
}
});
} finally {
await browser.close();
}
Puppeteer’s guide notes that page.pdf() waits for fonts by default, but explicitly waiting for document.fonts.ready makes the readiness condition visible in your code. PDF generation guide
Convert an HTML string
import puppeteer from 'puppeteer';
const html = `<!doctype html>
<html>
<head>
<meta charset="utf-8">
<style>
@page { size: A4; margin: 18mm; }
body { font-family: Arial, sans-serif; }
h1 { break-after: avoid; }
.page-break { break-before: page; }
</style>
</head>
<body>
<h1>Invoice</h1>
<p>Generated from an HTML string.</p>
<div class="page-break">Second page</div>
</body>
</html>`;
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setContent(html, { waitUntil: 'networkidle0' });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'invoice.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
} finally {
await browser.close();
}
Wait for application data
Network idle alone may fire before a client-side application finishes rendering. Wait for a stable selector or an application-specific promise.
await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-report-ready]', { timeout: 30_000 });
await page.evaluate(() => document.fonts.ready);
await page.pdf({ path: 'report.pdf', format: 'Letter', printBackground: true });
3. Convert HTML with Playwright
Playwright’s page.pdf() uses print media by default. To follow screen styles, call page.emulateMedia({ media: 'screen' }). Its PDF API includes paper size, width and height, margins, page ranges, print backgrounds, scale, CSS page-size preference, headers and footers, and tagged output options. Playwright documents that PDF generation is Chromium-only. See the Playwright page.pdf() reference.
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto('https://example.com', {
waitUntil: 'networkidle',
timeout: 60_000
});
await page.emulateMedia({ media: 'screen' });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'document.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '18mm', right: '16mm', bottom: '18mm', left: '16mm' }
});
} finally {
await browser.close();
}
4. PDF layout options that affect output
Print versus screen CSS
Both browser APIs default to print media. Put print-specific rules in @media print, or emulate screen media when your web layout is the source of truth.
@media print {
nav, .cookie-banner, .interactive-controls { display: none; }
a { color: inherit; text-decoration: none; }
}
@page {
size: A4 portrait;
margin: 18mm 16mm;
}
.invoice-item { break-inside: avoid; }
.chapter { break-before: page; }
Paper, dimensions, orientation, and margins
Use a named format such as A4 or Letter, or provide explicit width and height. Set orientation to landscape for wide tables. Keep enough margin for printers and avoid placing critical content in headers or footers.
Backgrounds and scale
Enable background graphics with printBackground: true in Puppeteer or Playwright. A scale below 1 fits more content but reduces readable size; test the resulting pages rather than assuming a particular scale will work for every document.
CSS page size
With preferCSSPageSize: true, the browser honors the document’s @page size instead of overriding it with the API format. Choose one source of truth for paper size to avoid surprises.
Page ranges and headers
Export only selected pages with the page-ranges option when generating previews or extracts. Headers and footers use browser templates and have restrictions on available styling and JavaScript; keep them simple and verify page numbers in the final PDF.
Fonts and external assets
Fonts, images, stylesheets, and scripts must be reachable from the rendering environment. Wait for document.fonts.ready, use absolute asset URLs when converting an HTML string, and ensure private assets are authenticated. A missing webfont can change line wrapping and therefore page breaks.
5. Hosted conversion with Adobe PDF Services
Adobe’s HTML-to-PDF operation accepts static and dynamic HTML, ZIP input, and a URL. Its documented REST example uses a POST request to https://pdf-services.adobe.io/operation/htmltopdf with an API key, bearer token, asset ID, and rendering options. Follow Adobe’s current authentication and upload instructions rather than copying credentials into source code.
curl -X POST "https://pdf-services.adobe.io/operation/htmltopdf" \
-H "x-api-key: $ADOBE_API_KEY" \
-H "Authorization: Bearer $ADOBE_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"assetID": "YOUR_ASSET_ID",
"renderingOptions": {
"pageSize": "A4"
}
}'
The exact asset upload, polling, and download steps depend on the provider workflow documented by Adobe. Confirm current quotas, retention, region, and pricing before selecting a hosted service.
6. Or skip the browser setup
ScreenshotNeo provides a website capture API that can return PNG, JPEG, WebP, or PDF. It handles the browser layer and exposes PDF options such as paper size, margins, landscape mode, and page ranges. The API also supports custom CSS and JavaScript, headers, cookies, user agents, authorization, waiting rules, blocking rules, caching, asynchronous jobs, bulk capture, and signed webhooks. See the ScreenshotNeo API documentation.
One request starts a capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For PDF output and the complete option names, use the documentation. Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and whether the request was billed. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
7. Reliability and production checklist
- Pin or regularly update the Chromium version used by your renderer.
- Set navigation, selector, and total-job timeouts.
- Wait for the selector that proves application data is rendered.
- Wait for fonts and critical images before export.
- Use retries only for transient navigation failures, with an idempotency strategy.
- Limit concurrent pages according to available CPU and memory.
- Isolate untrusted pages and restrict outbound network access where appropriate.
- Record browser errors, URL, render duration, page count, and output size.
- Compare representative PDFs after browser, CSS, or font changes.
8. Performance, cost, and scaling
Cold-starting a browser costs more than reusing a process, while too much concurrency causes memory pressure and navigation failures. Reuse a browser process, create isolated pages or contexts, cap concurrency, and close pages in a finally block. Wait only as long as the page needs: a selector or application-ready signal is usually more predictable than an arbitrary delay.
Hosted APIs move browser CPU, patching, and scaling into provider billing. Self-managed rendering moves those costs into your infrastructure and operations. Compare browser startup time, memory per concurrent page, external network access, font availability, print fidelity, data residency, security isolation, observability, quotas, retention, and vendor lock-in. Do not assume a fixed latency or universal pixel fidelity across engines.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Blank or partially rendered PDF | Export ran before client-side rendering finished | Wait for a readiness selector or application promise; then wait for fonts |
| Wrong colors or missing backgrounds | Print media or background printing is disabled | Emulate screen media when needed and enable printBackground |
| Text wraps differently | Webfont failed to load or viewport differs | Check font requests, await document.fonts.ready, and set the viewport explicitly |
| Images missing | Relative URLs, blocked requests, or lazy loading | Use absolute URLs, allow required assets, and trigger or wait for lazy content |
| Unexpected page breaks | Conflicting API and CSS page sizes or break rules | Choose one page-size source and use break-inside, break-before, and break-after |
| Navigation timeout | Slow dependency, blocked host, or never-ending request | Inspect network errors, set a realistic timeout, and wait for a specific selector instead of global idle |
| Playwright PDF call fails outside Chromium | PDF export is Chromium-only | Launch the Chromium browser supplied for the project |
| Header or footer overlaps content | Margins are too small for the template | Increase top or bottom margins and verify the rendered pages |
10. FAQ
Can an API convert an HTML string instead of a URL?
Yes. Set the page content directly with Puppeteer or Playwright, or use a hosted provider’s documented HTML or ZIP upload flow.
How do I preserve JavaScript-generated content?
Use a browser renderer, wait for the application’s ready state, and verify that required API calls and assets are reachable from the renderer.
Why does my PDF look different from the website?
PDF export defaults to print CSS. Screen media, fonts, viewport dimensions, page size, and background settings can all change layout.
Is Playwright’s PDF export cross-browser?
The documented PDF capability is Chromium-only.
Should I use a hosted API or run Chromium myself?
Use a hosted API when reducing browser operations is the priority. Run Puppeteer or Playwright when browser version, network access, deployment, and isolation must remain under your control.


