HTML to PDF Conversion: Libraries vs. Headless Browsers vs. APIs
Compare browser printing, document renderers, and managed APIs for HTML-to-PDF conversion, with runnable examples and a practical selection guide.
Choose a headless browser when the PDF should reproduce a live, JavaScript-driven web page. Choose a document renderer such as WeasyPrint when you are producing structured documents and its HTML/CSS support matches your templates. Choose a managed API when you want a service to handle rendering infrastructure and its privacy, reliability, limits, and cost meet your requirements. There is no universal winner: test representative pages against the same acceptance criteria before committing.
This guide compares the approaches, shows runnable starting points, and covers the operational decisions that usually determine whether a conversion pipeline is dependable.
1. Decide what the PDF must represent
Start with the document, not the library. A PDF made from a live application page has different needs from an invoice or a long report with predictable pagination.
| Approach | Good fit | Investigate before adopting |
|---|---|---|
| Headless browser: Puppeteer or Playwright | PDFs should reflect a browser-rendered page, including client-side JavaScript and the page’s print styles. | Browser installation and versioning, concurrency, memory and CPU use, print CSS, fonts, page breaks, and color handling. |
| Dedicated renderer: WeasyPrint | Document-oriented output where its HTML/CSS support and PDF features suit your templates. | Exact CSS support, pagination, font availability, deployment dependencies, and visual changes between versions. |
| Managed API | You want to outsource rendering operations to a service that accepts content or a URL and returns a PDF. | Data handling, security, availability, location, limits, expected-volume cost, failure behavior, and provider lock-in. |
For browser output, the intended result is normally the browser’s print rendering, not a screenshot of the screen. Puppeteer and Playwright use print CSS media for PDF generation by default. They also document that print colors can be adjusted by default, so a page that looks correct on screen may need print-specific styles or explicit color handling.
2. Use a headless browser for browser-rendered pages
Puppeteer with Node.js
Install Puppeteer in a Node.js project; its package setup provides a compatible browser installation. Save this as pdf-puppeteer.mjs and run it with node pdf-puppeteer.mjs.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com', {
waitUntil: 'networkidle2',
timeout: 60_000,
});
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
});
} finally {
await browser.close();
}
The PDF method uses print media by default. If the document must use screen CSS, emulate screen media before calling page.pdf(). For colors that must match the design, set print color adjustment in the page CSS, for example * { -webkit-print-color-adjust: exact; }, and inspect the generated PDF in your target viewers.
Playwright with Node.js
Install Playwright and its browser binaries for your deployment environment. This runnable example writes the returned PDF buffer to disk.
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com', {
waitUntil: 'networkidle',
timeout: 60_000,
});
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
displayHeaderFooter: false,
});
await writeFile('page.pdf', pdf);
} finally {
await browser.close();
}
Playwright’s PDF options include paper format, margins, header and footer templates, page ranges, scaling, background graphics, and preferCSSPageSize. Letter is the documented default format. Header/footer templates have constraints: scripts in the templates are not evaluated, and page styles are not visible inside them. Consult the [Playwright Page API](https://playwright.dev/docs/api/class-page) for the option details supported by your installed version.
Browser conversion details that matter
- Wait for the right condition. Network-idle waits can be unsuitable for pages with long polling or continuously active requests. For dynamic pages, wait for a meaningful selector or application-ready signal, then print.
- Set the print layout deliberately. Define page size, margins, page breaks, and hidden interactive elements in print CSS. Use
@media printand, where appropriate,@page. - Make assets available. Fonts and images must finish loading before PDF generation. A successful navigation does not prove every asset is ready.
- Control page access. If pages require authentication, configure the browser context or page before navigation. Treat credentials and rendered content as sensitive.
- Close resources. Close pages and browsers in cleanup paths, including after timeouts, so failed jobs do not accumulate browser processes.
3. Use a dedicated renderer for document-oriented output
WeasyPrint is a Python HTML/CSS-to-PDF renderer. It documents PDF hyperlinks, bookmarks, attachments, and forms. Those features can be useful when the output is a structured document rather than a reproduction of an interactive browser session. Confirm that its supported HTML and CSS cover your actual templates, especially pagination, fonts, and layout details.
Install WeasyPrint following its [official installation instructions](https://doc.courtbouillon.org/weasyprint/stable/first_steps.html), which account for platform-specific dependencies, then save the following as render.py:
from weasyprint import HTML
HTML(url='https://example.com').write_pdf('page.pdf')
For a local HTML file, use HTML(filename='document.html').write_pdf('document.pdf'). For controlled input, pass an explicit base URL so relative image, stylesheet, and font references resolve as intended:
from weasyprint import HTML
HTML(
string='''<!doctype html>
<html><head><meta charset="utf-8"><style>
@page { size: A4; margin: 12mm; }
body { font: 11pt sans-serif; }
</style></head><body><h1>Report</h1><p>Generated document.</p></body></html>''',
base_url='https://example.com/',
).write_pdf('report.pdf')
WeasyPrint warns that documents can render differently after a version change even when the API remains stable. Treat renderer upgrades as output changes: keep representative documents and compare their PDFs or rendered pages in regression review.
4. Evaluate a managed API
A managed HTML-to-PDF API can move browser or renderer operations out of your application deployment. That trades infrastructure work for a dependency on a provider. Before choosing one, review where submitted HTML and page content are processed, retention and deletion terms, authentication, network access, service limits, failure handling, availability commitments, and price at your expected volume. A provider’s own comparison or performance claims are vendor-authored; validate them independently for your workload.
For a ScreenshotNeo alternative, start with [ScreenshotNeo](https://screenshotneo.com): it offers website captures and PDF output through an API and an MCP server, and bills only clean shots. Its documented options include PDF paper size, margins, landscape, and page ranges. Check the [ScreenshotNeo documentation](https://screenshotneo.com/docs/) for the current PDF request configuration and authentication details.
5. Compare on the same production-like inputs
There is no neutral controlled evaluation in the cited research that establishes a universal cost or speed winner across browsers, dedicated renderers, and APIs. A vendor-published comparison reports particular Puppeteer and WeasyPrint timings, output sizes, and installation footprints for its own workloads; those figures are not general guarantees. Use your own acceptance criteria and workload instead.
- Choose representative inputs. Include short and long documents, complex tables, charts, large images, custom fonts, right-to-left or non-Latin text if relevant, and pages that load client-side data.
- Define correctness. Check page count, clipping, line wrapping, links, bookmarks or forms where needed, colors, font embedding/rendering, and expected page breaks.
- Measure the full pipeline. Record cold-start and warm latency, throughput at expected concurrency, memory, CPU, output size, retries, and failure rate. Keep the environment and input set the same.
- Exercise failure cases. Include a slow page, missing asset, invalid URL, failed navigation, and a page that never becomes network-idle.
- Repeat after upgrades. Pin browser or renderer versions where reproducibility matters, and re-run visual/document checks after changing versions or templates.
6. Reliability, security, and cost decisions
Reliability
- Set navigation and job timeouts. Return a clear failure for a page that never reaches the required ready state.
- Use bounded concurrency. Browser rendering can consume substantial resources; determine safe parallelism from measurements in your deployment environment.
- Make retries selective. Retrying a transient navigation failure may help, but retries cannot fix a consistently broken template or inaccessible resource. Use a retry limit and avoid duplicate downstream work.
- Track conversion outcomes separately from HTTP success. A response can be successful while the PDF is incomplete or visually wrong.
- Keep a small set of output fixtures for regression checks. Compare page count and rendered pages, since byte-for-byte PDF output can vary for reasons unrelated to visible content.
Security
- Converting arbitrary URLs can expose internal network services or local resources. Restrict destinations and redirects, and isolate workers from sensitive networks.
- Do not place secrets in URLs that may be logged. Protect cookies, authorization headers, generated PDFs, and temporary files.
- For a hosted API, assess the provider’s data handling and access controls before sending private pages or customer data.
- Limit input size, execution time, and output size to contain malformed or unusually expensive documents.
Cost
For self-hosted conversion, include engineering and operational work, browser or renderer deployment, compute, memory, scaling, monitoring, and upgrade regression review. For an API, model price against actual conversion volume and include retries, failed jobs, peak usage, and any limits or overages. Do not infer your costs from a benchmark’s render time or a provider’s sample price alone.
7. Troubleshooting common failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| PDF is blank or missing page content | Navigation completed before client rendering, or the page failed to load its data. | Wait for an application-ready selector or explicit page condition; capture console and navigation errors. |
| Styles differ from the browser view | PDF generation uses print media by default. | Add print CSS, or emulate screen media in Puppeteer if screen styling is the intended output. |
| Backgrounds or colors are missing | Background printing is disabled or print color adjustment changes colors. | Enable background printing and set print color adjustment where appropriate; inspect the resulting PDF. |
| Fonts or images are absent | Assets are inaccessible, relative URLs resolve against the wrong base, or rendering starts before they load. | Check network access and base URL, ensure fonts are installed/served, and wait for required assets. |
| Content is clipped or split badly | Paper size, margins, page breaks, or long unbreakable elements are not accounted for. | Set page dimensions and margins explicitly; add print-specific break rules and test long tables and code blocks. |
| Browser job hangs at network idle | The site maintains long-lived connections or periodic requests. | Wait for a specific selector or app signal instead of network idle; keep a hard timeout. |
| WeasyPrint output changes after upgrade | Rendering behavior can change across versions. | Compare representative documents, review the version’s release notes, and pin or roll back until templates are reviewed. |
| Managed API returns an error or unexpected PDF | Authentication, request options, URL access, limits, or provider behavior may differ from assumptions. | Check the provider’s current API docs and response details; retry only transient failures and verify data/URL access permissions. |
8. Or skip the browser setup
ScreenshotNeo can return PDF output from a website capture request. The call below demonstrates the supplied API request shape; consult the [PDF documentation](https://screenshotneo.com/docs/) for the PDF-specific request configuration supported by the API.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
For PDF output, configure the documented PDF option and output filename/handling for the response instead of assuming the example’s .webp filename represents a PDF. The same service also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for MCP clients such as Claude and Cursor.
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server lets AI agents take screenshots and capture PDFs.
- 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
9. FAQ
Can one renderer handle both web pages and formal documents?
It may, but choose based on the features and output requirements you actually need. A browser is a natural fit for browser-rendered pages; a document renderer may suit structured documents. Validate the same templates and acceptance criteria before consolidating on one tool.
Should I generate PDFs in the request path?
That depends on latency and workload. For slow or variable pages, a queued job can isolate rendering time from a user-facing request. Whichever pattern you use, define timeouts, concurrency limits, and how callers learn that a job failed.
Are PDF files visually identical across viewers?
Not necessarily. Fonts, color handling, and viewer behavior can affect presentation. Review output in the viewers and environments your recipients use.
Which approach is cheapest?
The dossier does not establish a universal cost winner. Compare the full cost of operating a self-hosted renderer with the provider’s charges and operational terms at your expected volume.
