wkhtmltopdf vs Puppeteer for Converting Web Pages to PDF
Compare wkhtmltopdf and Puppeteer for web-to-PDF conversion, with runnable examples, print controls, security guidance, and a practical migration checklist.
Short answer: For a new HTML-to-PDF workflow, start with Puppeteer if pages use modern JavaScript, current browser behavior, or print CSS. Keep wkhtmltopdf for an existing workflow only after checking its rendering limits, exact binary build, security boundary, and output compatibility. There is no controlled head-to-head benchmark in the sources used here, so measure speed and fidelity on your own representative pages.
This comparison covers converting web pages to PDF with both tools, including runnable examples, print behavior, deployment, security, troubleshooting, and a practical way to decide whether migration is worthwhile.
1. The core difference: browser engine and page behavior
wkhtmltopdf uses Qt WebKit. The project status page notes that QtWebKit was deprecated in 2015 and removed from Qt in 2016. The project’s downloads page lists 0.12.6 as its stable series, released June 11, 2020. Puppeteer automates a more modern browser engine, which makes it a stronger default for pages whose content depends on current JavaScript and CSS behavior. These are project-maintainer descriptions, not a measured comparison of output quality or speed. wkhtmltopdf project status, downloads and FAQ
| Need | Starting point | What to verify |
|---|---|---|
| Modern JavaScript-rendered content | Puppeteer | Wait for the application-specific content before printing. |
| Print CSS and detailed pagination controls | Puppeteer | Check page breaks, fonts, margins, colors, headers, and footers. |
| Existing output tied to wkhtmltopdf | Keep wkhtmltopdf during evaluation | Test the exact deployed binary and compare output compatibility. |
| Untrusted HTML | Neither is safe by default | Sanitize input and isolate the renderer with restricted network and filesystem access. |
| Latency, memory, or operating cost | Benchmark both in your target environment | Use the same host, OS image, fonts, inputs, concurrency, and resource limits. |
2. Convert a web page to PDF with Puppeteer
Install Puppeteer in a Node.js project, then save the following as save-page.js. The script opens a URL, waits for the page to load, waits for web fonts, and writes a PDF. Puppeteer’s Page.pdf() uses print media by default. The PDF API also supports paper size, margins, headers and footers, background graphics, page ranges, scaling, CSS page size priority, and other controls. Puppeteer Page.pdf documentation, PDF options reference
npm install puppeteer
// save-page.js
const puppeteer = require('puppeteer');
async function main() {
const url = process.argv[2] || 'https://example.com';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle2', timeout: 60000 });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
margin: { top: '18mm', right: '15mm', bottom: '18mm', left: '15mm' }
});
} finally {
await browser.close();
}
}
main().catch((error) => {
console.error(error);
process.exitCode = 1;
});
// Run: node save-page.js https://example.com
Wait for the content your application actually needs
networkidle2 can be a useful general wait condition, but applications with analytics, streaming requests, or long polling may never become idle. When the PDF depends on a known element, wait for that selector explicitly. A fixed delay is another option, but it is less reliable because it assumes a particular response time.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
await page.waitForSelector('[data-report-ready="true"]', { timeout: 30000 });
await page.evaluate(() => document.fonts.ready);
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });
Print media or screen media
Puppeteer uses print CSS for PDF output. Use @media print and @page in the site stylesheet when the PDF should have a print layout. If the PDF should preserve screen styling, emulate screen media before calling pdf(). Print colors may be adjusted by the browser; for exact colors, the Puppeteer documentation points to the CSS property -webkit-print-color-adjust.
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-style.pdf', printBackground: true });
/* In your stylesheet */
@media print {
.no-print { display: none !important; }
.report-section { break-inside: avoid; }
}
@page {
size: A4;
margin: 18mm;
}
html {
-webkit-print-color-adjust: exact;
}
Useful Puppeteer PDF options
| Option or feature | Use |
|---|---|
format |
Choose a named paper format such as A4 or Letter. |
width and height |
Set custom page dimensions instead of a named format. |
landscape |
Change page orientation. |
margin |
Set top, right, bottom, and left margins. |
printBackground |
Include CSS background graphics. |
displayHeaderFooter, headerTemplate, footerTemplate |
Add repeating page header and footer templates. |
pageRanges |
Render selected page ranges when appropriate. |
scale |
Scale page content to fit the selected layout. |
preferCSSPageSize |
Give CSS @page dimensions priority over the API format. |
| Font readiness | Wait for document fonts to load before generating output. |
Check the installed Puppeteer version’s API documentation for the exact option names and supported values in your deployment. PDFOptions reference
3. Convert a page with wkhtmltopdf
After installing a wkhtmltopdf build for your system, the simplest form is a URL followed by the destination PDF path:
wkhtmltopdf https://example.com page.pdf
A practical command can specify paper size, orientation, margins, and a JavaScript-controlled wait condition:
wkhtmltopdf \
--page-size A4 \
--orientation Portrait \
--margin-top 18mm \
--margin-right 15mm \
--margin-bottom 18mm \
--margin-left 15mm \
--window-status report-ready \
https://example.com/report \
report.pdf
The --window-status option waits for the page’s JavaScript to set the requested window status. For example, the page can run window.status = 'report-ready' after its report data has rendered. This is more targeted than guessing with a fixed delay, but it depends on page code running correctly in the renderer. wkhtmltopdf’s command documentation also describes page, cover, and table-of-contents objects, outlines, headers, and footers. Some features require the project’s patched Qt build; distro packages may omit those patches and behave differently. Test the exact binary or package you plan to deploy. wkhtmltopdf command-line usage
4. Which is better for HTML to PDF?
For a new workflow, Puppeteer is generally the better starting point when browser behavior, JavaScript execution, and print CSS are central to the output. The wkhtmltopdf project itself recommends considering Puppeteer or a wrapper for sites using dynamic JavaScript. That recommendation comes from the project’s status page; it is not a universal guarantee that Puppeteer will match every existing PDF. Project status and engine discussion
wkhtmltopdf may remain practical when an established system depends on its output, when its page/cover/TOC behavior is already integrated, or when migration introduces compatibility risk. Keep it only after confirming the engine limitations, exact build variant, security isolation, and output needs. A migration such as QFQ’s move toward Puppeteer shows one real path, not proof that every system should migrate. QFQ 25.6 documentation
5. JavaScript, CSS, and pagination edge cases
- Client-rendered content: Wait for the application’s ready marker or data element before printing. A navigation event only tells you about navigation progress, not that your app-specific content is complete.
- Lazy-loaded images: Scroll through the page or otherwise trigger the site’s loading behavior before printing, then wait for images and fonts. Long pages may have content that does not exist until scrolled into view.
- Print-only layout: Puppeteer uses print media by default. Validate print-specific display rules, page breaks, and hidden elements.
- Backgrounds and colors: Set
printBackground: truefor Puppeteer and use print color adjustment when exact colors matter. Verify the final PDF in the target viewer. - Page breaks: Use print CSS break properties and inspect pages containing large tables, images, or cards. Avoid relying on one browser’s pagination behavior to match another renderer exactly.
- Fonts: Install needed fonts in the container and wait for document fonts. Missing fonts can change line wrapping and pagination.
- Headers and footers: Puppeteer provides templates; wkhtmltopdf offers command options. Confirm page numbering and available space against the body margins.
- External assets: Ensure the renderer can resolve authenticated assets, redirects, and certificates. Browser navigation to the main URL does not guarantee every embedded resource will load.
6. Security: is wkhtmltopdf safe for user-submitted HTML?
No renderer should be treated as safe for arbitrary user-controlled HTML and JavaScript simply because it produces a PDF. The wkhtmltopdf downloads page gives an explicit warning: “Do not use wkhtmltopdf with any untrusted HTML – be sure to sanitize any user-supplied HTML/JS, otherwise it can lead to complete takeover of the server on which it is running!” wkhtmltopdf downloads page
For either renderer, treat conversion as a security boundary. Sanitize untrusted markup and scripts, run conversion in a separate process or container, restrict filesystem access, limit outbound network access, enforce CPU and memory limits, and set a timeout. Avoid passing secrets or broad credentials into a renderer that processes attacker-controlled content. These are deployment precautions; they do not imply that Puppeteer is automatically safe.
7. Performance, reliability, and operating cost
The researched sources provide no controlled comparison of speed, memory, fidelity, or total cost between Puppeteer and wkhtmltopdf. Do not infer a winner from engine age or feature lists. Benchmark the exact pages and deployment you care about.
- Choose representative pages: simple, long, JavaScript-heavy, image-heavy, and pages with complex print styles.
- Run both tools on the same host, OS image, fonts, network conditions, resource limits, and input URLs.
- Measure render latency, peak memory, CPU time, output size, failures, and differences in page count and visual layout.
- Repeat at the concurrency expected in production, including cold starts if relevant.
- Compare outputs with visual review or document-level checks for the fields and page breaks your workflow depends on.
Reliability usually depends on controlling the environment: pin the renderer version and OS image, provide required fonts, use explicit waits and timeouts, handle failed navigation and missing assets, and record conversion errors. Puppeteer adds a browser runtime to deploy and maintain. wkhtmltopdf’s behavior can vary by build because some features depend on patched Qt. Include package size, updates, process isolation, and operational support in the cost comparison. The sources do not establish a numeric cost advantage for either tool.
8. Migration checklist: wkhtmltopdf to Puppeteer
- Inventory page types, custom headers and footers, covers, tables of contents, JavaScript waits, and command-line flags.
- Record the exact wkhtmltopdf version and build variant currently deployed.
- Implement Puppeteer waits for real content readiness, font loading, and any required assets.
- Port paper size, margins, orientation, backgrounds, and pagination into PDF options and print CSS.
- Compare representative PDFs page by page, including long tables and pages with dynamic content.
- Test the security boundary and resource limits for untrusted or externally supplied content.
- Run the same workload at expected concurrency and compare latency, memory, failures, and operating cost.
- Roll out gradually with a way to identify mismatches and return to the existing renderer if needed.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF is missing client-rendered content | Printing begins before the application finishes rendering. | Wait for an app-specific selector or ready marker; increase the timeout only after adding a meaningful readiness condition. |
| Navigation times out on a page that appears usable | Persistent network requests prevent an idle condition. | Use a less restrictive navigation wait and then wait for the specific content needed by the PDF. |
| Fonts differ or pages reflow in deployment | Fonts are absent or not loaded when printing starts. | Install the fonts in the runtime and await document.fonts.ready. |
| Colors or backgrounds are missing | Print output omits backgrounds or print color adjustment changes colors. | Enable Puppeteer printBackground and set print color adjustment in CSS where exact colors matter. |
| wkhtmltopdf ignores an option or feature | The installed build lacks the required patched Qt features. | Check the exact binary/package and its documented build behavior; test the intended deployment artifact. |
| PDF pages break differently after migration | Different engines paginate CSS and content differently. | Review print CSS, fixed dimensions, large elements, and break rules; compare representative output rather than assuming pixel-identical conversion. |
| Conversion exposes server files or internal services | Untrusted content can access resources available to the renderer. | Sanitize input and isolate the conversion process with restricted network and filesystem access. |
| High memory use or slow jobs at concurrency | Browser processes and page complexity exceed deployment limits. | Measure peak memory and latency at target concurrency, cap parallel jobs, and set process-level resource limits. |
10. Or skip the browser setup
If your job is to capture a web page as an image or PDF rather than operate your own PDF-rendering runtime, ScreenshotNeo provides a website screenshot API and MCP server. For PDF rendering with a locally controlled Puppeteer workflow, keep using the code above; ScreenshotNeo’s API returns a screenshot or PDF from one request. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
# Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
// Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, with no card.
11. FAQ
Does wkhtmltopdf support modern JavaScript?
The project recommends considering Puppeteer or a wrapper for sites using dynamic JavaScript. Test your page in the exact wkhtmltopdf build you deploy; do not assume modern client-side rendering will complete correctly.
How do I print a Puppeteer page with CSS?
Use print stylesheets and @page rules for print output, which is the default media type for page.pdf(). Call page.emulateMediaType('screen') before PDF generation if you need screen styling.
Can Puppeteer produce exactly the same PDF as wkhtmltopdf?
Do not expect identical pagination or rendering across different engines. Compare the actual output and adjust print CSS and PDF options to meet your compatibility requirements.
Which tool should I choose if I cannot benchmark yet?
For a new JavaScript-heavy or print-CSS-driven workflow, begin with Puppeteer and validate representative PDFs before launch. For an established wkhtmltopdf workflow, retain the current path until output compatibility and deployment behavior have been checked.
