wkhtmltopdf vs Puppeteer for Website PDF Generation in India
Compare wkhtmltopdf and Puppeteer for website PDF generation in India, with runnable examples, print controls, deployment checks, and a practical migration guide.
Short answer: For a new website-to-PDF workflow, choose Puppeteer when your pages depend on modern browser CSS or JavaScript, or when you need configurable paper, margins, headers and footers, backgrounds, or page ranges. Keep wkhtmltopdf when an established, stable command-line conversion already meets your needs and you have consciously accepted its archived upstream status. There is no evidence in the reviewed official sources that India has a different technical winner; validate the languages, fonts, hosting, data handling, and compliance relevant to your deployment.
This is a recommendation based on documented capabilities and project lifecycle, not a measured speed, cost, or rendering benchmark. Both tools can produce PDFs; representative pages from your own application are the right basis for a migration decision.
What each tool does
wkhtmltopdf
wkhtmltopdf is a headless command-line tool that renders HTML to PDF or images using Qt WebKit. Its project overview describes uses such as invoice generation. Its upstream GitHub repository is archived and read-only, and the changelog’s 0.12.6 entry is dated 2020-06-11. That lifecycle status matters for new deployments: decide who owns dependency maintenance, binary pinning, security review, and eventual migration.
Puppeteer
Puppeteer automates a browser and provides Page.pdf() to generate PDF output. Its guide shows the basic sequence: launch a browser, navigate to a page, save a PDF, and close the browser. PDF generation uses print CSS media by default. Puppeteer’s documented PDF options include paper format and dimensions, margins, headers and footers, backgrounds, page ranges, scaling, timeouts, and font waiting. Check the API documentation for the Puppeteer version installed by your project.
Comparison at a glance
| Decision area | wkhtmltopdf | Puppeteer |
|---|---|---|
| Rendering approach | Headless command-line conversion using Qt WebKit. | Browser automation with a browser’s print-to-PDF path. |
| Good fit | An existing, stable conversion pipeline whose output is already acceptable. | A new workflow that relies on browser behavior, modern page content, or print configuration. |
| Print controls | Review the installed version’s usage documentation and validate every needed control. | Documented controls include paper, margins, headers and footers, backgrounds, page ranges, scaling, timeouts, and font readiness. |
| Upstream lifecycle | The upstream repository is archived and read-only. | Use the current project documentation and account for the browser runtime you deploy. |
| India-specific verdict | No distinct advantage established by the reviewed sources. | No distinct advantage established by the reviewed sources. |
| Performance and cost | Measure in your own environment; no fair head-to-head benchmark was established. | Measure in your own environment; no fair head-to-head benchmark was established. |
Choose based on the document you need
- Prefer Puppeteer for a new system if pages rely on contemporary browser CSS or JavaScript, or if you need the documented print controls exposed by
Page.pdf(). This is an inference from the tools’ documented designs, not a benchmark. - Keep wkhtmltopdf for a working legacy pipeline if its rendered output is acceptable, its dependencies are pinned and understood, and your team accepts responsibility for the archived upstream dependency.
- Run a comparison before switching if output fidelity, page breaks, operational cost, or reliability are important. Capture representative documents with both tools and compare the PDFs visually and operationally.
Generate a PDF with Puppeteer
The following Node.js example is a complete minimal script for printing a page to a PDF file. Install Puppeteer in your project with npm install puppeteer, save this as print.mjs, and run node print.mjs. Puppeteer manages a compatible browser installation for the package in the standard setup.
import puppeteer from 'puppeteer';
const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle0', timeout: 60_000 });
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
margin: { top: '16mm', right: '14mm', bottom: '16mm', left: '14mm' },
waitForFonts: true
});
} finally {
await browser.close();
}
Pass a different page as the first argument, for example node print.mjs https://your-site.example/report. The script uses the default print media behavior and asks Puppeteer to wait for fonts before producing output.
Configure print output
Here is a more configurable PDF call. Use the options supported by your installed Puppeteer version; the official PDF options reference is the definitive list for that version.
await page.pdf({
path: 'report.pdf',
format: 'A4',
landscape: false,
printBackground: true,
preferCSSPageSize: true,
displayHeaderFooter: true,
headerTemplate: '<div style="font-size:8px;width:100%;text-align:center">Report</div>',
footerTemplate: '<div style="font-size:8px;width:100%;text-align:center"><span class="pageNumber"></span> / <span class="totalPages"></span></div>',
margin: { top: '20mm', right: '12mm', bottom: '20mm', left: '12mm' },
pageRanges: '1-3',
scale: 1,
timeout: 60_000,
waitForFonts: true
});
formatselects a named paper size such as A4 or Letter. Alternatively, setwidthandheightto explicit dimensions.marginaccepts per-side dimensions. Leave enough room for any header or footer templates.printBackgroundincludes background graphics that print output may otherwise omit.displayHeaderFooterenables header and footer templates. Keep their markup self-contained and test page-number placeholders in the generated file.pageRangescan restrict output to selected pages. Confirm the resulting page count and range when content length changes.preferCSSPageSizelets CSS page-size declarations take precedence where appropriate. Use@pagerules when page layout belongs in the document stylesheet.scaleadjusts output scale. Avoid using it to hide layout problems without checking text size and page breaks.timeoutandwaitForFontsaffect readiness. A timeout cannot make a page’s data or scripts available if the page never finishes loading.
To render using screen CSS instead of the default print media, set screen media before generating the PDF:
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-styled.pdf', format: 'A4', printBackground: true });
Use print media for conventional documents. Use screen media only when the intended PDF should reproduce screen styling. Check color output, page breaks, and background handling with real pages.
Load dynamic content deliberately
The example waits for networkidle0, which can be unsuitable for pages that keep connections open or continuously poll. When the page has a reliable application-level completion signal, wait for that signal instead:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.waitForSelector('[data-report-ready="true"]', { timeout: 30_000 });
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });
Replace the selector with an element your application sets only after the content needed in the PDF is ready. If no such signal exists, wait for a known element or a bounded delay and verify the result; avoid an unbounded wait.
Convert a URL with wkhtmltopdf
Install a wkhtmltopdf package or binary appropriate to your operating system, then run the command-line converter. This example uses A4 paper, includes background graphics, sets margins, and prints a URL to a file:
wkhtmltopdf --page-size A4 --margin-top 16mm --margin-right 14mm --margin-bottom 16mm --margin-left 14mm --background https://example.com page.pdf
For an existing installation, consult the wkhtmltopdf usage documentation for options actually supported by that build. Do not assume command-line options or rendering behavior are identical across package builds. Since upstream is archived, record the exact binary and environment that produce accepted output.
How to evaluate an India deployment
The sources do not establish an India-specific winner, local price difference, or legal mandate. Treat India as a deployment context to validate against your actual application and requirements, rather than as a reason to assume either renderer wins.
- Assemble representative pages. Include the longest document, tables near page boundaries, dynamic content, charts, and pages with headers and footers.
- Check scripts and fonts. Use real content in every language your service supports, including the relevant Indic scripts. Verify glyph coverage, shaping, fallback fonts, line wrapping, and whether fonts are available in the deployed runtime.
- Check page layout. Review paper size, margins, page breaks, background graphics, and clipped or split content. Compare the resulting PDFs, not just browser screenshots.
- Test network and data boundaries. Confirm which page resources the renderer can reach, where rendering runs, what document data is present in that environment, and how your hosting and data-handling requirements apply.
- Measure the actual workload. Record latency, memory, concurrency, cold starts, container size, and failure behavior on the target deployment.
- Document ownership. For wkhtmltopdf, explicitly account for the archived upstream repository and the team’s plan for pinned binaries, review, and migration. For Puppeteer, account for operating the browser runtime and keeping it aligned with your application.
These are validation steps, not claims that either tool fails a particular Indian language, hosting setup, or compliance requirement.
Migration checklist
- Choose a fixed set of representative URLs and capture their current PDFs as baselines.
- Pin tool versions and record the browser or binary, operating system, installed fonts, and relevant configuration.
- Implement the same paper size, margins, background behavior, page ranges, and readiness condition in the candidate renderer.
- Compare page counts, text wrapping, fonts, page breaks, headers and footers, and visual differences.
- Run concurrent jobs under expected load and record memory, latency, cold-start behavior, and failed conversions.
- Decide how to handle navigation errors, timeouts, retries, temporary files, and incomplete output.
- Roll out gradually and keep a rollback path until the new PDFs meet your acceptance criteria.
Performance, reliability, and cost
The reviewed official sources do not provide a fair, current head-to-head benchmark or an India-specific cost comparison. Do not infer that one is faster or cheaper from the tool names or rendering approach. Measure using the same pages, output settings, hosting resources, concurrency, and failure policy.
- Performance: Measure per-document latency, memory, concurrency, startup time, and deployment size. Browser startup and reuse policy can affect Puppeteer measurements; include the pattern you plan to operate.
- Reliability: Use bounded navigation and PDF timeouts, detect navigation failures, and verify a PDF was produced before returning it. Retry only failures that are plausibly transient, with a limit and backoff.
- Cost: Include compute, memory, container storage, operational maintenance, dependency updates, and migration work. The research does not support a numerical cost verdict.
- Security and maintenance: Review the deployed binaries and browser, restrict access to sensitive pages and network resources as appropriate, and define who maintains the rendering dependency. The archived wkhtmltopdf repository makes upstream maintenance ownership an explicit decision.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| PDF is blank or missing page content | Navigation returned before client-side content was ready, or the page failed to load. | Check navigation errors and response status; wait for a meaningful ready selector or bounded application signal before printing. |
| Fonts or glyphs are missing | The runtime lacks the font, font loading has not completed, or fallback differs from development. | Install and pin required fonts in the deployment, verify script coverage, and wait for fonts before PDF generation. |
| Background colors or images are absent | Background printing is disabled or print CSS suppresses the element. | Enable printBackground in Puppeteer or --background in wkhtmltopdf, and inspect the print stylesheet. |
| Content is cut off or page breaks look wrong | Paper dimensions, margins, print CSS, or scaling do not fit the content. | Set explicit paper and margins, inspect @page and break rules, then compare page count and output at normal scale. |
| Puppeteer navigation times out | The page is slow, waits on ongoing network activity, or never reaches the chosen lifecycle condition. | Use a bounded timeout and a more suitable readiness condition; wait for an application selector when possible. |
| Header or footer overlaps the document | Insufficient top or bottom margin, or a template exceeds available space. | Increase the relevant margin and simplify the template; test both short and long documents. |
| Output differs between local and production | Different browser or binary, fonts, operating system, network access, or runtime configuration. | Pin and record the environment; reproduce using the production image and the same input page. |
| wkhtmltopdf output is unsuitable for a page | The page depends on rendering behavior that does not match the installed Qt WebKit based converter. | Test the actual page against Puppeteer or retain the old path only for pages whose output remains acceptable. |
Or skip the browser setup
If your task is to capture a website as an image or PDF rather than operate a renderer, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. The API returns a PNG, JPEG, WebP, or PDF from one GET request. See the ScreenshotNeo API documentation for parameters and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners are accepted like a visitor and removed before the shot, along with known newsletter popups and chat widgets; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Is wkhtmltopdf discontinued?
The upstream GitHub repository is archived and read-only. That is the relevant lifecycle fact for planning; check your own package source and deployment before making assumptions about a particular binary.
Does Puppeteer always wait for fonts?
The Puppeteer PDF guide says font loading is awaited by default. You should still ensure the required fonts exist in the runtime and inspect glyphs in the output.
Does India require one of these tools?
The reviewed sources establish no India-specific technical winner or legal mandate. Apply your own hosting, language, data-handling, and compliance requirements to the deployment.
Can I compare them with a single test page?
A single page can reveal obvious differences, but it cannot cover scripts, page breaks, dynamic content, load failures, and workload behavior. Use a small representative document set for a decision.
