wkhtmltopdf Alternatives That Support JavaScript-Heavy Websites
Compare JavaScript-capable wkhtmltopdf alternatives, including Puppeteer, Playwright, Gotenberg, and WeasyPrint, and learn how to choose and configure a reliable PDF workflow.
Short answer: For pages whose content appears only after JavaScript runs, use Puppeteer or Playwright with Chromium, or deploy Gotenberg as a Docker-based HTTP service. Each uses a browser engine capable of running page scripts. WeasyPrint is useful for print-focused, static or pre-rendered HTML, but its documentation describes its own Python CSS layout engine rather than a full browser engine, so do not assume it executes client-side JavaScript. Choose by testing your real pages and PDF requirements.
wkhtmltopdf renders HTML to PDF and images using Qt WebKit. Its GitHub repository was archived on January 2, 2023 and is read-only. That repository status is a reason to review your dependency, but it does not by itself establish a specific security flaw. Project site · Repository status.
Which alternative should you choose?
| Option | Choose it when | What you operate |
|---|---|---|
| Puppeteer + Chromium | You want a programmable browser API in JavaScript or Node.js and need browser behavior and PDF output. | Your application manages browser installation, process lifecycle, versions, and workload testing. |
| Playwright + Chromium | Your team already uses Playwright or prefers its browser automation framework. | Your application manages the browser and should validate PDF behavior in the selected browser. |
| Gotenberg | You want an HTTP conversion endpoint that your team can deploy as a container. | You deploy and operate the service and its dependencies. |
| WeasyPrint | Your source is static or pre-rendered HTML and print-focused CSS layout is the priority. | Your application and deployment manage the Python renderer, fonts, and CSS compatibility. |
| Managed HTML-to-PDF API | You prefer not to operate browser infrastructure. | Check each vendor’s JavaScript support, privacy, reliability, pricing, and contract terms directly. |
Puppeteer is documented as browser automation for Chrome and Firefox, including PDF generation. Playwright’s Page API documents browser page operations and PDF generation; check its current API and browser limitations when choosing an engine. Gotenberg’s quick start shows a Docker service and a Chromium URL-to-PDF route. These facts identify capabilities, not universal performance or fidelity rankings.
Sources: Puppeteer documentation, Puppeteer Page.pdf(), Playwright Page API, Gotenberg quick start, WeasyPrint documentation.
Before migrating: confirm what the page needs
- Check whether the content is rendered in the browser. If important text or data appears only after scripts run, prioritize Chromium automation or a Chromium-backed service.
- Define readiness. Decide what proves the page is ready: a specific selector, an application event, or a stable network state. A navigation response alone may arrive before client-side rendering finishes.
- Define the PDF contract. Record paper size, margins, orientation, page breaks, headers and footers, font availability, background graphics, and whether the output should reflect print or screen CSS.
- Choose the operating model. A browser library gives your application direct control; Gotenberg gives you a containerized endpoint you operate; a managed API can shift infrastructure work to a vendor, subject to its terms.
- Build a representative test set. Include long pages, slow scripts, lazy content, unusual fonts, authentication, and pages that sometimes fail. Compare output and failure behavior in your own environment.
Puppeteer: generate a PDF after client-side content is ready
Install Puppeteer in a Node.js project with npm install puppeteer. The following CommonJS script navigates to a page, waits for a meaningful element, and writes a PDF. Replace the URL and readiness selector with values from your application.
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
page.setDefaultNavigationTimeout(45_000);
page.setDefaultTimeout(15_000);
await page.goto('https://example.com/report', {
waitUntil: 'domcontentloaded',
timeout: 45_000
});
await page.waitForSelector('[data-report-ready="true"]', {
visible: true
});
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' }
});
} finally {
await browser.close();
}
})();
Save it as render.cjs and run node render.cjs. Puppeteer’s PDF output uses print CSS media by default. If the page is styled for screen instead, emulate screen media before producing the PDF; then verify the result. Print color adjustment may be needed when exact colors matter. Refer to the Page.pdf() options for the current supported settings.
// Optional: use screen CSS instead of the default print media.
await page.emulateMediaType('screen');
// Then produce the PDF with the same page.pdf(...) call.
Use one readiness strategy that reflects the site rather than stacking arbitrary waits. If there is no stable ready selector, an application-specific signal is preferable; a fixed delay is simple but can be wasteful on fast pages and insufficient on slow ones. Network-idle conditions can also be unsuitable for pages with long-lived requests. Verify the choice against the target page.
Playwright: use browser automation with explicit readiness
Install Playwright and its Chromium browser with npm install playwright and npx playwright install chromium. This runnable Node.js example waits for a page-specific selector and saves a PDF.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/report', {
waitUntil: 'domcontentloaded',
timeout: 45_000
});
await page.locator('[data-report-ready="true"]').waitFor({
state: 'visible',
timeout: 15_000
});
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' }
});
} finally {
await browser.close();
}
})();
Run with node render.js. Check the current Page API for PDF options and browser support. Do not assume PDF output is identical across engines; use the engine and version you intend to deploy and compare the resulting files.
Gotenberg: expose Chromium conversion over HTTP
Gotenberg is a Docker-based document conversion API. Its Chromium URL conversion endpoint is /forms/chromium/convert/url. Start with the official quick start, then send a form request containing the target URL. For example, with a locally running service on port 3000:
curl --fail --show-error \
--form url=https://example.com/report \
http://localhost:3000/forms/chromium/convert/url \
--output report.pdf
This makes conversion an HTTP operation, but the team still has to deploy and operate the container and its dependencies. Check Gotenberg’s current endpoint documentation for supported form fields and PDF settings. Confirm how the page’s scripts signal readiness and test authentication, network access, and timeouts in the deployment environment.
WeasyPrint: a print renderer for suitable HTML
WeasyPrint is a Python HTML/CSS rendering engine with its own CSS layout system, aimed at document and print output. It can be a good fit for static or pre-rendered HTML. Its documentation does not establish client-side JavaScript execution, so pages that build their content in JavaScript need a separate rendering step or a browser-based option.
For a simple local HTML file, install WeasyPrint according to its installation documentation, then run:
weasyprint report.html report.pdf
Validate CSS support, fonts, page breaks, and assets against your real documents. A successful conversion does not show that a JavaScript-dependent page was rendered; inspect the PDF content and compare it with the browser page.
PDF settings that commonly change the result
| Concern | What to decide and verify |
|---|---|
| Print versus screen CSS | Browser PDF generation commonly targets print media. Puppeteer documents print media as its default. Emulate screen only if that matches the intended document. |
| Paper size and CSS page size | Choose an explicit format or let CSS page rules control size; check how your selected API resolves conflicts. |
| Margins and page breaks | Set margins deliberately and test long tables, headings near page boundaries, and forced breaks. |
| Backgrounds and colors | Enable background printing where required and inspect color adjustment, since print rendering may alter colors. |
| Fonts and assets | Ensure fonts and images are reachable from the renderer and fully loaded before conversion. Compare line wrapping and page count. |
| Long or dynamic pages | Wait for lazy content and application rendering before PDF generation; define how much content and how many pages are acceptable. |
Or skip the browser setup
If you need a clean image capture of a page as part of a workflow, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It is for screenshots, not a drop-in PDF replacement. One GET request can return PNG, JPEG, WebP, or a PDF:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/report \
-o report.pdf
See the ScreenshotNeo API documentation for request options. Its clean-shot flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
The free plan includes 1,000 shots a month with no card. Paid plans start at $5 for 3,000 shots; all features are on every plan. Try ScreenshotNeo free: get 1,000 screenshots a month with no card.
cURL, Python, and Node.js for a self-operated browser
There is no universal cURL command that runs JavaScript locally: cURL makes HTTP requests but does not execute a browser page. To use cURL in a self-operated workflow, call an HTTP service such as Gotenberg. For a Python application, one practical pattern is to call a Node.js browser worker; the worker can use the Puppeteer or Playwright examples above. The following Python client submits a URL to a service endpoint and saves the returned PDF. It assumes that endpoint accepts a url form field and returns PDF bytes; adjust the endpoint and authentication to your deployment.
import requests
response = requests.post(
"http://localhost:3000/forms/chromium/convert/url",
files={"url": (None, "https://example.com/report")},
timeout=90,
)
response.raise_for_status()
with open("report.pdf", "wb") as output:
output.write(response.content)
For ScreenshotNeo’s equivalent one-call request in Python or Node.js, use the documented API examples and replace the target URL. These return a screenshot response in the requested format; choose PDF output when that is the desired artifact and configure the relevant parameters in the docs.
# Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/report"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
// Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/report'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Reliability, performance, and cost
- Measure on your own pages. Compare readiness time, total conversion time, memory, throughput, output fidelity, and failure rate in the environment you will operate. The available evidence does not establish a universal performance winner.
- Control browser lifecycle. Browser libraries make your application responsible for process management, browser versions, and workload testing. Decide whether to reuse browser processes or isolate jobs based on measured behavior and your failure model.
- Set bounded timeouts. Bound navigation, readiness waits, and service requests. Record whether a failure was navigation, readiness, or PDF generation so retries do not conceal a persistently broken page.
- Retry selectively. A transient network failure may merit a bounded retry; a missing selector or invalid page usually needs investigation. Avoid unbounded parallel browser jobs.
- Account for deployment costs. Self-hosted choices require infrastructure and operational time; a managed API has vendor-specific pricing and terms. This research does not establish comparable pricing or a cost ranking, so verify current costs and model volume, concurrency, and document size.
- Protect sensitive content. Review where authenticated pages, cookies, headers, and generated PDFs travel and persist. Check the service’s privacy and retention terms before sending confidential documents.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF is blank or missing dynamic content | Conversion started before the client-side app rendered, or scripts failed. | Wait for a page-specific ready selector or application signal; inspect browser console and network failures. |
| Navigation timeout | The page is slow, blocked, or keeps connections open; the chosen load condition may never settle. | Use a suitable navigation condition, set a bounded timeout, and wait separately for the content you need. Investigate access restrictions and failed requests. |
| PDF differs from the browser view | Print CSS is active, page size or margins differ, or fonts/backgrounds are missing. | Check print versus screen media, configure page dimensions and margins, enable backgrounds where needed, and verify font and asset loading. |
| CSS layout breaks across pages | Print rules, page size, margins, or content length interact differently than screen layout. | Add and test print-specific page-break rules; validate representative long documents at the target paper size. |
| Images or lazy content are absent | Capture began before images loaded or lazy content entered the viewport. | Wait for the relevant image or content selectors and ensure the page has triggered the content load before rendering. |
| Browser fails to launch in deployment | Browser installation, runtime dependencies, or deployment permissions differ from local development. | Install the browser and dependencies using the selected framework’s deployment instructions and reproduce the issue in the same container or host environment. |
| WeasyPrint output lacks interactive-page content | WeasyPrint is a document-oriented HTML/CSS renderer, not a full browser JavaScript runtime. | Pre-render the HTML or use a browser-based renderer for pages whose content depends on client-side scripts. |
| Gotenberg request returns an error | The service may be unavailable, the route or form fields may be wrong, or the target URL may not be reachable from its container. | Check the service logs and current endpoint documentation; verify container network access and the exact route and form input. |
Migration checklist
- Inventory URLs and identify which depend on JavaScript, authentication, or delayed data.
- Choose the renderer and browser version, then pin and document them for repeatable output.
- Define page size, margins, print or screen media, backgrounds, and readiness conditions.
- Compare a representative set of PDFs for missing content, layout, fonts, and page breaks.
- Set timeouts, concurrency limits, logs, and selective retry behavior.
- Measure throughput, memory, latency, and failures on your deployment; estimate infrastructure or vendor costs using actual workload data.
- Keep a rollback path until the new output meets the documents’ acceptance criteria.
FAQ
Does wkhtmltopdf run JavaScript?
wkhtmltopdf uses Qt WebKit, but this does not make it equivalent to a current full browser automation workflow. For pages whose output depends on modern client-side behavior, validate a Chromium-based alternative against the actual site.
Is Puppeteer faster than Playwright for PDF generation?
The cited documentation does not provide a comparable benchmark. Measure both with your pages, browser versions, and deployment settings if speed is a deciding factor.
Can I use WeasyPrint for a JavaScript-heavy site?
Only if the needed content is already present in the HTML it receives or has been pre-rendered. Its documented renderer is its own Python CSS layout engine, not a full browser engine.
Is Gotenberg managed hosting?
The documented setup is a Docker-based API. The team deploying it operates that service; confirm any third-party hosting arrangement separately.
Which option is best for every site?
There is no evidence-backed universal winner. Match the rendering model to the page, define output requirements, and compare candidates on representative documents.
