ScreenshotNeo

BlogComparisons

DocRaptor vs. wkhtmltopdf for Webpage PDFs

Compare DocRaptor’s hosted HTML-to-PDF API with the self-managed wkhtmltopdf CLI. Choose based on your PDF features, operations, security needs, and costs.

By the ScreenshotNeo team4 October 202612 min read

DocRaptor and wkhtmltopdf both turn HTML into PDFs, but they put the work in different places. DocRaptor is a hosted API: send it HTML or a URL and receive a PDF. wkhtmltopdf is a command-line program that you install and operate in your own environment. Choose DocRaptor when you want a managed conversion service or need its documented PDF capabilities, such as accessible tagging, mixed page layouts, or forms. Choose wkhtmltopdf when a self-run CLI fits your workflow and you can account for its archived project status and the security responsibilities of running it.

Neither choice is a universal winner. Rendering engines, deployment responsibilities, required PDF features, security boundaries, and the cost of your actual workload matter more than a general claim that one produces better PDFs. The sources cited here do not provide an apples-to-apples comparison of cost, speed, or rendering fidelity.

At a glance

Question DocRaptor wkhtmltopdf
How do you use it? Submit HTML content or a URL to a hosted API. Run a headless command-line tool in your environment.
Rendering engine Prince PDF engine; the API documents versioned pipelines. Qt WebKit.
Notable documented features Accessible PDF tagging, mixed page layouts, and PDF forms. Forms use Prince-specific CSS. CLI controls include headers, footers, page numbering, and JavaScript settings.
Maintenance signal Vendor documentation describes configurable pipelines and their engine versions. GitHub repository archived January 2, 2023; the project downloads page lists 0.12.6 as the stable series, released June 11, 2020.
Who operates conversion? The service handles the conversion endpoint; your application still needs to submit jobs and handle results. Your team installs, runs, monitors, and updates the executable and its environment.
Cost evidence Check current plan terms for your expected volume and usage. The project is LGPL-3.0, but licensing does not remove infrastructure, maintenance, or engineering costs.

These are documented differences, not results from a benchmark. DocRaptor describes its API and engine in its API reference and API overview. The wkhtmltopdf project describes its tool and status in its repository and downloads page.

How the operating models differ

DocRaptor: hosted conversion API

Your application sends a document’s HTML or a URL to DocRaptor. Its API documentation describes synchronous binary PDF responses, as well as hosted or asynchronous response patterns. You integrate an HTTP service and decide how your application handles authentication, request failures, retries, and returned documents. The exact request and response options are documented in the API reference.

DocRaptor says it uses the Prince PDF engine. Its API reference documents pipelines that map to Prince and JavaScript engine versions, with pipeline 10.1 listed as the default when the pipeline parameter is omitted. Pipeline behavior is version-sensitive; confirm the current reference and select a pipeline deliberately if consistent output matters to your application.

wkhtmltopdf: self-managed command line

wkhtmltopdf runs as a headless command-line tool and uses Qt WebKit, according to its project repository. Your application can invoke it as a process and read the PDF it writes. You also own its installation, environment, process lifecycle, resource limits, and operational response to failures.

The repository is archived and read-only as of January 2, 2023. The project’s downloads page lists 0.12.6 as its stable series, released June 11, 2020. Treat those as material maintenance facts when deciding whether to put it in a new production system. Check the package and platform you intend to deploy rather than assuming every distribution has identical behavior.

Documented PDF features to compare

When DocRaptor’s documented capabilities matter

DocRaptor’s vendor materials describe support for accessible PDF tagging and mixed page layouts. Its forms guide documents fillable PDF forms using the Prince-specific CSS property -prince-pdf-form; the guide notes that this functionality is not part of the HTML standard. These are vendor-documented capabilities, not independent comparative test results. Confirm that the feature fits your exact document and downstream reader requirements.

  • Accessible tagging: verify the output against your accessibility requirements and validation process.
  • Mixed page layouts: useful to investigate when one document needs different page orientations or layouts.
  • PDF forms: DocRaptor documents form creation using Prince-specific CSS, so the markup is not portable standard HTML behavior.
  • JavaScript: the API reference documents JavaScript processing options and says they are disabled by default. Confirm the setting and pipeline behavior for your content.

When wkhtmltopdf’s CLI controls matter

The wkhtmltopdf usage reference documents controls for PDF generation including headers, footers, page numbering, and JavaScript settings. Consult the usage reference for available switches and their syntax. These options establish that controls exist; they do not establish compatibility with modern browser features or parity with DocRaptor.

For either option, build a representative fixture set before committing: include your actual CSS, fonts, images, page breaks, long tables, and any scripts your documents depend on. Compare the resulting PDFs against your requirements. No source cited here establishes which renderer will be more faithful for your pages.

Runnable examples

These examples show the basic integration shape. Replace credentials and input content, and consult each product’s current documentation for optional fields, authentication details, and response behavior. Avoid putting credentials directly in source control.

DocRaptor with cURL

curl https://docraptor.com/docs -X POST \
  -u YOUR_API_KEY: \
  -H "Content-Type: application/json" \
  -d '{"type":"pdf","document_content":"<html><body><h1>Invoice</h1></body></html>","name":"invoice.pdf"}' \
  -o invoice.pdf

Use the endpoint, authentication format, request fields, and response mode specified by the current DocRaptor API reference. The example illustrates a synchronous binary response; production integrations should check HTTP status and handle API errors before treating the output as a PDF.

DocRaptor with Python

import requests

api_key = "YOUR_API_KEY"
payload = {
    "type": "pdf",
    "document_content": "<html><body><h1>Invoice</h1></body></html>",
    "name": "invoice.pdf",
}

response = requests.post(
    "https://docraptor.com/docs",
    auth=(api_key, ""),
    json=payload,
    timeout=120,
)
response.raise_for_status()
with open("invoice.pdf", "wb") as pdf_file:
    pdf_file.write(response.content)

DocRaptor with Node.js

const apiKey = process.env.DOCRAPTOR_API_KEY;
if (!apiKey) throw new Error('Set DOCRAPTOR_API_KEY');

const response = await fetch('https://docraptor.com/docs', {
  method: 'POST',
  headers: {
    'Authorization': `Basic ${Buffer.from(`${apiKey}:`).toString('base64')}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    type: 'pdf',
    document_content: '<html><body><h1>Invoice</h1></body></html>',
    name: 'invoice.pdf',
  }),
});
if (!response.ok) {
  throw new Error(`DocRaptor returned HTTP ${response.status}: ${await response.text()}`);
}
const pdf = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('invoice.pdf', pdf));

For production, use the current API docs to verify endpoint and payload details, and use the API’s supported asynchronous pattern if a request should not hold an application connection open.

wkhtmltopdf from the command line

After installing a compatible package for your environment, a basic URL-to-PDF invocation is:

wkhtmltopdf https://example.com page.pdf

To render a local HTML file:

wkhtmltopdf ./page.html page.pdf

A simple header, footer, and page numbering example is:

wkhtmltopdf \
  --header-center "Report" \
  --footer-center "Page [page] of [topage]" \
  https://example.com report.pdf

Options and placeholder support depend on the installed build. See the project’s usage reference. Treat options that enable JavaScript as an explicit decision: wait behavior and script execution can affect whether content has finished rendering before conversion.

Calling wkhtmltopdf from Python

import subprocess

subprocess.run(
    ["wkhtmltopdf", "https://example.com", "page.pdf"],
    check=True,
    timeout=120,
)

check=True makes a nonzero exit status an error; the timeout prevents the caller from waiting indefinitely. In a service, also capture diagnostic output and decide how to clean up partial output after a failed process.

Calling wkhtmltopdf from Node.js

import { execFile } from 'node:child_process';
import { promisify } from 'node:util';

const execFileAsync = promisify(execFile);
const { stdout, stderr } = await execFileAsync(
  'wkhtmltopdf',
  ['https://example.com', 'page.pdf'],
  { timeout: 120_000, maxBuffer: 10 * 1024 * 1024 },
);
console.log(stderr);

Use an argument array rather than building a shell command from user input. Handle process errors and timeouts, and make sure the destination path is controlled by your application.

Security and trust boundaries

The wkhtmltopdf project’s downloads page warns: “Do not use wkhtmltopdf with any untrusted HTML – be sure to sanitize any user-supplied HTML/JS, otherwise it can lead to complete takeover of the server it is running on!” If your input includes user-submitted or otherwise untrusted HTML, treat that warning as a serious design constraint. Sanitization and safe execution need to be part of your system design; the project warning itself does not specify a complete mitigation plan.

Do not infer from this warning that the dossier establishes a comprehensive security comparison for DocRaptor. It does not. Assess the data you send to any hosted API, the trust boundary around your content, and each provider’s current security and data-handling terms. For self-managed conversion, examine what the renderer process can access in its environment and how your application limits its inputs and execution. The sources here are not a security audit of either option.

Performance, reliability, and cost

Performance

There is no verified speed benchmark here, so do not choose based on an assumed throughput or latency advantage. Measure your own documents and workload. Include small and large pages, script-heavy content, remote assets, and peak concurrency. Track end-to-end time, failures, output size, and any queueing or process contention that matters to your application.

Reliability

With a hosted API, your integration depends on network access and the service’s current API behavior. Design for failed requests and timeouts, and consult the documented synchronous or asynchronous response options. With a CLI, reliability also depends on your installed package, runtime environment, process management, and ability to respond to failures. In both cases, retain enough diagnostics to identify the input and stage that failed without exposing secrets.

For either approach, decide how to handle retries before production. Retrying a transient network failure may be reasonable; blindly retrying every invalid document or deterministic rendering error can add load without fixing the cause. Make output writes atomic or otherwise ensure a partial file is not mistaken for a completed PDF.

Cost

The available sources do not support a current, apples-to-apples total-cost comparison. DocRaptor’s current plan terms should be checked directly for your volume and requirements; do not treat an undated search-result price as current. wkhtmltopdf’s LGPL-3.0 license does not make operation cost-free: include compute, deployment, maintenance, and engineering time in your estimate. Compare cost per successfully generated document under your real workload, including operational work and failures.

Decision checklist

  1. List required PDF features. Identify whether you need accessible tagging, mixed page layouts, forms, headers and footers, page numbers, or JavaScript-driven content.
  2. Choose who operates conversion. Decide whether a hosted API suits your system or whether you want to own the CLI runtime and its ongoing care.
  3. Review project status. For wkhtmltopdf, account for the archived repository and the stable release listed by the project.
  4. Map trust boundaries. Identify whether HTML or JavaScript is user-controlled and determine how each architecture handles that risk.
  5. Validate representative documents. Generate PDFs from your own fixtures and inspect the required visual, accessibility, and form behavior.
  6. Estimate your real costs. Check current service terms and include self-managed infrastructure and maintenance effort.

In short: choose DocRaptor if a hosted conversion API and its documented Prince-based PDF capabilities match your requirements. Consider wkhtmltopdf if running a CLI yourself fits your deployment and you accept the project’s maintenance status and security responsibilities. Validate both against the actual documents and operating constraints that matter to you.

Or skip the browser setup

If the job is to capture a webpage as an image or PDF, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It is an alternative to try first when you want a screenshot workflow rather than configuring a browser capture stack. It is not a substitute for a full HTML-to-PDF conversion service when you need document-specific features such as fillable forms or accessible PDF tagging.

One GET request can return a screenshot or PDF. For example, this cURL request saves a WebP screenshot of a webpage:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same endpoint can be called from Python:

import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Or Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options and response details. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Sign up for the free plan.

Troubleshooting

Symptom Likely cause What to check
DocRaptor request is rejected Authentication, endpoint, payload, or document settings do not match the current API requirements. Check the response status and error body, then verify credentials and fields against the current API reference.
PDF response is missing or treated as text The integration assumes a successful binary PDF response when the request failed or uses a different response pattern. Check HTTP status before saving bytes; review the selected synchronous or asynchronous API flow.
Content rendered before JavaScript completed Script processing may be disabled by default or the chosen pipeline/settings may not match the page. Review the DocRaptor JavaScript options and pipeline in the current reference. For wkhtmltopdf, consult its JavaScript and related CLI switches; do not assume behavior matches a current desktop browser.
wkhtmltopdf executable is not found The binary is absent or not on the process PATH. Install the intended package in the runtime environment and verify the executable path and permissions.
wkhtmltopdf exits with an error Invalid arguments, inaccessible input, runtime dependencies, or a rendering failure. Capture stderr and exit status; verify the URL or file is reachable from the process environment and confirm option syntax in the usage reference.
Output omits images, styles, or fonts Assets may not be reachable from the converter, or the document may rely on resources that are not ready when capture starts. Check asset URLs, access requirements, and load timing from the converter’s environment. Test a self-contained fixture to isolate the issue.
Conversion hangs or times out A slow or unreachable resource, script behavior, or a stuck process/request may prevent completion. Set an application-level timeout, capture diagnostics, and inspect external resource dependencies. Use the documented asynchronous API pattern where appropriate for DocRaptor.
Unexpected page breaks or layout Renderer engines can interpret document styling differently; source material does not establish cross-engine parity. Reduce the document to a reproducible fixture and validate the required layout in the selected engine. Avoid assuming a CSS feature works identically across renderers.
Partial or corrupt output file The process or request failed after the destination was created. Only publish the output after success and validation; remove or replace partial files on failure.

FAQ

Is wkhtmltopdf still maintained?

The GitHub repository is archived and read-only, dated January 2, 2023. The project downloads page lists 0.12.6 as its stable series, released June 11, 2020. Check the package you plan to use and weigh that status against your maintenance needs.

Does DocRaptor use a browser engine?

DocRaptor says it uses the Prince PDF engine. Its API documentation also describes pipelines mapped to Prince and JavaScript engine versions; it should not be described as equivalent to a current desktop browser on the evidence cited here.

Can wkhtmltopdf safely render user-submitted HTML?

The project explicitly warns against using it with untrusted HTML unless supplied HTML and JavaScript are sanitized. Treat user-controlled content as a security boundary and design accordingly.

Which one is cheaper?

The cited sources do not establish an apples-to-apples cost. Check current DocRaptor terms and compare them with the infrastructure and engineering costs of operating wkhtmltopdf for your workload.