ScreenshotNeo

BlogComparisons

PDFCrowd vs. wkhtmltopdf for Converting HTML to PDF

Compare PDFCrowd’s hosted API with wkhtmltopdf’s self-managed CLI, including inputs, JavaScript, layout controls, operations, and how to choose.

By the ScreenshotNeo team4 October 20269 min read

PDFCrowd is a hosted HTML-to-PDF API; wkhtmltopdf is an open-source command-line converter that your team runs. Choose PDFCrowd when its hosted HTTP workflow and documented PDF controls fit your needs and you accept sending conversion requests to a third party. Choose wkhtmltopdf when self-managed deployment and its LGPLv3 license fit your environment, and representative pages render correctly with its engine. Neither is a universal winner for fidelity, speed, security, or price.

This comparison is based on the products’ documentation, not a benchmark or a claim that either tool was tested here. Check current service terms, limits, and supported output requirements before adopting either option.

At a glance

Decision PDFCrowd wkhtmltopdf
Operating model Hosted API over HTTP, with official client libraries. Binary or source build run as a command-line process.
Inputs URL, HTML text, or uploaded HTML file; local assets can be packaged with HTML in a supported archive. URL or document, with command-line controls for resources, cookies, headers, and local-file access.
JavaScript Product overview lists custom JavaScript and readiness controls. JavaScript is enabled by default; the manual documents a default 200 ms delay and controls to disable JavaScript or set a delay.
Layout Page size, orientation, margins, page breaks, headers, footers, page numbers, print CSS, and custom CSS. Page and margin settings, print media selection, headers, footers, page numbering, and other CLI options.
Special PDF requirements Documents watermarks, password protection, PDF/A, and tagged PDF output. The reviewed manual establishes common conversion options, not parity for those requirements.
Operations and licensing Provider runs the renderer; request rate and concurrency depend on the license. Your team runs and maintains the converter and its runtime. Project identifies the license as LGPLv3.

Sources: PDFCrowd product overview, PDFCrowd HTTP documentation, PDFCrowd parameter documentation, wkhtmltopdf project, and wkhtmltopdf manual.

How the operating models differ

PDFCrowd: managed conversion over HTTP

Your application sends a conversion request to PDFCrowd and receives the resulting PDF. The HTTP guide describes POST requests to a versioned endpoint, authenticated with a PDFCrowd username and API key using HTTP Basic authentication. That removes the need for your application to install and operate the converter, while making the request dependent on the hosted service and its current account limits.

For URL conversion, the target and its resources must be reachable from PDFCrowd’s servers. For local assets, the documentation describes uploading the HTML and related resources together in a supported archive. The documented maximum upload size is 300 MB; rate and concurrency depend on the license. Recheck those service details when planning a deployment.

wkhtmltopdf: self-managed command line

wkhtmltopdf runs as a process in an environment you manage. The project describes its tools as headless, so a display service is not required. Your team owns installation or building, runtime compatibility, updates, monitoring, capacity, and handling failed conversions. Its manual exposes controls for cookies, headers, print media, local-file access, links, forms, page setup, and JavaScript.

Self-managed execution can suit private or local input workflows, but it also makes deployment and operations part of the conversion feature. Review the LGPLv3 license against how you package and distribute the software.

Inputs, network access, and private pages

  • Public URL: Either tool can be considered for URL input. For PDFCrowd, the URL and its resources must be accessible from the provider’s servers. For wkhtmltopdf, the process must be able to reach the URL from its runtime environment.
  • Private URL or authenticated page: Check the authentication and network path before choosing. wkhtmltopdf documents HTTP cookie and header options. PDFCrowd’s HTTP guide describes API authentication and URL conversion constraints; verify the current supported way to provide access to the target page before relying on it.
  • Generated HTML and local assets: PDFCrowd documents HTML input and packaging local resources with the HTML in an archive. With wkhtmltopdf, verify that the process can resolve every referenced file and that local-file access settings permit the required reads.
  • Large inputs: The PDFCrowd HTTP documentation states a 300 MB maximum upload size. Treat this as a documented service limit to recheck. For either option, reduce unnecessary assets and test the largest expected document.

For sensitive data, inspect data handling, retention, and contractual terms for the hosted service; for self-managed conversion, inspect the process’s filesystem and network permissions. The available documentation here does not support a blanket security ranking.

JavaScript and rendering readiness

Both products document JavaScript-related controls, but those controls do not prove equivalent support for modern sites or equal output fidelity.

  • PDFCrowd: Its overview lists custom JavaScript and readiness controls. Use the documented readiness mechanism that matches the page’s asynchronous behavior, and verify that the expected content is present before conversion.
  • wkhtmltopdf: JavaScript is enabled by default. The manual documents a default delay of 200 milliseconds, a configurable delay, and an option to disable JavaScript. A fixed delay may not match a page whose data arrives later or varies by load.

Test pages with client-side data, delayed images, web fonts, and asynchronous components. Check the resulting PDF itself for missing content; a successful process exit or HTTP response does not establish that the page finished rendering as intended.

Layout and PDF output requirements

Both tools expose common page-layout controls. PDFCrowd documents page size, orientation, margins, page breaks, headers and footers, page numbers, print CSS, and custom CSS. wkhtmltopdf documents page and margin settings, print media, headers and footers, and page numbering.

For either converter, validate output with the real stylesheet and representative content. Check page breaks around long tables and images, font availability, background printing, repeated headers, links, and whether content is clipped at the selected paper size. CSS that looks correct in a browser is not sufficient evidence of matching PDF pagination.

If the output must be watermarked, password protected, PDF/A, or tagged, PDFCrowd lists those capabilities in its product overview. Do not assume wkhtmltopdf has feature parity based on the common settings in its manual; confirm the precise requirement with the relevant documentation and an output inspection.

Runnable starting points

These minimal examples show the integration shapes described in the official docs. Add the page settings and readiness controls required by your document, then check the returned or written PDF. Keep credentials outside source control.

PDFCrowd HTTP request with cURL

curl --user "$PDFCROWD_USERNAME:$PDFCROWD_API_KEY" \
  -F "url=https://example.com/report" \
  -o report.pdf \
  "https://api.pdfcrowd.com/convert/24.04/"

Use the current endpoint and parameter names from the PDFCrowd HTTP guide for your account and conversion type. For HTML text or packaged assets, follow that guide’s corresponding request format.

wkhtmltopdf command line

wkhtmltopdf \
  --page-size A4 \
  --margin-top 15mm \
  --margin-right 12mm \
  --margin-bottom 15mm \
  --margin-left 12mm \
  --javascript-delay 1000 \
  https://example.com/report \
  report.pdf

This example sets a one-second JavaScript delay; choose a value based on the page and verify readiness in the output. The manual documents additional options including cookies, custom headers, print media, and local-file access. Consult the wkhtmltopdf manual for exact syntax and behavior.

How to choose: a validation checklist

  1. Inventory your input. Decide whether conversion starts from a public URL, authenticated URL, generated HTML, or HTML plus local assets.
  2. List required PDF properties. Record paper size, orientation, margins, page breaks, headers and footers, accessibility or archival requirements, and any password or watermark needs.
  3. Test the difficult page. Include JavaScript-driven content, custom fonts, large images, long tables, and the page most likely to expose pagination problems.
  4. Test failure cases. Include a slow resource, unavailable asset, invalid input, and a page that fails to reach its expected ready state. Decide how the application detects, retries, reports, or quarantines each failure.
  5. Compare operational ownership. For PDFCrowd, assess third-party data handling, plan limits, and API dependency. For wkhtmltopdf, include packaging, upgrades, monitoring, runtime capacity, and license review.
  6. Measure your own workload. Compare output correctness, end-to-end latency, concurrency behavior, and total cost at your actual volume. No universal speed or price ranking follows from the documentation summarized here.

Performance, reliability, and cost

There is no supported universal performance ranking in the reviewed sources. Conversion time depends on input size, page behavior, resource loading, configuration, and the chosen deployment. Measure representative documents under expected concurrency rather than extrapolating from a single small page.

With PDFCrowd, rate and concurrency depend on the license; confirm the current plan limits and expected usage. A hosted service shifts renderer operations to the provider but introduces a network request and third-party dependency. With wkhtmltopdf, infrastructure and capacity are yours to size and monitor. Account for upgrades, process failures, and any isolation required for untrusted input.

Calculate total cost with your actual request volume. For the hosted option, check current licensing and service terms. For self-managed use, include engineering time and compute, plus deployment and maintenance. The dossier does not establish a current dollar comparison.

Troubleshooting

Symptom Likely cause What to check
PDFCrowd cannot load a URL or assets are missing The URL or resources are not reachable from the hosted conversion environment, or local assets were not included correctly. Confirm public reachability and resource URLs. For local files, use the documented HTML-plus-assets archive workflow and check its size limit.
PDFCrowd request is rejected or throttled Authentication, request format, account limits, rate, or concurrency. Check Basic-auth credentials, endpoint and parameters against the current HTTP guide; verify the plan’s limits.
wkhtmltopdf output is blank or missing dynamic content The page did not finish rendering before capture, scripts are disabled, or required resources are inaccessible. Check JavaScript settings and delay, network access, and resource paths. Inspect the rendered PDF and adjust readiness based on the page.
Local images or styles are absent in wkhtmltopdf File paths are wrong or local-file access is restricted. Verify paths from the process’s working environment and consult the manual’s local-file access options. Grant only the access the conversion needs.
Content is clipped or pagination is wrong Paper size, margins, print CSS, page-break rules, or renderer behavior differs from expectations. Set page dimensions and margins explicitly; test the actual stylesheet, long content, images, headers, and footers.
Conversion works locally but fails in production Different runtime, fonts, permissions, network access, or process capacity. Match the production environment in validation, capture converter logs and exit status, and verify installed fonts and resource access.

ScreenshotNeo as an alternative to try first

If your goal is to capture a web page as a PDF rather than convert supplied HTML with PDF-specific controls, ScreenshotNeo is a website screenshot API and MCP server. It can return a screenshot or PDF from one GET request. It is not a claim of feature parity with PDFCrowd’s documented PDF/A, tagged-PDF, password, or watermark controls; match the output requirements before choosing.

For a web page screenshot, ScreenshotNeo removes known consent banners, newsletter popups, and chat widgets before capture; only clean shots are billed, while bot checks, blank pages, failed loads, and cache hits cost nothing. Its MCP server lets AI agents use screenshot tools. Plans include 1,000 shots per month free without a card; paid plans start at $5 for 3,000. See the ScreenshotNeo documentation and sign up for 1,000 free screenshots a month, no card required.

Or skip the browser setup

For a webpage capture as a PDF, call the ScreenshotNeo API; see the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free and get 1,000 screenshots a month, no card required.

FAQ

Can wkhtmltopdf run without a graphical desktop?

Yes. The project describes the command-line tools as headless and says they do not require a display service.

Does a documented JavaScript delay guarantee a complete page?

No. A delay is a timing control, not proof that every asynchronous component has finished. Validate the output for the actual page.

Which one should I use?

Use the operating model, input path, PDF requirements, license, and results from representative documents to decide. The documentation does not establish one overall winner.