Html2Pdf.app vs Playwright for Automated Webpage PDF Capture
Compare hosted Html2Pdf.app with Playwright’s browser-based PDF generation, including runnable examples, rendering controls, operations, security, and costs.
Short answer: Choose Html2Pdf.app when you want to send a public URL or HTML to a hosted conversion API and receive PDF bytes. Choose Playwright when you want to run and control the browser yourself, integrate PDF creation into existing browser automation, and manage the runtime and deployment. Neither choice guarantees a particular rendering result: verify representative pages, fonts, dynamic content, and page breaks in the environment you plan to use.
This guide compares the documented integration and controls, gives runnable examples for both options, and covers validation, security, operations, and cost. The official documentation establishes capabilities; it does not provide a matched performance or fidelity benchmark.
1. The practical difference
| Question | Html2Pdf.app | Playwright |
|---|---|---|
| Where does rendering run? | In a hosted service using headless Chromium, according to the vendor documentation. | In a browser launched by your Playwright application or job. |
| What do you send? | An authenticated POST request containing a public URL or raw HTML and conversion options. | Browser automation navigates to a page; then page.pdf() returns a PDF buffer. |
| Who owns browser operations? | The service runs the conversion infrastructure. Your application still owns request handling, credentials, retries, and output storage. | Your team owns browser installation, execution environment, concurrency, resource limits, and operational recovery. |
| When does it fit? | A backend needs a conversion endpoint and the service’s controls and data handling fit its requirements. | The application already uses Playwright, needs browser-level workflow control, or requires its own rendering environment. |
These are different operating models rather than interchangeable performance tiers. The reviewed official sources contain no same-page, same-environment benchmark, so do not infer that either one is faster, more accurate, or cheaper for your workload.
2. Html2Pdf.app: hosted PDF conversion
The documented synchronous flow is a POST to https://api.html2pdf.app/v1/generate with an X-API-Key header and JSON input. The successful response contains PDF bytes. Input can be a public webpage URL or raw HTML in the html field. Keep the API key in backend code or a trusted job; the vendor advises against putting it in browser JavaScript, client-side templates, or public repositories.
Runnable cURL example
curl --fail --silent --show-error \
--request POST 'https://api.html2pdf.app/v1/generate' \
--header 'X-API-Key: YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{"url":"https://example.com","format":"A4","orientation":"portrait","printBackground":true}' \
--output page.pdf
Replace the placeholder key and URL. Treat the response as binary: save it to a file or stream it to your storage layer instead of parsing it as JSON. The vendor documents a URL or raw HTML input; use the input form and option names in its current API documentation for your account.
Runnable Python example
import os
import requests
api_key = os.environ["HTML2PDF_API_KEY"]
endpoint = "https://api.html2pdf.app/v1/generate"
payload = {
"url": "https://example.com",
"format": "A4",
"orientation": "portrait",
"printBackground": True,
}
response = requests.post(
endpoint,
headers={"X-API-Key": api_key, "Content-Type": "application/json"},
json=payload,
timeout=90,
)
response.raise_for_status()
with open("page.pdf", "wb") as pdf:
pdf.write(response.content)
Install the dependency with python -m pip install requests and set HTML2PDF_API_KEY in the process environment. Pick a timeout appropriate to your own pages and job limits; the example is not a service guarantee.
Runnable Node.js example
import { writeFile } from 'node:fs/promises';
const apiKey = process.env.HTML2PDF_API_KEY;
if (!apiKey) throw new Error('Set HTML2PDF_API_KEY');
const response = await fetch('https://api.html2pdf.app/v1/generate', {
method: 'POST',
headers: {
'X-API-Key': apiKey,
'Content-Type': 'application/json',
},
body: JSON.stringify({
url: 'https://example.com',
format: 'A4',
orientation: 'portrait',
printBackground: true,
}),
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) {
throw new Error(`Conversion failed: HTTP ${response.status} ${await response.text()}`);
}
await writeFile('page.pdf', Buffer.from(await response.arrayBuffer()));
Use a Node.js version that supports the built-in fetch and AbortSignal.timeout, or substitute your standard HTTP client and timeout mechanism.
Raw HTML input
For generated documents, the API also documents raw HTML as an input. The HTML may refer to external stylesheets, fonts, or images; those resources must be reachable by the renderer and can affect the output. The vendor’s documentation warns that media selection, available resources and fonts, and JavaScript timing can change the result. Validate the exact document template and resource setup you intend to send. Consult the service documentation for the precise JSON schema and supported option names.
3. Playwright: generate a PDF in your browser workflow
Playwright’s Page.pdf() returns a PDF buffer. Its documented default uses print CSS media. For screen styles, call page.emulateMedia({ media: 'screen' }) before generating the PDF. You control navigation and can wait for application-specific readiness before calling the PDF method.
Runnable JavaScript example
Install Playwright and its Chromium browser using the official setup instructions for your environment. This example targets Node.js with the Playwright package:
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com', {
waitUntil: 'networkidle',
timeout: 60_000,
});
// Optional: use screen CSS instead of the default print CSS.
// await page.emulateMedia({ media: 'screen' });
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '15mm', right: '12mm', bottom: '15mm', left: '12mm' },
});
await writeFile('page.pdf', pdf);
} finally {
await browser.close();
}
networkidle is one possible readiness condition, not a universal signal that a page is complete. Pages with analytics, long polling, or continuously active requests may never become idle. For those pages, wait for a specific content selector or application-ready signal with an explicit timeout, then capture. Conversely, a page that looks loaded can still be missing late fonts, images, or client-rendered content.
PDF options to consider
Playwright’s PDF API documents controls for page format or explicit dimensions, margins, page ranges, print backgrounds, scaling, and whether CSS page size takes precedence. Its API reference is the source of truth for exact option types and behavior. Html2Pdf.app documents format, margins, orientation, media mode, scale, and header/footer templates. Match options by the output requirement rather than assuming similarly named controls behave identically.
- Print or screen styles: Playwright defaults to print media; emulate screen media when that is the intended design. Html2Pdf.app exposes media mode; check the current API schema.
- Paper geometry: Select a named paper format or dimensions and set margins. CSS
@pagerules may also influence layout; validate how your chosen path applies them. - Backgrounds: Enable background printing when colored blocks or background images carry meaning.
- Scale and page ranges: Use scaling carefully because it changes text size and pagination. Request only needed page ranges for large documents where supported.
- Headers and footers: Html2Pdf.app documents header/footer templates. Playwright provides its own PDF options; consult the API reference for supported templates and substitution rules.
4. How to choose
- Map the runtime. If you already operate Playwright workers and need PDF output inside the same browser workflow, Playwright may fit naturally. If you want an HTTP conversion integration and accept a hosted renderer, evaluate Html2Pdf.app.
- List required controls. Write down media mode, paper size, margins, page ranges, scale, backgrounds, headers, and dynamic-content readiness needs. Confirm each exact control in the relevant API.
- Check data handling. Decide whether sending the URL or HTML to a hosted service is allowed. Html2Pdf.app says generated PDFs are temporarily processed rather than permanently stored, while selected request metadata and a source URL may be retained in logs. Review current vendor documentation and your own data rules before sending sensitive content.
- Build a representative corpus. Include long and short pages, web fonts, images, charts, tables, unusual page breaks, authenticated pages if relevant, and dynamic sections. Compare content correctness and pagination in the actual deployment environment.
- Model the whole cost. For the API, check current plan limits and price. For Playwright, account for compute, browser images, storage, queueing, engineering, and on-call ownership. The available sources do not provide a like-for-like total cost calculation.
Html2Pdf.app’s product page displays free and paid plans, but prices and quotas can change; check the current page before purchase. Do not use vendor-displayed volume, speed, or uptime claims as an independent comparison. The reviewed documentation supplies no controlled comparative measurements.
5. Rendering validation, reliability, and performance
Validate the document, not just the HTTP response
- Check that the response is a valid PDF and has nonzero size.
- Inspect page count, first and last pages, and known page-break boundaries.
- Verify selectable text, font appearance, image presence, colors, and headers or footers.
- Test slow resources and JavaScript-rendered sections; record whether the content is present at capture time.
- Repeat captures in the production runtime after changing browser versions, fonts, templates, or CSS.
These checks are a suggested validation procedure, not reported benchmark results. Html2Pdf.app specifically cautions that media mode, fonts and resources, and JavaScript timing affect results. Playwright documents its print-media default and controls, but the reference does not promise that output will match another renderer.
Reliability and performance choices
For a hosted API, your service depends on successful network calls and the provider’s current service behavior. Set bounded timeouts, log request identifiers or safe diagnostic context, and retry only errors that are plausibly transient. Avoid blind retries that can duplicate work or exceed quotas. If the documented callback workflow suits long jobs, it can move waiting out of the request path.
For Playwright, keep browser processes and pages bounded, close them in cleanup paths, and isolate jobs so a failed navigation does not strand resources. Limit concurrency according to available memory and CPU, and measure queue delay and conversion duration with your own representative workload. Neither the dossier nor official API docs provide matched latency, throughput, or resource measurements, so capacity planning requires workload-specific observation.
6. Async generation and callbacks with Html2Pdf.app
The service documents a callback URL option for background generation. The callback includes a base64-encoded document and may echo a caller-provided state; delivery may be retried up to three times. That means the receiver should be prepared for duplicate delivery and should make processing idempotent.
- Create a callback endpoint reachable by the service and associate an opaque job identifier in state.
- Submit the conversion with the callback URL using the current documented request schema.
- Validate the callback according to the service’s current security guidance before trusting its payload.
- Decode the returned base64 document, verify it is a PDF, and store it against the job identifier.
- Make duplicate callback handling safe; acknowledge successful processing so retries do not create duplicate downstream effects.
Consult the current API documentation for the precise callback field names, request shape, authentication or validation mechanism, and response expectations. Do not expose credentials in the callback URL or log sensitive document data.
7. Security and data handling
- Protect credentials: Keep Html2Pdf.app’s private API key in a server-side secret store or trusted worker environment. Do not include it in browser code or source control.
- Assess source access: A hosted renderer must be able to fetch a public URL and its resources. Do not assume it can access a private network or authenticated page unless the vendor documents a supported and secure mechanism.
- Review retention: Html2Pdf.app describes temporary processing of generated PDFs and possible retention of selected request metadata and source URL in logs. Treat URLs as potentially sensitive and review current policy for your use case.
- Control output access: Store generated PDFs with access controls and retention periods appropriate to their contents.
- Isolate Playwright jobs: If rendering untrusted pages, treat navigation and downloaded content as untrusted input. Use your organization’s browser isolation and network egress controls.
8. Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| API request is rejected | Missing or invalid API key, malformed JSON, or a field that does not match the current schema. | Check the X-API-Key header, JSON content type, endpoint, and exact current request fields. Keep the key server-side. |
| Saved output is not a PDF | The request returned an error body but the client saved it as if it were binary output. | Check HTTP status before writing; on failure, inspect the response body securely. Save successful response bytes unchanged. |
| Fonts or images are missing | Resources were unavailable to the renderer, loaded too late, or referenced through inaccessible URLs. | Check resource URLs and network access; allow for font and image loading before capture. Test in the actual rendering environment. |
| Content is missing from a PDF | Client-side rendering had not completed when capture began. | Wait for a content selector or application-ready condition. Avoid relying only on a fixed short delay. |
| Layout differs from the browser screenshot | PDF generation uses print CSS, or page size, margins, scale, and CSS page rules differ. | Check the selected media mode. In Playwright, use screen media explicitly if needed. Set and verify page geometry and print backgrounds. |
| Playwright navigation times out | The site is slow, unavailable, or keeps requests active so the chosen load condition does not finish. | Use an appropriate timeout and wait condition; for active pages, wait for the specific content needed rather than global network idle. |
| Playwright process hangs or exhausts resources | Pages or browser processes are not closed, or concurrency exceeds available resources. | Close pages and browsers in finally cleanup, cap concurrent jobs, and observe memory and CPU under representative load. |
| Callback appears more than once | The documented callback delivery can be retried up to three times. | Make job completion idempotent and safely handle duplicate delivery. |
9. Or skip the browser setup
If your task is to capture a webpage as an image or PDF and you do not need to own browser execution, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API returns PNG, JPEG, WebP, or PDF. For a screenshot capture, one GET request can look like this; see the ScreenshotNeo API documentation for parameters and PDF options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
10. Frequently asked questions
Does Playwright produce PDFs from pages that need JavaScript?
It can generate a PDF from the browser page after your automation has navigated and waited for the content you need. Choose a readiness condition that matches the application; a generic load or network-idle event may not mean all page content is ready.
Can Html2Pdf.app convert raw HTML without a URL?
Its documentation describes both public URL input and raw HTML input. External resources and script timing can affect the final document, so validate the template and resource availability.
Which one should I use for confidential documents?
That depends on your data policy and deployment. Review the hosted service’s current processing and logging terms, or assess the controls and isolation of your own Playwright environment before sending or rendering sensitive material.
Can I conclude one is cheaper or faster from the published documentation?
No. The reviewed sources do not provide a matched benchmark or equivalent total-cost model. Compare current service pricing with your own infrastructure and engineering costs on a representative workload.
Sources
- Html2Pdf.app API documentation for request flow, rendering options, callbacks, and data-handling guidance.
- Html2Pdf.app product page for current plans and commercial details.
- Playwright Page.pdf() API reference for PDF defaults and options.
- Playwright emulateMedia() API reference for selecting screen media.
