Generate PDF Documents from HTML with an API
Generate reliable PDFs from HTML with Gotenberg or Playwright, including JavaScript pages, assets, print layout, accessibility, and failure handling.
Short answer
Use a Chromium-backed PDF endpoint. For local templates, POST multipart index.html and its assets to Gotenberg /forms/chromium/convert/html. For a deployed page, POST its URL to /forms/chromium/convert/url; Chromium executes JavaScript before producing the PDF. If you need application-level control, launch Chromium with Playwright and call page.pdf().
Choose an input route
| Source | Route | Use it when |
|---|---|---|
| Local HTML and assets | Gotenberg /forms/chromium/convert/html |
You own a template, images, fonts, and stylesheets. |
| Deployed page | Gotenberg /forms/chromium/convert/url |
The page is an SPA or loads data in the browser. |
| Custom service | Playwright page.pdf() |
You need authentication, data injection, custom waits, or per-request logic. |
Browser rendering preserves modern CSS and JavaScript. A string-to-PDF library is only suitable when your markup does not depend on browser layout, fonts, or client-side rendering.
1. Convert local HTML with Gotenberg
Run Gotenberg where your application can reach it. Upload a complete index.html; upload images, fonts, and stylesheets as additional multipart files and reference them by their filenames.
<!doctype html>
<html lang='en'>
<head>
<meta charset='utf-8'>
<style>
@page { size: A4; margin: 18mm 16mm 20mm; }
body { font: 11pt/1.45 Arial, sans-serif; color: #222; }
.keep { break-inside: avoid; }
.new-page { break-before: page; }
</style>
</head>
<body>
<h1>Invoice 1042</h1>
<p>Issued 2026-10-01</p>
<section class='keep'><h2>Line items</h2><p>Consulting — 8 hours — $1,200</p></section>
<p>Total: $1,200</p>
</body>
</html>
cURL
curl --request POST http://localhost:3000/forms/chromium/convert/html \
--form files=@/path/to/index.html \
--form files=@/path/to/logo.png \
--output invoice.pdf
Python
import requests
with open('index.html', 'rb') as html:
response = requests.post(
'http://localhost:3000/forms/chromium/convert/html',
files={'files': ('index.html', html, 'text/html')},
timeout=120,
)
response.raise_for_status()
with open('invoice.pdf', 'wb') as output:
output.write(response.content)
Node.js
import { readFile, writeFile } from 'node:fs/promises';
const form = new FormData();
form.append('files', new Blob([await readFile('index.html')], { type: 'text/html' }), 'index.html');
const response = await fetch('http://localhost:3000/forms/chromium/convert/html', { method: 'POST', body: form });
if (!response.ok) throw new Error(`${response.status} ${await response.text()}`);
await writeFile('invoice.pdf', Buffer.from(await response.arrayBuffer()));
2. Convert a deployed URL
The URL route is intended for JavaScript-heavy pages, SPAs, and pages that fetch data after navigation.
curl --request POST http://localhost:3000/forms/chromium/convert/url \
--form url=https://example.com/invoice/1042 \
--output invoice.pdf
import requests
r = requests.post('http://localhost:3000/forms/chromium/convert/url', data={'url': 'https://example.com/invoice/1042'}, timeout=120)
r.raise_for_status()
open('invoice.pdf', 'wb').write(r.content)
Wait for a selector or expression that proves the data is ready. A fixed delay is less reliable because it can be too short on a slow run and wasteful on a fast run.
3. Control paper size and pagination
- CSS page size: Define
@pageand enablepreferCssPageSizewhen the document owns its paper dimensions. Otherwise set paper width and height in the request. - Margins and orientation: Set explicit margins and choose portrait or landscape.
- Backgrounds: Set
printBackground=truefor colored sections, gradients, and background images. - Scale: Change scale only after fixing CSS dimensions; scaling can make text unreadably small.
- Breaks: Use
break-inside: avoid,break-before: page, andbreak-after: pagefor cards, tables, and chapters. - Page ranges: Export selected pages for previews or partial documents when supported.
Use semantic h1 through h6 headings. Avoid fixed-height containers and essential content positioned across page boundaries.
4. Build an endpoint with Playwright
Playwright PDF generation is Chromium-only. The API reference recommends calling page.emulateMedia() before page.pdf() when you need screen media.
import { chromium } from 'playwright';
const browser = await chromium.launch();
try {
const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
await page.goto('https://example.com/invoice/1042', { waitUntil: 'networkidle' });
await page.waitForSelector('[data-ready="true"]');
await page.emulateMedia({ media: 'screen' });
await page.pdf({
path: 'invoice.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '18mm', right: '16mm', bottom: '20mm', left: '16mm' }
});
} finally {
await browser.close();
}
5. Reliability and security
- Set a bounded conversion timeout and terminate workers that exceed it.
- Choose policies for
failOnHttpStatusCodes,failOnResourceHttpStatusCodes, andfailOnResourceLoadingFailed. - Restrict outbound URLs when users submit arbitrary HTML or links to reduce SSRF risk.
- Record request IDs, renderer logs, source versions, and failed resource URLs.
- Limit upload size, page count, and concurrency; image-heavy Chromium jobs use substantial memory.
- Reuse a browser process where appropriate, but create an isolated Playwright context per job and close it in
finally.
6. Accessibility and document governance
Enable generateDocumentOutline for bookmarks. The outline is based on semantic headings and also enables tagged PDF generation. Gotenberg documents PDF/A and PDF/UA post-processing, metadata, encryption, watermarks, stamps, and page ranges. PDF/A and encryption are mutually exclusive, and some post-processing can rasterize table cells, so validate the output required by your archive or accessibility policy.
7. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Blank or incomplete SPA | Capture happened before data rendered | Use the URL route and wait for a selector or expression. |
| Missing images or fonts | Asset was not uploaded or is unreachable | Upload each asset, match filenames, or use reachable absolute URLs. |
| Background colors missing | Print backgrounds disabled | Set printBackground=true and check print CSS. |
| Wrong paper size | CSS page size ignored | Enable preferCssPageSize or set explicit dimensions. |
| Cards or rows split | Uncontrolled page breaks | Use break-inside: avoid, remove rigid heights, and add explicit chapter breaks. |
| Conversion hangs | Polling page or resource never settles | Use a bounded timeout, deterministic readiness condition, tracing, and resource blocking. |
| Unexpected HTTP success | Error status policy is permissive | Configure the three failOn... controls for your policy. |
| Bookmarks absent | No semantic headings or outline option | Use h1–h6 and enable generateDocumentOutline. |
| PDF/A and encryption rejected | Those options conflict | Choose archival conformance or encryption. |
| Playwright PDF unsupported | Non-Chromium browser | Launch Chromium. |
8. Performance and cost
- Browser startup and rendering dominate latency; keep a warm browser process and cap concurrent pages.
- Wait for the actual readiness signal instead of a large blanket delay.
- Block analytics, ads, and unused resource types when they do not affect the document.
- Cache deterministic PDFs by template version, input hash, and asset version.
- Self-hosting means paying for CPU, memory, Chromium updates, queues, and observability. Compare that operational cost with hosted API pricing and your peak concurrency.
Or skip the browser setup
ScreenshotNeo is a hosted website capture API that can return PNG, JPEG, WebP, or PDF. Its documentation lists PDF paper size, margins, landscape mode, page ranges, and other options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account.
FAQ
Can these APIs execute JavaScript?
Yes. Gotenberg’s URL route and Playwright use Chromium, so client-side rendering runs before export.
Should I upload HTML or send a URL?
Upload HTML for private templates and local assets. Send a URL for a deployed page or SPA that the renderer can reach.
How do I keep an invoice section together?
Wrap it in a container with break-inside: avoid and remove fixed heights that force overflow.
What architecture works for production?
Queue jobs, run bounded Chromium workers, store the PDF, expose a status or download endpoint, and retain renderer errors for diagnosis.


