ScreenshotNeo

BlogHTML to image & PDF

Generate PDF Documents from HTML with an API

Generate reliable PDFs from HTML with Gotenberg or Playwright, including JavaScript pages, assets, print layout, accessibility, and failure handling.

By the ScreenshotNeo team1 October 20266 min read

Short answer

Use a Chromium-backed PDF endpoint. For local templates, POST multipart index.html and its assets to Gotenberg /forms/chromium/convert/html. For a deployed page, POST its URL to /forms/chromium/convert/url; Chromium executes JavaScript before producing the PDF. If you need application-level control, launch Chromium with Playwright and call page.pdf().

Choose an input route

Source Route Use it when
Local HTML and assets Gotenberg /forms/chromium/convert/html You own a template, images, fonts, and stylesheets.
Deployed page Gotenberg /forms/chromium/convert/url The page is an SPA or loads data in the browser.
Custom service Playwright page.pdf() You need authentication, data injection, custom waits, or per-request logic.

Browser rendering preserves modern CSS and JavaScript. A string-to-PDF library is only suitable when your markup does not depend on browser layout, fonts, or client-side rendering.

1. Convert local HTML with Gotenberg

Run Gotenberg where your application can reach it. Upload a complete index.html; upload images, fonts, and stylesheets as additional multipart files and reference them by their filenames.

<!doctype html>
<html lang='en'>
<head>
  <meta charset='utf-8'>
  <style>
    @page { size: A4; margin: 18mm 16mm 20mm; }
    body { font: 11pt/1.45 Arial, sans-serif; color: #222; }
    .keep { break-inside: avoid; }
    .new-page { break-before: page; }
  </style>
</head>
<body>
  <h1>Invoice 1042</h1>
  <p>Issued 2026-10-01</p>
  <section class='keep'><h2>Line items</h2><p>Consulting — 8 hours — $1,200</p></section>
  <p>Total: $1,200</p>
</body>
</html>

cURL

curl --request POST http://localhost:3000/forms/chromium/convert/html \
  --form files=@/path/to/index.html \
  --form files=@/path/to/logo.png \
  --output invoice.pdf

Python

import requests

with open('index.html', 'rb') as html:
    response = requests.post(
        'http://localhost:3000/forms/chromium/convert/html',
        files={'files': ('index.html', html, 'text/html')},
        timeout=120,
    )
response.raise_for_status()
with open('invoice.pdf', 'wb') as output:
    output.write(response.content)

Node.js

import { readFile, writeFile } from 'node:fs/promises';

const form = new FormData();
form.append('files', new Blob([await readFile('index.html')], { type: 'text/html' }), 'index.html');
const response = await fetch('http://localhost:3000/forms/chromium/convert/html', { method: 'POST', body: form });
if (!response.ok) throw new Error(`${response.status} ${await response.text()}`);
await writeFile('invoice.pdf', Buffer.from(await response.arrayBuffer()));

2. Convert a deployed URL

The URL route is intended for JavaScript-heavy pages, SPAs, and pages that fetch data after navigation.

curl --request POST http://localhost:3000/forms/chromium/convert/url \
  --form url=https://example.com/invoice/1042 \
  --output invoice.pdf
import requests
r = requests.post('http://localhost:3000/forms/chromium/convert/url', data={'url': 'https://example.com/invoice/1042'}, timeout=120)
r.raise_for_status()
open('invoice.pdf', 'wb').write(r.content)

Wait for a selector or expression that proves the data is ready. A fixed delay is less reliable because it can be too short on a slow run and wasteful on a fast run.

3. Control paper size and pagination

  • CSS page size: Define @page and enable preferCssPageSize when the document owns its paper dimensions. Otherwise set paper width and height in the request.
  • Margins and orientation: Set explicit margins and choose portrait or landscape.
  • Backgrounds: Set printBackground=true for colored sections, gradients, and background images.
  • Scale: Change scale only after fixing CSS dimensions; scaling can make text unreadably small.
  • Breaks: Use break-inside: avoid, break-before: page, and break-after: page for cards, tables, and chapters.
  • Page ranges: Export selected pages for previews or partial documents when supported.

Use semantic h1 through h6 headings. Avoid fixed-height containers and essential content positioned across page boundaries.

4. Build an endpoint with Playwright

Playwright PDF generation is Chromium-only. The API reference recommends calling page.emulateMedia() before page.pdf() when you need screen media.

import { chromium } from 'playwright';

const browser = await chromium.launch();
try {
  const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
  await page.goto('https://example.com/invoice/1042', { waitUntil: 'networkidle' });
  await page.waitForSelector('[data-ready="true"]');
  await page.emulateMedia({ media: 'screen' });
  await page.pdf({
    path: 'invoice.pdf',
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true,
    margin: { top: '18mm', right: '16mm', bottom: '20mm', left: '16mm' }
  });
} finally {
  await browser.close();
}

5. Reliability and security

  1. Set a bounded conversion timeout and terminate workers that exceed it.
  2. Choose policies for failOnHttpStatusCodes, failOnResourceHttpStatusCodes, and failOnResourceLoadingFailed.
  3. Restrict outbound URLs when users submit arbitrary HTML or links to reduce SSRF risk.
  4. Record request IDs, renderer logs, source versions, and failed resource URLs.
  5. Limit upload size, page count, and concurrency; image-heavy Chromium jobs use substantial memory.
  6. Reuse a browser process where appropriate, but create an isolated Playwright context per job and close it in finally.

6. Accessibility and document governance

Enable generateDocumentOutline for bookmarks. The outline is based on semantic headings and also enables tagged PDF generation. Gotenberg documents PDF/A and PDF/UA post-processing, metadata, encryption, watermarks, stamps, and page ranges. PDF/A and encryption are mutually exclusive, and some post-processing can rasterize table cells, so validate the output required by your archive or accessibility policy.

7. Troubleshooting

Symptom Cause Fix
Blank or incomplete SPA Capture happened before data rendered Use the URL route and wait for a selector or expression.
Missing images or fonts Asset was not uploaded or is unreachable Upload each asset, match filenames, or use reachable absolute URLs.
Background colors missing Print backgrounds disabled Set printBackground=true and check print CSS.
Wrong paper size CSS page size ignored Enable preferCssPageSize or set explicit dimensions.
Cards or rows split Uncontrolled page breaks Use break-inside: avoid, remove rigid heights, and add explicit chapter breaks.
Conversion hangs Polling page or resource never settles Use a bounded timeout, deterministic readiness condition, tracing, and resource blocking.
Unexpected HTTP success Error status policy is permissive Configure the three failOn... controls for your policy.
Bookmarks absent No semantic headings or outline option Use h1–h6 and enable generateDocumentOutline.
PDF/A and encryption rejected Those options conflict Choose archival conformance or encryption.
Playwright PDF unsupported Non-Chromium browser Launch Chromium.

8. Performance and cost

  • Browser startup and rendering dominate latency; keep a warm browser process and cap concurrent pages.
  • Wait for the actual readiness signal instead of a large blanket delay.
  • Block analytics, ads, and unused resource types when they do not affect the document.
  • Cache deterministic PDFs by template version, input hash, and asset version.
  • Self-hosting means paying for CPU, memory, Chromium updates, queues, and observability. Compare that operational cost with hosted API pricing and your peak concurrency.

Or skip the browser setup

ScreenshotNeo is a hosted website capture API that can return PNG, JPEG, WebP, or PDF. Its documentation lists PDF paper size, margins, landscape mode, page ranges, and other options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account.

FAQ

Can these APIs execute JavaScript?

Yes. Gotenberg’s URL route and Playwright use Chromium, so client-side rendering runs before export.

Should I upload HTML or send a URL?

Upload HTML for private templates and local assets. Send a URL for a deployed page or SPA that the renderer can reach.

How do I keep an invoice section together?

Wrap it in a container with break-inside: avoid and remove fixed heights that force overflow.

What architecture works for production?

Queue jobs, run bounded Chromium workers, store the PDF, expose a status or download endpoint, and retain renderer errors for diagnosis.