ScreenshotNeo

BlogHTML to image & PDF

How to Generate Dynamic PDFs with an API

Build a production-ready PDF API with HTML templates, Puppeteer, PDFKit, ReportLab, or a hosted converter.

By the ScreenshotNeo team1 October 202610 min read

Short answer: accept and validate JSON, map it to a versioned template, render the document with an appropriate engine, and return the bytes with Content-Type: application/pdf. Use a browser renderer such as Puppeteer when you already own HTML/CSS templates, a direct library such as PDFKit or ReportLab when you need programmatic layout and streaming, and a hosted conversion API when you want the rendering infrastructure managed for you.

A reliable PDF endpoint also controls pagination, fonts, external assets, timeouts, memory, security, and template versions. This guide shows complete implementations and the production decisions behind them.

1. The data-to-PDF pipeline

Keep the responsibilities separate:

  1. Validate and authorize input. Check the request schema and confirm that the caller may access the referenced records.
  2. Build a template context. Load data by ID on the server rather than trusting arbitrary client-supplied totals or permissions.
  3. Render. Pass the context to a versioned HTML/CSS template or a programmatic document builder.
  4. Finalize. Wait for fonts, images, and other required resources, then generate the PDF.
  5. Return or store. Send the bytes for small documents, or put large files in object storage and return a short-lived URL.

Expose a stable contract. A synchronous endpoint can return the document directly:

POST /v1/invoices/{id}/pdf
Accept: application/json

HTTP/1.1 200 OK
Content-Type: application/pdf
Content-Disposition: inline; filename="invoice-123.pdf"

For slow reports, return 202 Accepted with a job ID and provide a status endpoint or signed completion webhook.

2. HTML and CSS with Puppeteer (Node.js)

Puppeteer launches Chromium, loads your HTML, and calls page.pdf(). Its PDF method uses print CSS by default; call page.emulateMediaType('screen') when your screen styles are the intended output. Set paper size, margins, backgrounds, and header/footer behavior explicitly. See the Puppeteer PDF API.

Install

npm install express puppeteer zod

Complete endpoint

import express from 'express';
import puppeteer from 'puppeteer';
import { z } from 'zod';

const app = express();
app.use(express.json({ limit: '64kb' }));

const InvoiceInput = z.object({
  customerName: z.string().min(1).max(200),
  invoiceNumber: z.string().regex(/^[A-Z0-9-]+$/),
  issuedAt: z.string().datetime(),
  dueAt: z.string().datetime(),
  currency: z.string().length(3),
  items: z.array(z.object({
    description: z.string().min(1).max(500),
    quantity: z.number().positive().finite(),
    unitPrice: z.number().nonnegative().finite()
  })).min(1).max(500),
  taxRate: z.number().min(0).max(1)
});

const escapeHtml = (value) => String(value)
  .replaceAll('&', '&')
  .replaceAll('<', '<')
  .replaceAll('>', '>')
  .replaceAll('"', '"')
  .replaceAll("'", ''');

function renderInvoice(data) {
  const subtotal = data.items.reduce((sum, item) =>
    sum + item.quantity * item.unitPrice, 0);
  const tax = subtotal * data.taxRate;
  const total = subtotal + tax;
  const money = new Intl.NumberFormat('en-US', {
    style: 'currency', currency: data.currency
  });
  const rows = data.items.map(item => `
    
      ${escapeHtml(item.description)}
      ${item.quantity}
      ${money.format(item.unitPrice)}
      ${money.format(item.quantity * item.unitPrice)}
    `).join('');

  return `

  

Invoice ${escapeHtml(data.invoiceNumber)}

Bill to
${escapeHtml(data.customerName)}
Issued
${escapeHtml(data.issuedAt)}
Due
${escapeHtml(data.dueAt)}
${rows}
DescriptionQtyUnit priceAmount
Subtotal${money.format(subtotal)}
Tax${money.format(tax)}
Total${money.format(total)}
`; } let browserPromise; function getBrowser() { if (!browserPromise) browserPromise = puppeteer.launch({ args: ['--no-sandbox', '--disable-setuid-sandbox'] }); return browserPromise; } app.post('/v1/invoices/:id.pdf', async (req, res) => { const parsed = InvoiceInput.safeParse(req.body); if (!parsed.success) return res.status(400).json({ error: 'invalid_input' }); const html = renderInvoice(parsed.data); const browser = await getBrowser(); const page = await browser.newPage(); try { await page.setContent(html, { waitUntil: 'networkidle0', timeout: 30000 }); await page.emulateMediaType('print'); await page.evaluate(() => document.fonts.ready); const pdf = await page.pdf({ format: 'A4', printBackground: true, preferCSSPageSize: true, displayHeaderFooter: false, timeout: 30000 }); res.set({ 'Content-Type': 'application/pdf', 'Content-Disposition': `inline; filename="invoice-${req.params.id}.pdf"`, 'Content-Length': pdf.length }).send(pdf); } catch (error) { console.error('pdf_render_failed', { id: req.params.id, error }); if (!res.headersSent) res.status(504).json({ error: 'pdf_render_failed' }); } finally { await page.close(); } }); app.listen(3000);

Reuse a browser process, but create a fresh page per request. Close every page in finally, cap concurrent renders, and close the browser during graceful shutdown. Do not accept arbitrary navigation URLs from users without network isolation and an allow-list.

Headers and footers

const pdf = await page.pdf({
  format: 'A4',
  printBackground: true,
  displayHeaderFooter: true,
  headerTemplate: '<div></div>',
  footerTemplate: '<div style="font-size:8px;width:100%;text-align:center">Page <span class="pageNumber"></span> of <span class="totalPages"></span></div>',
  margin: { top: '24mm', right: '16mm', bottom: '20mm', left: '16mm' }
});

3. Direct generation with PDFKit

PDFKit avoids a browser runtime. The trade-off is that your code owns line wrapping, pagination, font registration, tables, and image placement. Its PDFDocument is a readable stream, so it can write directly to an HTTP response.

import express from 'express';
import PDFDocument from 'pdfkit';

const app = express();
app.get('/report.pdf', async (req, res, next) => {
  try {
    const summary = await buildSummary();
    res.setHeader('Content-Type', 'application/pdf');
    res.setHeader('Content-Disposition', 'inline; filename="report.pdf"');
    const doc = new PDFDocument({ size: 'A4', margin: 50 });
    doc.pipe(res);
    doc.fontSize(20).text('Quarterly report');
    doc.moveDown().fontSize(11).text(summary, { width: 495 });
    doc.addPage().fontSize(16).text('Details');
    doc.end();
  } catch (error) { next(error); }
});
app.listen(3000);

4. Python with ReportLab

ReportLab supports direct drawing, Platypus flowables, and RML templates. Its json2pdf pattern keeps data extraction separate from a versioned template, which makes fixture-based regression testing practical. See the ReportLab documentation.

pip install fastapi uvicorn reportlab pydantic
from io import BytesIO
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from pydantic import BaseModel, Field
from reportlab.lib.pagesizes import A4
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, Table, TableStyle
from reportlab.lib import colors

app = FastAPI()

class Item(BaseModel):
    description: str = Field(min_length=1, max_length=500)
    quantity: float = Field(gt=0)
    unit_price: float = Field(ge=0)

class Invoice(BaseModel):
    invoice_number: str
    customer_name: str
    items: list[Item] = Field(min_length=1, max_length=500)

@app.post('/v1/invoices/{invoice_id}.pdf')
def invoice_pdf(invoice_id: str, invoice: Invoice):
    output = BytesIO()
    doc = SimpleDocTemplate(output, pagesize=A4, rightMargin=40,
                            leftMargin=40, topMargin=40, bottomMargin=40)
    styles = getSampleStyleSheet()
    story = [Paragraph(f'Invoice {invoice.invoice_number}', styles['Title']),
             Paragraph(invoice.customer_name, styles['Normal']), Spacer(1, 18)]
    rows = [['Description', 'Quantity', 'Unit price', 'Amount']]
    for item in invoice.items:
        rows.append([item.description, str(item.quantity),
                     f'{item.unit_price:.2f}',
                     f'{item.quantity * item.unit_price:.2f}'])
    table = Table(rows, repeatRows=1)
    table.setStyle(TableStyle([
        ('BACKGROUND', (0, 0), (-1, 0), colors.lightgrey),
        ('GRID', (0, 0), (-1, -1), 0.25, colors.grey),
        ('ALIGN', (1, 1), (-1, -1), 'RIGHT'),
        ('VALIGN', (0, 0), (-1, -1), 'TOP')
    ]))
    story.append(table)
    doc.build(story)
    output.seek(0)
    return StreamingResponse(output, media_type='application/pdf',
        headers={'Content-Disposition': f'inline; filename="invoice-{invoice_id}.pdf"'})

5. cURL, Python, and Node.js clients

Any HTTP client can consume the endpoint as binary data. Never decode the response as JSON.

curl -X POST http://localhost:3000/v1/invoices/123.pdf \
  -H 'Content-Type: application/json' \
  -d '{"invoice_number":"INV-123","customer_name":"Acme","items":[{"description":"Consulting","quantity":2,"unit_price":125}]}' \
  -o invoice.pdf
import requests
payload = {
    'invoice_number': 'INV-123', 'customer_name': 'Acme',
    'items': [{'description': 'Consulting', 'quantity': 2, 'unit_price': 125}]
}
r = requests.post('http://localhost:3000/v1/invoices/123.pdf', json=payload, timeout=60)
r.raise_for_status()
open('invoice.pdf', 'wb').write(r.content)
const payload = {
  invoice_number: 'INV-123', customer_name: 'Acme',
  items: [{ description: 'Consulting', quantity: 2, unit_price: 125 }]
};
const res = await fetch('http://localhost:3000/v1/invoices/123.pdf', {
  method: 'POST', headers: { 'content-type': 'application/json' },
  body: JSON.stringify(payload)
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('node:fs').writeFileSync('invoice.pdf', Buffer.from(await res.arrayBuffer()));

6. Hosted conversion APIs

Managed services accept HTML, URLs, or office documents and run the conversion infrastructure for you. Adobe PDF Services documents dynamic HTML, ZIP, URL, and other inputs. HTMLPDF.dev is an example of an API contract accepting either url or raw html plus paper size, orientation, margins, timeout, and output format. PDF Generator API provides versioned templates with text, tables, barcodes, and an expression language.

Before choosing one, review authentication, quotas, retention, data residency, regional processing, retries, maximum document size, font support, webhook signing, and contractual availability. A hosted API reduces browser operations but adds network latency, vendor cost, and dependency risk.

7. Pagination, fonts, and assets

  • Declare page size and margins with @page or engine options.
  • Use thead { display: table-header-group; } so long tables repeat their header.
  • Use break-inside: avoid for invoice rows, cards, and signature blocks, while accepting that very tall elements must split.
  • Wait for document.fonts.ready and for important images to load before rendering.
  • Bundle or pin fonts and assets. Remote resources can disappear, change, or delay a job.
  • Test empty fields, long names, Unicode, right-to-left text, huge tables, images, and page breaks.

8. Security controls

  • Escape every value inserted into HTML. Prefer text nodes and templating auto-escaping.
  • Do not allow arbitrary user URLs in a browser renderer. Use an allow-list, network egress restrictions, and request-level timeouts.
  • Protect internal metadata endpoints and cloud instance credentials from SSRF.
  • Limit JSON size, item count, image dimensions, render duration, and concurrent browser pages.
  • Keep API keys and signing secrets out of templates and logs.
  • Sanitize filenames and set Content-Disposition deliberately.

9. Reliability and performance

Launching Chromium for every request increases latency and memory use; a warm browser with isolated pages is usually more efficient. Bound concurrency with a queue so a burst cannot exhaust memory. Reuse immutable templates and cache stable assets. For large PDFs, stream directly or store them in object storage instead of buffering multiple copies.

Record template and engine versions, render duration, output size, timeout count, and failure reason. Measure these values on your own representative documents; no cross-vendor benchmark establishes a universal winner. Retry only failures that are safe to repeat, and use idempotency keys when a job can create a billable or externally visible artifact.

10. Cost and operating choices

Approach You operate Best fit Main trade-off
Puppeteer Chromium, fonts, sandbox, queue High HTML/CSS fidelity Memory, cold starts, browser security
PDFKit Layout and pagination code Programmatic documents and streaming More manual layout work
ReportLab Python templates and fonts Reports and data-driven layouts Less browser CSS reuse
Hosted API Credentials, contract, integration Managed conversion infrastructure Vendor cost, latency, privacy review

11. Troubleshooting

Blank or partially rendered pages

Cause: the page was captured before fonts, images, or client-side data finished loading. Fix: wait for network idle plus explicit readiness signals, then await document.fonts.ready and critical image completion.

Unexpected page breaks

Cause: implicit paper size, margins, or break rules. Fix: set @page, engine margins, print backgrounds, table header repetition, and break-inside rules explicitly.

Fonts look different

Cause: the runtime lacks the font or loads a fallback. Fix: package a licensed font, register it, wait for the font promise, and test Unicode and right-to-left samples.

Images are missing

Cause: blocked network access, expired signed URLs, CORS assumptions, or a timeout. Fix: use reachable assets, allow only required hosts, increase the bounded timeout, and verify responses before rendering.

Requests time out or the process runs out of memory

Cause: unbounded pages, huge images, too much concurrency, or a browser per request. Fix: reuse the browser, queue jobs, cap input dimensions and duration, close pages in finally, and restart unhealthy workers.

PDF downloads as corrupted data

Cause: treating binary bytes as text or writing an error body into the file. Fix: use binary response handling, check the HTTP status, and set Content-Type: application/pdf.

HTML injection or server-side request forgery

Cause: unescaped values or arbitrary navigation. Fix: escape output, validate schemas, allow-list destinations, block private network ranges, and isolate the renderer.

12. Production checklist

  • Validate and authorize every request.
  • Version templates, fixtures, fonts, and renderer dependencies.
  • Set page format, margins, print backgrounds, headers, and footers explicitly.
  • Wait for fonts, images, and application readiness.
  • Bound time, memory, input size, and concurrency.
  • Escape HTML and restrict remote navigation.
  • Test long tables, page breaks, Unicode, images, empty values, and failures.
  • Stream large outputs or use short-lived object-storage URLs.
  • Log latency, output size, engine version, template version, and failure reason.
  • Measure your real workload before selecting a renderer or hosted provider.

13. Or skip the browser setup

ScreenshotNeo provides a website screenshot API that can also return PDFs. Its GET endpoint handles the capture infrastructure and supports PDF paper size, margins, landscape mode, and page ranges. Read the ScreenshotNeo API docs for the complete option list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

14. FAQ

Should I render invoices in a browser?

Choose a browser when the source design is already HTML/CSS or needs web-font and print-CSS fidelity. Choose PDFKit or ReportLab when deterministic programmatic layout and a smaller runtime matter more.

Can an API return a PDF directly?

Yes. Set Content-Type: application/pdf and send the binary bytes, or return a job resource for long-running generation.

How do I support large reports?

Use bounded asynchronous jobs, stream or store the result, and provide a short-lived download URL. Avoid holding multiple full copies in memory.

How do I make output reproducible?

Pin the engine and fonts, version templates, freeze locale and timezone, use fixed assets, and retain representative regression fixtures.

What is the most common production failure?

Rendering before fonts, images, or client-side data are ready. Add explicit readiness checks and bounded timeouts instead of relying on a single delay.