Formatting and Transforming Data for PDF Generation
Shape, validate, and render structured data into reliable PDFs with ReportLab, WeasyPrint, and ScreenshotNeo.

Direct answer: normalize and validate your data first, then choose a PDF representation that matches the document. Use ReportLab when a Python-native drawing and layout model fits your report; use WeasyPrint when your report is naturally HTML and CSS. Render representative short and long datasets, inspect pagination and formatting, and keep source data separate from generated files.
1. Start with a data-to-document pipeline
A dependable PDF workflow has four boundaries:
- Input: load JSON, database rows, CSV records, or API responses.
- Transformation: normalize types, fill or flag missing values, format dates and numbers, and calculate derived fields.
- Presentation: apply page size, typography, spacing, colors, tables, charts, and page-break rules.
- Verification: inspect the resulting PDF for clipping, wrapping, links, repeated headers, and required PDF features.
Keeping transformation separate from presentation lets you change a template without changing business data. It also makes validation and regression checks easier.
Define output assumptions before coding
- Audience: internal analyst, customer, regulator, or print reader.
- Page size and orientation: choose explicitly. ReportLab measures page sizes in points; do not rely on an implicit default. See the ReportLab user guide.
- Destination: download, email attachment, archive, browser preview, or print.
- Required features: links, bookmarks, forms, attachments, page numbers, charts, or signatures.
- Locale rules: currency, decimal separator, time zone, calendar, and translated labels.
2. Normalize and validate structured data
Do not let a renderer decide how malformed values should appear. Convert input into a predictable internal shape first.

from dataclasses import dataclass
from datetime import date
from decimal import Decimal
from typing import Iterable
@dataclass
class InvoiceLine:
description: str
quantity: Decimal
unit_price: Decimal
@property
def total(self) -> Decimal:
return self.quantity * self.unit_price
@dataclass
class Invoice:
number: str
issued_on: date
customer: str
lines: list[InvoiceLine]
@property
def subtotal(self) -> Decimal:
return sum((line.total for line in self.lines), Decimal('0'))
def parse_invoice(raw: dict) -> Invoice:
if not raw.get('number'):
raise ValueError('number is required')
if not raw.get('customer'):
raise ValueError('customer is required')
lines = []
for index, item in enumerate(raw.get('lines', []), start=1):
try:
quantity = Decimal(str(item['quantity']))
unit_price = Decimal(str(item['unit_price']))
except (KeyError, ValueError, TypeError) as exc:
raise ValueError(f'line {index} has invalid numeric data') from exc
if quantity < 0 or unit_price < 0:
raise ValueError(f'line {index} cannot be negative')
lines.append(InvoiceLine(str(item['description']), quantity, unit_price))
if not lines:
raise ValueError('at least one line is required')
return Invoice(
number=str(raw['number']),
issued_on=date.fromisoformat(raw['issued_on']),
customer=str(raw['customer']),
lines=lines,
)
Use decimal arithmetic for money, preserve the original source record for auditing, and decide explicitly whether missing values become an empty cell, a visible placeholder, or a validation error.
Format values at the presentation boundary
from decimal import Decimal
def money(value: Decimal, currency='$') -> str:
return f'{currency}{value:,.2f}'
def quantity(value: Decimal) -> str:
return f'{value.normalize():f}'
def display_date(value) -> str:
return value.strftime('%Y-%m-%d')
# Keep raw values in the model; format only when building a paragraph or table cell.
Long labels, nulls, very large numbers, negative values, and non-ASCII names should all appear in representative fixtures before production use.
3. Generate PDFs with ReportLab
ReportLab provides Python interfaces for drawing directly onto PDF pages and higher-level layout constructs for reports with text, tables, and charts. Its pdfgen canvas is the lower-level page-painting interface. The table guide documents row-height calculation, page splitting, and repeated rows at page breaks.
Complete ReportLab example
from io import BytesIO
from reportlab.lib import colors
from reportlab.lib.enums import TA_RIGHT
from reportlab.lib.pagesizes import A4
from reportlab.lib.styles import getSampleStyleSheet, ParagraphStyle
from reportlab.lib.units import mm
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, Table, TableStyle, PageBreak
def build_invoice_pdf(invoice):
output = BytesIO()
doc = SimpleDocTemplate(
output,
pagesize=A4,
rightMargin=18 * mm,
leftMargin=18 * mm,
topMargin=18 * mm,
bottomMargin=18 * mm,
title=f'Invoice {invoice.number}',
author='PDF generator',
)
styles = getSampleStyleSheet()
styles.add(ParagraphStyle(name='Right', parent=styles['Normal'], alignment=TA_RIGHT))
story = [
Paragraph(f'Invoice {invoice.number}', styles['Title']),
Paragraph(f'Customer: {invoice.customer}', styles['Normal']),
Paragraph(f'Issued: {invoice.issued_on:%Y-%m-%d}', styles['Normal']),
Spacer(1, 8 * mm),
]
rows = [[Paragraph('Description', styles['Normal']),
Paragraph('Qty', styles['Right']),
Paragraph('Unit price', styles['Right']),
Paragraph('Total', styles['Right'])]]
for line in invoice.lines:
rows.append([
Paragraph(line.description, styles['Normal']),
Paragraph(quantity(line.quantity), styles['Right']),
Paragraph(money(line.unit_price), styles['Right']),
Paragraph(money(line.total), styles['Right']),
])
rows.append(['', '', Paragraph('Subtotal', styles['Right']),
Paragraph(f'{money(invoice.subtotal)}', styles['Right'])])
table = Table(rows, colWidths=[88 * mm, 18 * mm, 34 * mm, 34 * mm], repeatRows=1)
table.setStyle(TableStyle([
('BACKGROUND', (0, 0), (-1, 0), colors.HexColor('#e8eef5')),
('GRID', (0, 0), (-1, -1), 0.35, colors.HexColor('#9aa7b5')),
('VALIGN', (0, 0), (-1, -1), 'TOP'),
('ALIGN', (1, 1), (-1, -1), 'RIGHT'),
('LEFTPADDING', (0, 0), (-1, -1), 6),
('RIGHTPADDING', (0, 0), (-1, -1), 6),
('TOPPADDING', (0, 0), (-1, -1), 5),
('BOTTOMPADDING', (0, 0), (-1, -1), 5),
]))
story.append(table)
doc.build(story)
return output.getvalue()
if __name__ == '__main__':
from datetime import date
from decimal import Decimal
invoice = Invoice('INV-1001', date.today(), 'Example customer', [
InvoiceLine('Implementation and data preparation', Decimal('2'), Decimal('125.00')),
InvoiceLine('PDF template work with a deliberately long description that must wrap', Decimal('1'), Decimal('240.00')),
])
with open('invoice.pdf', 'wb') as file:
file.write(build_invoice_pdf(invoice))
For long tables, set explicit column widths, use repeatRows=1, and test rows with long wrapped content. The documented splitting behavior helps, but the actual template still needs inspection.
When to use the canvas
Use canvas.Canvas when you need exact coordinates, custom drawings, labels, or a fixed-page form. Track the current y-coordinate yourself, start a new page when it reaches the bottom margin, and choose the page size explicitly. Flowables and tables are usually easier to maintain for variable-length reports.
4. Generate PDFs from HTML and CSS with WeasyPrint
WeasyPrint accepts HTML and CSS and can write a PDF to a path or return PDF bytes. Its documentation also warns that unsupported CSS properties may produce warnings, so verify the CSS features your template depends on. Start with semantic markup and print-specific CSS.
Complete WeasyPrint example
from pathlib import Path
from weasyprint import HTML
html = '''
Invoice INV-1001
Description Qty Unit price Total
Implementation and data preparation 2 $125.00 $250.00
Long descriptions wrap inside the cell instead of clipping 1 $240.00 $240.00
'''
HTML(string=html, base_url=Path('.').resolve().as_uri()).write_pdf('invoice-weasyprint.pdf')
Use a base_url when templates reference local stylesheets, fonts, or images. Render with real data and review warnings; a successful process exit does not prove that every CSS rule produced the intended result.
5. Choose the representation that fits
| Question | ReportLab | WeasyPrint |
|---|---|---|
| Authoring model | Python drawing and document-layout objects | HTML structure and CSS presentation |
| Best fit | Programmatic layouts, native tables, charts, and precise drawing | Templates shared with web content or teams comfortable with HTML/CSS |
| Pagination | Flowables, table splitting, and repeated rows | Print CSS, page rules, and supported HTML/CSS features |
| Main risk | Manual coordinate logic can become brittle | Unsupported CSS or unexpected print layout behavior |
Neither source provides a controlled speed, fidelity, or cost benchmark against the other. Measure your own representative documents if those factors determine the choice.
6. Pagination and formatting checklist
- Set page size, margins, and orientation explicitly.
- Give tables intentional column widths and allow long text to wrap.
- Repeat table headers on every page of a long table.
- Keep rows together where splitting would make the meaning unclear.
- Reserve space for headers, footers, page numbers, and totals.
- Test empty, one-row, many-row, and very-long-label datasets.
- Check fonts and glyph coverage for every supported language.
- Open the generated PDF in more than one viewer when links or print fidelity matter.

7. Or skip the browser setup
If your source is already a web page, ScreenshotNeo can capture a PDF with one HTTP request instead of maintaining a browser-rendering stack. Read the ScreenshotNeo API documentation for all options.
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o report.pdf
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com', 'format': 'pdf'},
timeout=90,
)
r.raise_for_status()
open('report.pdf', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com', format: 'pdf' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('report.pdf', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. It also provides custom CSS and JavaScript, wait conditions, request blocking, headers, cookies, user-agent, authorization, timezone, geolocation, PDF paper size, margins, landscape mode, page ranges, caching, async jobs, bulk capture, signed links, and an MCP server with take_screenshot, get_page_info, and capture_pdf for AI agents. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
8. Troubleshooting
Text is clipped or overlaps
Cause: fixed coordinates, narrow columns, or a font with different metrics. Fix: use flowable layout or wrapping cells, widen the column, reduce font size deliberately, and test the longest real value.
Table headers disappear after a page break
Cause: the header was not marked as repeating. Fix: use ReportLab’s repeatRows=1 or print CSS with a table header group in WeasyPrint.
CSS appears to be ignored
Cause: the property is unsupported or the stylesheet cannot be resolved. Fix: read WeasyPrint warnings, provide a correct base_url, and replace unsupported rules with documented print CSS.
Images or fonts are missing
Cause: relative URLs have no base, files are inaccessible, or the runtime lacks the font. Fix: use absolute or resolvable paths, set base_url, package assets with the job, and verify font licensing and glyph coverage.
Dates or money are wrong
Cause: formatting happened implicitly in the template or values were converted through binary floating point. Fix: normalize types first, use decimal arithmetic for money, and format with an explicit locale policy.
PDF generation is slow or memory-heavy
Cause: huge tables, oversized images, repeated full-document renders, or an unbounded queue. Fix: paginate input, resize images, stream or batch jobs where supported, cache identical outputs, and measure with representative documents. The research sources do not establish a universal performance number.
ScreenshotNeo returns a non-success response
Check the HTTP status and response headers, confirm the URL is reachable, and inspect the page verdict. Add a selector wait, delay, network-idle wait, custom headers, cookies, or authorization when the target requires them. Failed loads and bot checks are not billed.
9. Reliability, cost, and operational design
- Store the input data and generated PDF separately so a template can be rerun.
- Record template version, renderer version, locale, time zone, and input identifier with each artifact.
- Use deterministic fixtures for regression checks after library or template upgrades.
- Validate output existence and basic page count before publishing or emailing.
- Retry transient network or asset failures with a limit; do not silently replace missing data.
- For ScreenshotNeo, use caching with a chosen TTL when repeated captures are acceptable, and use async jobs with signed webhooks for long or bulk workflows.
- Estimate cost from actual billed captures. ScreenshotNeo bills only clean shots; cache hits and failed categories listed in its response are not billed.
10. FAQ
Should I generate a PDF directly from JSON?
Usually transform JSON into a validated document model first. That keeps validation and presentation rules explicit.
Is HTML/CSS always easier than drawing?
No. HTML/CSS is a natural fit for markup-driven reports; ReportLab can be clearer for Python-native layouts, tables, charts, or exact drawing.
How do I handle a table with thousands of rows?
Paginate deliberately, repeat headers, test memory use, and consider separate files or an export format better suited to raw data when a reader does not need every row in one PDF.
Can a PDF template be reused for changing data?
Yes. Keep the template and transformation code separate, validate each input, and rerun the same representative fixtures whenever either changes.
Where can I find ScreenshotNeo options?
See the ScreenshotNeo documentation for capture, PDF, waiting, blocking, authentication, caching, and async parameters.


