ScreenshotNeo

BlogHTML to image & PDF

How to Convert HTML to PDF in Python with WeasyPrint

Convert HTML to PDF in Python with WeasyPrint, including assets, print CSS, fonts, security, troubleshooting, and production patterns.

By the ScreenshotNeo team1 October 20268 min read

WeasyPrint converts HTML and CSS into PDF without launching a browser. The shortest working example is:

from weasyprint import HTML

HTML(string='<h1>Hello, PDF</h1>').write_pdf('output.pdf')

Use the explicit string=, filename=, or url= argument that matches your input. Pass a base_url when inline HTML refers to relative images, stylesheets, or fonts. The examples and behavior below follow the official WeasyPrint first-steps guide and API reference.

1. Install WeasyPrint

Install it inside the Python environment that will render your documents:

python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install weasyprint

WeasyPrint 70.0 documentation lists Python 3.10 or newer and Pango 1.44 or newer. Operating-system packages differ, so check the installation instructions for your target system and verify native dependencies in the same environment used by production.

Pin the version in deployed applications and record the Python, Pango, font, and operating-system versions. Rendering can change when any of these change.

2. Convert a string, file, or URL

HTML held in memory

from weasyprint import HTML

markup = '''
<!doctype html>
<html>
  <head><meta charset='utf-8'></head>
  <body><h1>Invoice</h1><p>Thank you.</p></body>
</html>
'''

HTML(string=markup).write_pdf('invoice.pdf')

Use the named string= argument. An ambiguous positional string can be interpreted as a filename or URL instead of markup.

Local HTML file

from weasyprint import HTML

HTML(filename='templates/report.html').write_pdf('report.pdf')

Relative resources are resolved from the HTML file location.

Remote URL

from weasyprint import HTML

HTML(url='https://example.com/article').write_pdf('article.pdf')

The default fetcher handles file and HTTP URLs, but it does not provide advanced cookie or authentication handling. Use a custom URL fetcher when the source requires controlled headers or credentials.

Return PDF bytes

from weasyprint import HTML

pdf_bytes = HTML(string='<h1>Report</h1>').write_pdf()
with open('report.pdf', 'wb') as output:
    output.write(pdf_bytes)

Omit the target to receive bytes. You can also pass a path or file object to write_pdf().

3. Resolve images, CSS, and fonts

Inline markup has no useful filesystem location. Set base_url so relative references can be fetched:

from pathlib import Path
from weasyprint import HTML

root = Path(__file__).parent
markup = '''
<html>
  <head><link rel='stylesheet' href='css/print.css'></head>
  <body><img src='images/logo.png' alt='Logo'></body>
</html>
'''

HTML(string=markup, base_url=str(root)).write_pdf('branded.pdf')

You can also include a <base href='...'> element. Check every image, stylesheet, and font URL, including case-sensitive path names and URL encoding.

Custom CSS and print rules

from weasyprint import HTML

markup = '''
<style>
  @page { size: A4; margin: 18mm 15mm 20mm; }
  @page :first { margin-top: 10mm; }
  body { font-family: sans-serif; font-size: 10pt; }
  h1 { break-after: avoid; }
  .keep-together { break-inside: avoid; }
  .page-break { break-before: page; }
</style>
<h1>Quarterly report</h1>
<div class='keep-together'>A section that should stay together.</div>
'''

HTML(string=markup).write_pdf('report.pdf')

WeasyPrint uses print media by default. Use @page for paper size and margins, and print-oriented break properties for pagination. Test long tables, headings near page bottoms, floats, and nested flex or grid layouts because this is a paginated renderer rather than a full browser engine.

4. Fonts and international text

Fonts available through the system font configuration can be embedded and are subset by default. Install the fonts in the runtime image, verify the required weights and styles, and test representative glyphs for every language you generate. For @font-face, pass one shared FontConfiguration while applying CSS:

from weasyprint import HTML, CSS
from weasyprint.text.fonts import FontConfiguration

font_config = FontConfiguration()
css = CSS(string='''
@font-face {
  font-family: ReportFont;
  src: url('fonts/report-regular.woff2');
}
body { font-family: ReportFont, sans-serif; }
''', font_config=font_config, base_url='.')

HTML(string='<h1>Résumé — 日本語</h1>', base_url='.').write_pdf('fonts.pdf', stylesheets=[css], font_config=font_config)

5. A production-friendly conversion function

from pathlib import Path
from weasyprint import HTML


def html_to_pdf(markup: str, output: str, asset_root: str | None = None) -> None:
    kwargs = {'string': markup}
    if asset_root:
        kwargs['base_url'] = str(Path(asset_root).resolve())
    HTML(**kwargs).write_pdf(output)


html_to_pdf(
    '<!doctype html><h1>Monthly statement</h1>',
    'statement.pdf',
    asset_root='public',
)

For many documents, keep a long-lived Python process and call the API repeatedly instead of starting a new process for every PDF. This avoids repeated startup overhead according to the WeasyPrint documentation. Add your own queue, timeout, logging, and output validation around the function.

6. Handling remote resources safely

WeasyPrint’s documentation warns that untrusted HTML or CSS can create security problems. A document can trigger long or resource-intensive rendering and may reach files or network resources available to the process.

  • Run rendering in a dedicated unprivileged process or container; never run it as root.
  • Restrict filesystem access to an asset directory.
  • Allow only the URL schemes and hosts your application needs.
  • Set CPU, memory, process, file-size, and wall-clock limits.
  • Treat SVG files and CSS as untrusted input too.
  • Use a custom URL fetcher to enforce protocol, host, authentication, and path rules.
  • Decide whether missing images or stylesheets should fail the job; fetcher errors are generally logged as warnings by default.

Do not pass user-controlled URLs directly to HTML(url=...) without SSRF protections. Consider blocking private IP ranges, redirects to internal hosts, and unexpectedly large responses.

7. Choosing the input and output form

Input Call Use when
Markup string HTML(string=markup) Your application generated the HTML.
Local file HTML(filename=path) A template and its assets already exist on disk.
Remote page HTML(url=url) The source is an accessible HTTP or HTTPS document.

For output, provide a filename for direct persistence, a file object for controlled storage, or no target when you need bytes for an HTTP response or object-store upload.

8. Troubleshooting

Module or native-library installation errors

Cause: WeasyPrint or a required system library is missing from the active environment.

Fix: Activate the deployment virtual environment, install the documented OS dependencies, confirm Python and Pango versions, and run python -m pip show weasyprint from that same environment.

Images or styles are missing

Cause: Relative URLs have no correct base, the file is outside the allowed path, or the fetcher cannot access it.

Fix: Set base_url or a <base> element, use resolvable URLs, inspect warnings, and verify permissions and URL encoding.

Fonts show as fallback boxes or the wrong typeface

Cause: The font is not installed, the declared source cannot be fetched, or required glyphs are absent.

Fix: Install and register the font in the runtime, pass a shared FontConfiguration, verify each weight and style, and test the actual deployment image.

Pages break in unexpected places

Cause: Print pagination differs from screen layout, or content cannot fit in the available page box.

Fix: Set @page size and margins, use break-before, break-after, and break-inside, keep headings with following content, and test tables across multiple pages.

Remote pages need login cookies

Cause: The default HTTP fetcher does not support advanced cookies or authentication.

Fix: Fetch the protected HTML and assets yourself, or implement a custom URL fetcher that supplies only the required credentials and enforces host and protocol restrictions.

Conversion hangs or consumes too much memory

Cause: Very large documents, recursive resources, expensive CSS, huge images, or untrusted URLs.

Fix: Enforce input and resource limits, isolate the process, reject excessive document sizes, set job timeouts, and use a bounded worker queue.

The PDF does not match a browser screenshot

Cause: WeasyPrint is designed for print and PDF and does not implement every browser feature.

Fix: Use print CSS and supported layout features, inspect warnings, simplify unsupported effects, and choose a browser renderer when exact browser behavior is a requirement.

9. Performance, reliability, and cost planning

  • Performance: Reuse a long-lived Python process, avoid repeatedly downloading identical assets, resize oversized images before rendering, and keep CSS and DOM trees focused.
  • Reliability: Pin versions, keep fonts and assets in a known runtime, log warnings, validate that output bytes are non-empty, and retain representative PDF fixtures for regression checks.
  • Concurrency: Use bounded workers and memory limits. Measure your own documents; the official documentation does not provide a universal throughput benchmark.
  • Cost: WeasyPrint itself is software. Your practical costs come from compute, storage, network fetches, fonts, and operational isolation. Size workers from real document workloads.

10. Or skip the browser setup

If your goal is a screenshot or PDF of a live URL rather than server-side HTML-to-PDF rendering, ScreenshotNeo provides a single API request. Read the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing result. ScreenshotNeo also offers an MCP server so Claude, Cursor, and other MCP clients can take screenshots, inspect pages, and capture PDFs. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to get started.

11. FAQ

Does WeasyPrint require Selenium or Playwright?

No. The documented Python API converts HTML and CSS directly. A browser automation stack is only needed when you require browser-specific behavior that WeasyPrint does not implement.

Can I return a PDF from a web endpoint?

Yes. Call write_pdf() without a target, then return the resulting bytes with application/pdf and a suitable content-disposition header.

Why is base_url important?

It gives relative images, stylesheets, and fonts a location from which they can be resolved when the HTML is supplied as a string.

Should I use a custom fetcher for every project?

Use one whenever documents can reference untrusted or authenticated resources, or whenever you need strict protocol, host, timeout, or path controls.

Is WeasyPrint a general-purpose browser engine?

No. It is a print and PDF renderer. Validate complex layouts and browser-only features before adopting it for a particular document.