ScreenshotNeo

BlogHTML to image & PDF

HTML to PDF in Python: Code Examples

Convert HTML to PDF in Python with runnable WeasyPrint and Playwright examples, deployment guidance, troubleshooting, and a browser-free API option.

By the ScreenshotNeo team1 October 20267 min read

Python has two well-supported ways to convert HTML to PDF: WeasyPrint renders HTML and CSS directly, while Playwright drives a real Chromium page and calls page.pdf(). Use WeasyPrint for controlled, document-style HTML; use Playwright when your output depends on browser navigation, JavaScript, or browser layout behavior.

This guide gives complete examples, installation steps, configuration options, deployment advice, troubleshooting, and a browser-free ScreenshotNeo alternative.

1. Choose the rendering approach

Requirement Best starting point Reason
Generated reports with known HTML and CSS WeasyPrint Simple Python API and no browser process.
Pages that need JavaScript or navigation Playwright Renders a real Chromium page before creating the PDF.
Screen-accurate output from a deployed URL Playwright or ScreenshotNeo Both can render a browser page; ScreenshotNeo removes common overlays before capture.

There is no sourced universal winner for fidelity or speed. Test representative documents, including fonts, images, page breaks, links, and any required PDF conformance, before selecting a production path.

2. WeasyPrint: direct HTML and CSS to PDF

WeasyPrint exposes the shortest conversion path: create an HTML object and call write_pdf(). Its documented API accepts a string, URL, filename, or file object, and can return PDF bytes when no destination is supplied. See the WeasyPrint First Steps documentation for current requirements and platform instructions.

Install

python -m pip install weasyprint

The current documentation lists Python 3.10 or newer and Pango 1.44 or newer, plus platform-specific native libraries. Follow the installation instructions for your operating system and pin the version used in deployment.

Convert an HTML string

from weasyprint import HTML

html = """


  
    
    
  
  
    

Monthly report

Generated from HTML with Python.

""" HTML(string=html).write_pdf("report.pdf")

Read from a file, URL, or return bytes

from pathlib import Path
from weasyprint import HTML

# Local file. base_url lets relative CSS, images, and fonts resolve.
html = HTML(filename="report.html", base_url=str(Path("report.html").parent))
html.write_pdf("report.pdf")

# Return bytes for an HTTP response or object storage upload.
pdf_bytes = HTML(string="<h1>Invoice</h1>").write_pdf()
Path("invoice.pdf").write_bytes(pdf_bytes)

Set base_url whenever your HTML references relative assets such as images/logo.png or styles/report.css. For remote documents, pass a URL only when the referenced resources and network policy are suitable for your environment.

Useful WeasyPrint document controls

  • @page controls paper size, margins, and named pages.
  • page-break-before, page-break-after, and break-inside help keep sections and table rows together.
  • Embed or make fonts available to the rendering environment; a missing font changes wrapping and pagination.
  • Use absolute URLs or a correct base_url for images, stylesheets, and font files.
  • Generate a PDF in memory when your web framework streams the result, then set an appropriate PDF content type.

3. Playwright: render a browser page, then create a PDF

Playwright’s Python API creates or navigates a browser page and calls page.pdf(). The API uses print CSS media by default. To apply screen styles instead, call page.emulate_media(media="screen") before generating the PDF. See the Page API reference, Python library installation guide, and browser installation documentation.

Install the package and browser

python -m pip install playwright
python -m playwright install chromium

The Python package alone is not enough: the Chromium browser binaries must also be installed in the image, host, or build step that runs your code.

Convert HTML content

from playwright.sync_api import sync_playwright

html = """
<!doctype html>
<html>
  <head>
    <style>
      @page { size: A4; margin: 16mm; }
      body { font-family: Arial, sans-serif; }
    </style>
  </head>
  <body>
    <h1>Monthly report</h1>
    <p>Rendered by Chromium.</p>
  </body>
</html>
"""

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.set_content(html, wait_until="networkidle")
    page.pdf(path="report.pdf", format="A4", print_background=True)
    browser.close()
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(viewport={"width": 1440, "height": 900}, device_scale_factor=1)
    page.goto("https://example.com", wait_until="networkidle")
    page.emulate_media(media="screen")
    page.pdf(
        path="page.pdf",
        format="A4",
        print_background=True,
        margin={"top": "12mm", "right": "12mm", "bottom": "12mm", "left": "12mm"},
    )
    browser.close()

Important Playwright PDF options

  • format selects a standard paper size; alternatively provide width and height.
  • landscape=True rotates the page.
  • margin accepts CSS length strings for each edge.
  • print_background=True includes background colors and images.
  • page_ranges="1-3" limits output to selected pages.
  • prefer_css_page_size=True lets CSS @page size win over a supplied format.
  • display_header_footer=True enables header and footer templates; use the documented template variables for page numbers and titles.

4. A production-friendly Playwright function

from pathlib import Path
from playwright.sync_api import sync_playwright

def html_to_pdf(url: str, output: str, *, media: str = "print") -> None:
    with sync_playwright() as p:
        browser = p.chromium.launch()
        try:
            page = browser.new_page()
            page.goto(url, wait_until="networkidle", timeout=60_000)
            if media in {"screen", "print"}:
                page.emulate_media(media=media)
            page.pdf(
                path=output,
                format="A4",
                print_background=True,
                prefer_css_page_size=True,
                margin={"top": "15mm", "right": "15mm", "bottom": "15mm", "left": "15mm"},
            )
        finally:
            browser.close()

html_to_pdf("https://example.com/report", "report.pdf", media="screen")

For dynamic pages, wait for a meaningful selector instead of assuming network idle means all application work is complete:

page.goto(url, wait_until="domcontentloaded")
page.wait_for_selector("#report-ready", state="visible", timeout=30_000)
page.pdf(path="report.pdf", print_background=True)

5. Security and input handling

Do not treat arbitrary user-supplied HTML or CSS as safe to render. The WeasyPrint documentation explicitly warns: “Using WeasyPrint with untrusted HTML or untrusted CSS may lead to various security problems.” Validate or sanitize markup, restrict network access, control allowed URL schemes, and isolate rendering workers. Browser rendering also needs limits for navigation, scripts, memory, and outbound requests.

6. Deployment, performance, and reliability

  • Dependencies: WeasyPrint needs native text and layout libraries; Playwright needs browser binaries. Build and test the same container or image used in production.
  • Startup cost: Reusing a browser process and creating pages per job generally avoids repeatedly launching Chromium. Always close pages and browsers in error paths.
  • Concurrency: Bound simultaneous renders. PDFs with large images, complex CSS, or long pages can consume substantial CPU and memory.
  • Timeouts: Set navigation and selector timeouts. Return a clear failure when a page never reaches its ready condition.
  • Determinism: Pin Python, library, browser, and font versions. External assets and JavaScript can change output between runs.
  • Validation: Open generated files, check that they begin with the PDF signature, and verify page count, fonts, links, and expected content for representative inputs.
  • Cost: Self-hosted conversion costs the compute, storage, browser or native dependencies, and operational time required by your workload. Measure your own documents; the research does not establish a universal speed comparison.

7. Troubleshooting common failures

Symptom Likely cause Fix
ModuleNotFoundError: weasyprint Package is not installed in the active environment. Install it with the same Python interpreter used to run the script.
WeasyPrint cannot load or import native libraries Missing Pango or another platform dependency. Follow the OS-specific installation steps in the current WeasyPrint documentation.
Playwright reports that an executable is missing Browser binaries were not installed. Run python -m playwright install chromium during image setup.
Styles look different from the web page PDF uses print media by default. Call page.emulate_media(media="screen") or add print-specific CSS intentionally.
Images or CSS are missing Relative URLs have no base, or the renderer cannot reach remote assets. Set WeasyPrint base_url; use resolvable URLs; check network and certificate policy.
Blank or incomplete dynamic page PDF was created before the application finished rendering. Wait for a ready selector or explicit application state before calling page.pdf().
Unexpected page breaks Content dimensions, fonts, or CSS break rules changed. Pin fonts, use @page and break properties, and test the actual document.
Process hangs or is killed Unbounded navigation, scripts, concurrency, or document size. Set timeouts, restrict inputs and requests, cap concurrency, and isolate workers.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API that can return PNG, JPEG, WebP, or PDF from one GET request. It handles the hosted browser and provides PDF controls such as paper size, margins, landscape mode, and page ranges. Before capture, cookie and consent banners, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the verdict and billing status in X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all parameters. The same endpoint can also accept custom headers, cookies, user agents, authorization, waits, CSS, JavaScript, blocking rules, caching, signed links, asynchronous jobs, and bulk capture.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

An MCP server also lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

9. FAQ

Can WeasyPrint execute JavaScript?

Choose a browser workflow such as Playwright when the page depends on JavaScript execution or browser navigation.

Why does Playwright output look different from my screen?

page.pdf() uses print media by default. Emulate screen media when that is the intended design, and keep print-specific CSS where appropriate.

Should I render untrusted HTML in the same process as my application?

No. Sanitize and validate input, restrict network access, and isolate rendering workers because HTML, CSS, and browser scripts can create security and resource risks.

How do I preserve relative assets?

For WeasyPrint, provide a correct base_url. For Playwright, navigate to a page whose asset URLs resolve in the browser context.