ScreenshotNeo

BlogHTML to image & PDF

Convert an HTML File to PDF with Python

Convert a local HTML file to PDF with Python using WeasyPrint, handle assets and failures, and compare a hosted screenshot-to-PDF option.

By the ScreenshotNeo team1 October 20267 min read

Use WeasyPrint for a local HTML file:

from weasyprint import HTML

HTML(filename="input.html").write_pdf("output.pdf")

Install it in the Python environment that will run the conversion:

python -m pip install weasyprint

This is the documented Python API pattern. The rest of this guide covers installation, relative assets, print CSS, page control, security, troubleshooting, batch jobs, and a hosted alternative when you do not want to manage a browser or renderer.

1. Install WeasyPrint

Install into an activated virtual environment or your deployment environment:

python -m venv .venv
source .venv/bin/activate
python -m pip install weasyprint
weasyprint --info

Current WeasyPrint documentation lists Python 3.10 or later and Pango 1.44 or later, along with Python and native dependencies. On Linux, your distribution package manager may provide the native libraries more easily than pip alone. Check the official installation and first-steps documentation for the operating system and version you deploy.

2. Convert a local HTML file

Given this layout:

project/
  input.html
  styles.css
  images/
    logo.png
  convert.py

Use an explicit filename:

from weasyprint import HTML

HTML(filename="input.html").write_pdf("output.pdf")

You can also pass the path positionally:

from weasyprint import HTML

HTML("input.html").write_pdf("output.pdf")

Run it from the directory containing the file:

python convert.py

The result is written as output.pdf. Create the destination directory first if it does not exist.

3. Make relative CSS, images, and fonts resolve

HTML files commonly reference resources with relative URLs such as styles.css or images/logo.png. Keep the resource paths consistent with the HTML file’s location and inspect the generated PDF for missing assets.

<!doctype html>
<html>
<head>
  <meta charset="utf-8">
  <link rel="stylesheet" href="styles.css">
</head>
<body>
  <h1>Invoice</h1>
  <img src="images/logo.png" alt="Company logo">
</body>
</html>

For a script that may run from another working directory, construct an absolute input path and pass it to HTML:

from pathlib import Path
from weasyprint import HTML

base = Path(__file__).resolve().parent
HTML(filename=str(base / "input.html")).write_pdf(str(base / "output.pdf"))

Verify URLs, file permissions, image formats, and font paths when an asset is absent. The documented filename input does not guarantee that every arbitrary resource arrangement or CSS feature will render as intended, so review a representative PDF before relying on it.

4. Add print CSS for predictable pages

Use a print media block for page size, margins, colors, and page breaks:

@page {
  size: A4;
  margin: 18mm 16mm 20mm;
}

@media print {
  body {
    color: #111;
    background: white;
  }

  .page-break {
    break-before: page;
  }

  .avoid-split {
    break-inside: avoid;
  }
}

Keep long tables and cards together where possible with break-inside: avoid, and insert deliberate section breaks with break-before: page. Always inspect page breaks because a renderer’s supported CSS subset and the available space determine the final layout.

5. Convert HTML held in a string

If your application already has HTML text, pass it to HTML(string=...):

from weasyprint import HTML

html = """
<!doctype html>
<html>
  <body><h1>Report</h1><p>Generated document</p></body>
</html>
"""

HTML(string=html).write_pdf("report.pdf")

For user supplied content, treat the HTML and CSS as untrusted input. The WeasyPrint documentation warns that using untrusted HTML or untrusted CSS can lead to security problems. Apply the project’s security guidance, isolate the renderer where appropriate, and constrain network and filesystem access.

6. Choose output bytes instead of a file

write_pdf can return PDF bytes when you omit the output filename. This is useful for an HTTP response or object storage upload:

from weasyprint import HTML

pdf_bytes = HTML(filename="input.html").write_pdf()

with open("output.pdf", "wb") as pdf_file:
    pdf_file.write(pdf_bytes)

Set the response content type to application/pdf in your web framework and choose a download filename separately.

7. Command-line conversion

The installed package also provides a command-line interface:

weasyprint input.html output.pdf

Use the Python API when conversion is part of application logic, needs validation, or must return bytes. Use the command line for a simple build or automation step.

8. What WeasyPrint can and cannot guarantee

WeasyPrint supports PDF content such as text, raster and vector graphics, hyperlinks, bookmarks, attachments, and forms according to its API documentation. That capability list does not mean every source document will transfer exactly.

The project also states that generated-document validity is not guaranteed for every combination of HTML, CSS, and PDF features. Treat the renderer as an implementation with limits: test your actual templates and inspect fonts, images, links, page breaks, forms, and print-specific layout. Do not assume that an arbitrary modern web page will be pixel-perfect.

9. Troubleshooting

Symptom Likely cause Fix
Import error or native-library error WeasyPrint or a required system dependency is missing from the active environment. Run python -m pip install weasyprint, check Python and Pango versions, install the platform’s native packages, then run weasyprint --info.
CSS is missing The stylesheet URL is wrong relative to the HTML file or the process working directory. Check the HTML directory, use a stable absolute filename, and verify the stylesheet path.
Images do not appear Image paths, permissions, or formats are invalid. Open each referenced path independently, use correct relative URLs, and check file permissions and supported image data.
Fonts fall back The requested font is unavailable or its URL cannot be resolved. Install or provide the font in the deployment environment, verify its URL, and inspect the PDF on the target system.
Unexpected page breaks Content height, print CSS, or unsupported layout rules change pagination. Add explicit @page rules, use break-before/break-inside, simplify the layout, and inspect a representative output.
Blank or incomplete output The input is malformed, a resource failed, or the chosen HTML/CSS feature is outside implementation limits. Reduce the document to a minimal case, validate the HTML, check resource paths, and compare the result with the documented feature limits.
Conversion hangs or consumes excessive resources Large documents, expensive assets, or uncontrolled untrusted input. Set an application-level timeout, limit input size and remote access, isolate untrusted work, and process large jobs separately.

10. Batch conversion and performance

For repeated conversions, keep a long-lived Python API process instead of starting a new interpreter for every file. The WeasyPrint documentation describes this as a way to avoid repeated startup costs; it does not establish a universal speedup percentage.

from pathlib import Path
from weasyprint import HTML

source_dir = Path("html")
output_dir = Path("pdf")
output_dir.mkdir(exist_ok=True)

for source in source_dir.glob("*.html"):
    destination = output_dir / (source.stem + ".pdf")
    HTML(filename=str(source)).write_pdf(str(destination))

For reliable batch jobs:

  • Reuse the process and log the source and destination for every conversion.
  • Write to a temporary filename, then rename it after a successful conversion.
  • Keep representative fixtures for page breaks, fonts, images, and long tables.
  • Bound concurrency based on memory usage rather than assuming more workers are always faster.
  • Record renderer and system dependency versions so a deployment change can be diagnosed.

11. Security checklist

  • Do not render untrusted HTML or CSS without reading and applying WeasyPrint’s security guidance.
  • Restrict access to local files and network resources when processing user input.
  • Limit document size, resource count, and execution time.
  • Run untrusted conversion in an isolated worker with minimal permissions.
  • Do not embed secrets in HTML, CSS, URLs, or generated logs.

12. Or skip the browser setup

If your source is a public web page and you want a hosted capture instead of maintaining a Python renderer, ScreenshotNeo exposes a GET endpoint that returns a PDF or image. It handles consent banners before capture, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you control each step. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and whether the request was billed.

See the ScreenshotNeo API documentation for the full option list, including PDF paper size, margins, landscape mode, page ranges, waits, custom headers and cookies, blocking rules, caching, async jobs, bulk capture, and signed links.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const pdf = Buffer.from(await res.arrayBuffer());
await fs.promises.writeFile('shot.pdf', pdf);

Use this route when you need cookie banners, popups, and chat widgets removed before the shot; failed loads and bot checks never billed; or an MCP server that lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

13. Practical decision guide

Need Route
Local HTML with controlled files and Python deployment WeasyPrint’s HTML(filename=...).write_pdf(...).
HTML generated in memory WeasyPrint’s HTML(string=...).write_pdf().
Public URLs, consent cleanup, and hosted capture ScreenshotNeo’s PDF endpoint.
Many URLs or AI-agent workflows ScreenshotNeo bulk capture or its MCP server.

FAQ

Does WeasyPrint require a browser?

No browser session is shown in the basic Python API. It is a Python package with Python and native dependencies, including Pango requirements documented by the project.

Can I convert a URL instead of a local file?

WeasyPrint can work with HTML inputs, but this guide’s verified minimal route is a local filename. For a hosted URL capture with PDF output, use ScreenshotNeo.

Why does the PDF differ from Chrome?

Different renderers support different HTML, CSS, and print features. Test your templates and follow the implementation limits documented by WeasyPrint.

How do I make a PDF in a web request?

Generate bytes with HTML(...).write_pdf(), return them with the application/pdf content type, and enforce request and input limits around the conversion.

Is there a universal WeasyPrint performance number?

No benchmark is established by the supplied documentation. A long-lived API process can avoid repeated startup costs, but measure your own documents and deployment.