How to Convert Raw HTML to PDF in Python
Convert an in-memory HTML string to a PDF in Python with WeasyPrint, including CSS, assets, security, troubleshooting, and production options.

Direct answer: For an HTML string already held in Python, use WeasyPrint’s named string argument and then call write_pdf(). The method can return PDF bytes or write directly to a file.
from weasyprint import HTML
html = '''
<!doctype html>
<html>
<head>
<meta charset="utf-8">
<style>
@page { size: A4; margin: 20mm; }
body { font-family: sans-serif; color: #222; }
h1 { color: #1457a6; }
</style>
</head>
<body>
<h1>Invoice</h1>
<p>Generated from an in-memory HTML string.</p>
</body>
</html>
'''
pdf_bytes = HTML(string=html).write_pdf()
with open('invoice.pdf', 'wb') as pdf_file:
pdf_file.write(pdf_bytes)
The named argument matters. HTML(string=html) tells WeasyPrint that the value is markup. Passing a string positionally can make it interpret the value as a filename or URL. Calling write_pdf() without a destination returns a byte string; passing a path writes the file directly. See the WeasyPrint first-steps documentation and ScreenshotNeo documentation for the APIs used in this guide.
1. Install WeasyPrint
Install the Python package in a virtual environment:
python -m venv .venv
source .venv/bin/activate
python -m pip install weasyprint
On some operating systems, WeasyPrint also needs native libraries for text layout, fonts, and image processing. Follow the installation instructions for your platform in the project documentation. In a deployment image, install those system packages explicitly and pin the Python package version so an environment rebuild does not silently change rendering.
2. Convert an HTML string to PDF bytes
Returning bytes is useful when the PDF will be sent through an HTTP response, stored in object storage, attached to an email, or passed to another Python component.

from weasyprint import HTML
html = '''
<html>
<head>
<meta charset="utf-8">
<title>Report</title>
<style>
body { font-family: Arial, sans-serif; line-height: 1.45; }
.total { font-size: 20px; font-weight: 700; }
</style>
</head>
<body>
<h1>Monthly report</h1>
<p>This document was generated without creating an HTML file first.</p>
<p class="total">Total: $125.00</p>
</body>
</html>
'''
pdf_bytes = HTML(string=html).write_pdf()
# Example: save the returned bytes
with open('report.pdf', 'wb') as output:
output.write(pdf_bytes)
Keep the output in binary mode. Opening the destination with text mode can corrupt the PDF.
3. Write directly to a destination file
If you do not need the bytes in memory, provide a path to write_pdf():
from weasyprint import HTML
html = '<h1>Invoice</h1><p>Paid</p>'
HTML(string=html).write_pdf('invoice.pdf')
The destination directory must already exist and the running process must have permission to write there. For a web service, prefer a temporary directory or a controlled storage layer instead of accepting arbitrary paths from a request.
4. Make relative CSS, images, and fonts resolve
An HTML string has no filename, so relative references such as styles/report.css or images/logo.png need a base URL. Pass base_url that points to the directory containing those resources:

from pathlib import Path
from weasyprint import HTML
html = '''
<link rel="stylesheet" href="css/report.css">
<img src="images/logo.png" alt="Company logo">
<h1>Report</h1>
'''
project_dir = Path('/srv/my-report')
HTML(string=html, base_url=project_dir.as_uri()).write_pdf('report.pdf')
WeasyPrint’s default fetcher can retrieve local files and HTTP resources, but its HTTP client does not provide advanced features such as cookies or authentication. If assets require special headers, signed URLs, or application credentials, use a custom URL fetcher or make the resources available through a controlled, accessible location. Resource loading is part of PDF generation; a valid HTML string does not guarantee that every external asset will resolve.
Embedding data directly
For small images or fonts, data URLs remove a separate network request:
html = '''
<style>
@font-face {
font-family: ReportFont;
src: url(data:font/woff2;base64,BASE64_FONT_DATA);
}
</style>
<img src="data:image/png;base64,BASE64_IMAGE_DATA">
'''
Embedding increases the HTML size, so use it selectively.
5. Control page size, margins, and page breaks
PDF layout is controlled with CSS at-rules. The @page rule sets paper size and margins; page-break properties keep related content together.
html = '''
<style>
@page {
size: Letter;
margin: 18mm 16mm 22mm;
@bottom-right {
content: "Page " counter(page) " of " counter(pages);
font-size: 9pt;
}
}
h1 { break-before: page; }
.invoice-total { break-inside: avoid; }
table { width: 100%; border-collapse: collapse; }
tr { break-inside: avoid; }
</style>
'''
Use the paper size expected by your recipients, such as A4 or Letter. Long tables can still split across pages; design table headers and row content so a split remains readable. Test documents with unusually long text, empty sections, and large tables.
6. Generate a PDF in a web endpoint
Because write_pdf() returns bytes, a framework can send the result without an intermediate file. Here is a minimal Flask example:
from flask import Flask, Response, request
from weasyprint import HTML
app = Flask(__name__)
@app.post('/pdf')
def create_pdf():
html = request.get_data(as_text=True)
if not html.strip():
return {'error': 'HTML body is required'}, 400
pdf_bytes = HTML(string=html).write_pdf()
return Response(
pdf_bytes,
mimetype='application/pdf',
headers={'Content-Disposition': 'attachment; filename="document.pdf"'}
)
if __name__ == '__main__':
app.run()
In production, place limits on request size and rendering time. Do not allow callers to choose arbitrary filesystem destinations.
7. Choose the right renderer
WeasyPrint is a direct fit when your input is an HTML string and you want a Python API for HTML/CSS-to-PDF conversion. It is not a browser automation tool, so confirm that the CSS and PDF features your templates need are supported.
pdfkit and wkhtmltopdf
If an existing project uses wkhtmltopdf, pdfkit exposes a string workflow:
import pdfkit
html = '<h1>Invoice</h1>'
pdfkit.from_string(html, 'invoice.pdf')
pdfkit is a wrapper; the separate wkhtmltopdf executable must be installed or configured. The pdfkit README documents executable discovery and configuration. Account for that binary in local setup, containers, and deployment.
Playwright Python
Use Playwright when the document must be rendered by a browser page or when your workflow already needs browser automation:
from playwright.sync_api import sync_playwright
html = '<html><body><h1>Report</h1></body></html>'
with sync_playwright() as playwright:
browser = playwright.chromium.launch()
page = browser.new_page()
page.set_content(html)
page.pdf(path='report.pdf', format='A4')
browser.close()
Playwright’s page.pdf() uses print media by default. To use screen styles, call page.emulate_media(media='screen') before generating the PDF. The Playwright Page API documents the available PDF controls.
8. Handle dynamic content and external resources
An HTML string can contain markup, styles, and references, but it does not automatically provide application state. Resolve template variables before rendering, and make sure images, stylesheets, and fonts are reachable from the renderer.
- Use absolute HTTPS URLs or a correct
base_urlfor relative references. - Bundle critical styles and fonts when deterministic output matters.
- Do not assume JavaScript will run in the same way as it does in a browser. If the page needs browser execution before printing, use a browser renderer such as Playwright.
- Use explicit character encoding, normally UTF-8, and include
<meta charset="utf-8">. - Define fallback fonts for characters that may not exist in the primary font.
9. Security for untrusted HTML
WeasyPrint’s documentation warns that untrusted HTML or CSS can create security problems. Treat submitted markup as hostile input. Sanitize HTML, constrain CSS, restrict network access, and prevent access to sensitive local files. A custom URL fetcher can enforce an allowlist of hosts and schemes. Run rendering with a least-privileged account and impose CPU, memory, input-size, and wall-clock limits.
Do not interpolate unescaped user values into HTML templates. Escape text, validate URLs, and keep secrets out of the document and its resource URLs. If users can provide CSS, review features that can trigger expensive layout or unexpected resource fetching.
10. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| The string is treated as a file | The constructor received a positional string. | Use HTML(string=html). |
| Images or CSS are missing | Relative URLs have no base directory, or the resource is inaccessible. | Pass base_url, use absolute URLs, or embed critical assets. |
| Fonts show as boxes or fallbacks | The font is unavailable or does not contain the required glyphs. | Install the font, provide a reachable @font-face, and add a fallback. |
| HTTP assets return unauthorized | Default fetching lacks your application’s cookies or auth headers. | Use public or signed asset URLs, prefetch assets, or implement a custom URL fetcher. |
| Layout differs after an upgrade | Rendering behavior or supported features changed between releases. | Pin versions and render representative fixtures during upgrades. |
| PDF generation fails in a container | Required native libraries or fonts are absent. | Follow the platform installation guide and include system dependencies in the image. |
| Pages are unexpectedly blank | Invalid markup, inaccessible resources, or CSS that creates an unexpected layout. | Reduce the document to a minimal case, validate inputs, and inspect resource URLs. |
| Requests hang | A remote resource is slow or unreachable. | Use bounded fetches, an allowlist, local assets, and an application-level timeout. |
11. Performance, reliability, and cost
Rendering time depends on document size, CSS complexity, fonts, images, and remote resources. Avoid downloading the same large assets repeatedly; cache immutable assets or bundle them. For high-volume jobs, queue work and limit concurrent renderers so one large document cannot exhaust memory.
For reliability, keep a small set of representative HTML fixtures and compare generated PDFs after dependency changes. Include long text, tables spanning pages, right-to-left or non-Latin characters if your application uses them, missing optional images, and authenticated assets. WeasyPrint documents limitations and notes that generated documents are not guaranteed to be valid for every combination of HTML, CSS, and PDF features; treat version upgrades as rendering changes that require review.
There is no universal speed or fidelity winner among WeasyPrint, wkhtmltopdf, and Playwright based on the cited documentation. Choose according to the input model, CSS needs, browser requirements, installation footprint, and resource handling you can operate. Measure your own templates if latency or throughput determines the design.
12. Or skip the browser setup
If your real goal is to capture a public web page as an image or PDF rather than convert an in-memory HTML string, ScreenshotNeo provides a single HTTP request. It handles cookie and consent banners before capture, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets each cleanup step be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
Read the full ScreenshotNeo API documentation for options such as full-page capture, CSS-selector element capture, dark mode, device presets, retina scale, PDF paper size and margins, custom CSS and JavaScript, waits, blocked resources, headers, cookies, user agents, timezone, geolocation, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);
An MCP server also exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I return PDF bytes without creating a file?
Yes. Call HTML(string=html).write_pdf() without a destination and use the returned bytes in a response, storage client, or email attachment.
Why should I use the named string argument?
It unambiguously identifies the value as HTML markup. A positional string can be interpreted as a filename or URL.
Does WeasyPrint execute JavaScript?
Choose a browser renderer such as Playwright when your output depends on browser JavaScript execution. WeasyPrint is intended for HTML and CSS rendering.
How do I preserve screen styles in Playwright?
Call page.emulate_media(media='screen') before page.pdf(); PDF generation otherwise uses print media by default.
Should I accept arbitrary HTML from users?
Only with strict sanitization, resource restrictions, and execution limits. Untrusted HTML and CSS can create security and availability risks.


