ScreenshotNeo

BlogComparisons

Best Python Libraries for Converting HTML to Images

Compare Playwright, html2image, and WeasyPrint for Python HTML-to-image work, with runnable code, troubleshooting, and a hosted API option.

By the ScreenshotNeo team1 October 20266 min read

Short answer: For browser-faithful HTML screenshots in Python, start with Playwright. It supports viewport, full-page, and element screenshots and can save PNG, JPEG, or WebP files or return bytes. Choose html2image for a small fixed-size wrapper around Chrome or Chromium. Choose WeasyPrint when the real output is a print-oriented PDF; raster images then require another conversion step.

This guide answers “How do I take a screenshot of an HTML page with Python?” and covers installation, rendering choices, deployment, troubleshooting, and a hosted alternative.

1. Choose the library by the rendering job

Library Best fit Strengths Constraint
Playwright Python Interactive, browser-rendered pages JavaScript, responsive layout, full-page and element capture, waits, PNG/JPEG/WebP, file or bytes output Install the Python package and compatible browser binaries.
html2image Simple fixed-size captures HTML/CSS strings, local files, or URLs through headless Chrome/Chromium Requires a browser; its documented API does not request full-page screenshots. Process only trusted content.
WeasyPrint Print layout and pagination HTML to PDF document generation PDF-first; raster output needs a separate step.

These are different workflows. Compare JavaScript needs, full-page versus viewport or element scope, input form, output format, browser setup, and whether a PDF intermediate is acceptable. The reviewed official sources provide no fair speed or fidelity benchmark, so validate representative pages yourself.

2. Playwright: the default for browser-faithful screenshots

Install

python -m pip install playwright
python -m playwright install chromium

The second command downloads browser binaries. Pin the package and browser versions together in CI or containers.

Capture a URL

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1440, "height": 900}, device_scale_factor=1)
    page.goto("https://example.com", wait_until="networkidle", timeout=60_000)
    page.screenshot(path="page.png", full_page=True, type="png")
    browser.close()

Use full_page=True for the entire scrollable document. Omit it for the current viewport. Use type="jpeg" with a quality value for smaller lossy files, or type="webp" where supported.

Capture an element and return bytes

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(viewport={"width": 1280, "height": 800})
    page.goto("https://example.com/dashboard", wait_until="domcontentloaded")
    page.locator("main .chart").screenshot(path="chart.webp", type="webp")
    image_bytes = page.locator("main .chart").screenshot(type="png")
    print(f"{len(image_bytes)} bytes")
    browser.close()

Locator screenshots suit cards, charts, invoices, and other regions where a full document adds noise.

Async API and deterministic waits

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page(viewport={"width": 1440, "height": 900})
        await page.goto("https://example.com", wait_until="domcontentloaded", timeout=60_000)
        await page.locator(".hero").wait_for(state="visible", timeout=15_000)
        await page.screenshot(path="hero.png", full_page=False)
        await browser.close()

asyncio.run(main())

Prefer a selector or application-ready signal over an arbitrary sleep. networkidle can be unsuitable for pages with analytics or open connections; use domcontentloaded plus a meaningful locator when needed.

HTML string and custom CSS

from pathlib import Path
from playwright.sync_api import sync_playwright

html = Path("template.html").read_text(encoding="utf-8")
with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(viewport={"width": 1200, "height": 800})
    page.set_content(html, wait_until="load")
    page.add_style_tag(content="body { background: white; } .debug { display:none !important; }")
    page.screenshot(path="render.png", full_page=True)
    browser.close()

For local assets, serve the directory over HTTP or verify that the browser can read correctly formed file:// URLs.

3. html2image: a compact fixed-size wrapper

python -m pip install html2image
from html2image import Html2Image

hti = Html2Image(output_path="captures", size=(1280, 720))
hti.screenshot(
    html="<h1>Invoice</h1><p>Ready</p>",
    css="body { font-family: sans-serif; padding: 40px; }",
    save_as="invoice.png",
)
hti.screenshot(url="https://example.com", save_as="example.png")

Set dimensions explicitly; the project documents a 1920×1080 default. Install and expose a supported Chrome or Chromium browser. Its documentation says there is no full-page screenshot request. Process only trusted content because unsanitized input can lead to malicious code execution.

4. WeasyPrint: when the deliverable is a PDF

from weasyprint import HTML

HTML(string="""
  <html><body><h1>Report</h1><p>Print layout</p></body></html>
""").write_pdf("report.pdf")
HTML(filename="report.html").write_pdf("report.pdf")

Use WeasyPrint for page breaks, paper size, margins, and print CSS. It is not evidenced as a direct PNG, JPEG, or WebP API here. Rasterize the PDF with a separate renderer when an image is required.

5. Production checklist

  1. Define viewport, full-document, or element scope and the required format.
  2. Pick Playwright for browser behavior, html2image for simple fixed-size jobs, or WeasyPrint for PDF pagination.
  3. Pin Python, package, and browser versions.
  4. Set viewport and device scale factor explicitly.
  5. Wait for a stable selector or ready signal.
  6. Write to a temporary path, validate the image, then move it atomically.
  7. Bound navigation and capture timeouts and close contexts in cleanup code.
  8. Isolate untrusted HTML and restrict its network or file access.

6. Troubleshooting

Symptom Cause Fix
Executable does not exist Playwright browsers were not installed. Run python -m playwright install chromium during the image build.
Blank or partial image Capture ran before data, fonts, or images loaded. Wait for a visible selector or ready signal and verify resource URLs.
Full page is clipped Viewport mode was used or content is in a nested scroller. Use full_page=True or screenshot the scrolling locator.
Element not found Selector changed, iframe boundary, or element is not mounted. Check the selector, wait for visibility, or access the correct frame.
CI output differs Browser, fonts, viewport, timezone, or scale differs. Pin versions and rendering parameters.
html2image cannot start Chrome Browser missing or undiscoverable. Install supported Chrome/Chromium and configure its path.
html2image stops at viewport Its API documents no full-page request. Use Playwright for full-document capture.
Untrusted HTML executes code Active content or local resources are available. Follow the trusted-content warning and isolate jobs.

7. Performance, reliability, and cost

  • Browser startup is expensive; reuse a browser process where safe and create isolated contexts for cookies and viewports.
  • Limit concurrency to available CPU and memory.
  • Use bounded timeouts and retry only transient navigation failures.
  • PNG preserves pixels; JPEG and WebP can reduce storage when lossy compression is acceptable.
  • Self-hosting shifts cost to browser binaries, CPU, memory, containers, and maintenance. No reviewed source establishes a speed benchmark.

8. Or skip the browser setup

ScreenshotNeo is a hosted website screenshot API. One request returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and X-Page-Verdict and X-Billed identify the result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for full-page and CSS-selector captures, dark mode, device presets or custom viewports, retina scale, PDF paper settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, async webhooks, bulk capture of up to 100 URLs, usage data, and OpenAPI details.

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account.

9. FAQ

Is Playwright faster than html2image?

The reviewed sources provide no fair benchmark. Measure your own pages, including startup and asset loading.

Can these libraries capture JavaScript-heavy sites?

Playwright is the browser automation choice here. html2image also uses headless Chrome or Chromium; WeasyPrint follows a PDF rendering model.

Which option creates PNG without a PDF step?

Playwright and html2image. WeasyPrint requires rasterization.

Should every page use full-page mode?

No. Use viewport captures for fixed visual checks and element captures for components.

What should I use for untrusted HTML?

Isolate the renderer and enforce your security policy; html2image explicitly advises trusted content only.