ScreenshotNeo

BlogHow-to

How to Generate an Image from HTML in Python

Render HTML as a browser page or document, then save PNG, JPEG, or WebP images with complete Python examples, options, fixes, and deployment notes.

By the ScreenshotNeo team1 October 20267 min read

Use Playwright when your HTML should look like a browser page. It runs Chromium, applies browser CSS and JavaScript, waits for page state, and saves a PNG, JPEG, or WebP screenshot. Use a locator screenshot for one component, full_page=True for the entire document, or omit the path to keep image bytes in memory. The official setup requires both the Python package and browser binaries: Playwright installation documentation.

For print-style, document-oriented output, evaluate WeasyPrint instead. It lays out and paginates HTML and CSS without a browser, but JavaScript-dependent pages need Playwright.

1. Install Playwright and Chromium

python -m venv .venv
source .venv/bin/activate       # Windows: .venv\\Scripts\\activate
pip install playwright
playwright install chromium

The second command downloads browser binaries. Include that step in your deployment image or build process. The synchronous and asynchronous APIs are both supported.

2. Generate a PNG from an HTML string

from playwright.sync_api import sync_playwright

html = """
<!doctype html>
<html>
  <head>
    <meta charset="utf-8">
    <style>
      body { font-family: system-ui, sans-serif; margin: 40px; }
      h1 { color: #1f2937; }
    </style>
  </head>
  <body><h1>Hello from HTML</h1></body>
</html>
"""

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(viewport={"width": 1280, "height": 720})
    page.set_content(html)
    page.screenshot(path="output.png", full_page=True)
    browser.close()

page.set_content() loads the supplied markup. Use page.goto() for a URL. The screenshot API is documented in Playwright’s Python screenshot guide.

3. Control the output format, size, and quality

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(
        viewport={"width": 1440, "height": 900},
        device_scale_factor=2,
    )
    page.set_content("<main>Retina output</main>")

    page.screenshot(path="image.png", type="png", full_page=True)
    page.screenshot(path="image.jpg", type="jpeg", quality=85)
    page.screenshot(path="image.webp", type="webp", quality=80)
    browser.close()
  • type supports PNG, JPEG, and WebP.
  • quality applies to JPEG and WebP.
  • device_scale_factor controls CSS pixels versus device pixels.
  • Transparent backgrounds are available for applicable image types; set the page background to transparent when needed.
  • JPEG does not preserve transparency.

4. Capture a full page, one element, or image bytes

Full-page capture

page.screenshot(path="long-page.png", full_page=True)

One component

card = page.locator(".product-card")
card.screenshot(path="card.png")

The locator must match a visible, stable element. Playwright scrolls it into view, but covered content is not captured. For a scrollable container, the screenshot contains the content currently scrolled into view rather than an automatically expanded version.

Keep bytes in memory

png_bytes = page.screenshot(full_page=True)
# send png_bytes to storage, an HTTP response, or an image processor

5. Render dynamic HTML reliably

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto("https://example.com", wait_until="domcontentloaded")
    page.wait_for_selector(".report", state="visible")
    page.wait_for_load_state("networkidle")
    page.screenshot(path="report.png", full_page=True)
    browser.close()

Choose a wait condition that matches the page. A selector is usually more precise than a fixed delay. Use a short, deliberate delay only when an animation or client-side render needs time. Disable animations when deterministic output matters:

page.add_style_tag(content="""
*, *::before, *::after {
  animation: none !important;
  transition: none !important;
  caret-color: transparent !important;
}
""")

For lazy-loaded images, scroll through the page before capturing, or wait for the image selectors you need. Keep browser version, fonts, viewport, assets, and page state consistent if you need repeatable pixels across machines.

6. Use asynchronous Python for concurrent jobs

import asyncio
from playwright.async_api import async_playwright

async def render():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page(viewport={"width": 1280, "height": 720})
        await page.set_content("<h1>Async HTML</h1>")
        data = await page.screenshot(type="png", full_page=True)
        await browser.close()
        return data

image_bytes = asyncio.run(render())
open("async-output.png", "wb").write(image_bytes)

Reuse a browser process for a batch, but create isolated pages or contexts for separate jobs. Close pages and browsers in a finally block in long-running workers.

7. Render a URL with headers, cookies, or a custom viewport

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(
        viewport={"width": 1366, "height": 768},
        locale="en-US",
        timezone_id="UTC",
        extra_http_headers={"Authorization": "Bearer TOKEN"},
    )
    context.add_cookies([{
        "name": "session",
        "value": "VALUE",
        "domain": "example.com",
        "path": "/",
    }])
    page = context.new_page()
    page.goto("https://example.com/dashboard", wait_until="networkidle")
    page.screenshot(path="dashboard.png", full_page=True)
    browser.close()

Keep secrets out of HTML and source control. For pages that depend on a specific locale, timezone, or user agent, set those values in the browser context so the rendering environment is explicit.

8. Document-style HTML with WeasyPrint

from weasyprint import HTML

html = """
<html>
  <body>
    <h1>Invoice</h1>
    <p>Document layout with pagination.</p>
  </body>
</html>
"""

HTML(string=html, base_url=".").write_png("invoice.png")

WeasyPrint accepts HTML from strings, URLs, filenames, or file objects. Supply base_url when a string contains relative images, stylesheets, or fonts. Confirm that the HTML and CSS you use are supported. Long or specially crafted documents can take a long time to render, so measure your real input; the inspected documentation provides no controlled speed or visual-fidelity benchmark against Playwright.

9. Choose the right method

Requirement Use Reason
Browser CSS, JavaScript, or an existing website Playwright Chromium renders the page before capture.
One stable component Playwright locator screenshot Captures the selected visible element.
In-memory processing Playwright screenshot without a path Returns image bytes.
Paginated reports or print layout WeasyPrint Provides document layout and pagination.

10. Troubleshooting

“Executable doesn’t exist” or browser launch failure

Cause: the Python package is installed but browser binaries are not. Fix: run playwright install chromium during setup, and verify the deployment user can read the browser cache.

Blank or incomplete screenshot

Cause: capture happened before client-side rendering or lazy assets completed. Fix: wait for a meaningful selector, a load state, or the specific images; avoid relying only on a short timeout.

Element screenshot throws because no element is found

Cause: the selector is wrong, hidden, or rendered later. Fix: check the locator, wait for state="visible", and ensure overlays do not cover it.

Relative images or CSS are missing in WeasyPrint

Cause: an HTML string has no base URL. Fix: pass base_url or use absolute resource URLs.

Fonts differ between local and production

Cause: different installed fonts or browser images. Fix: package the same fonts and browser version, then set explicit font stacks.

Huge output or slow full-page capture

Cause: very tall pages, large assets, animations, or expensive scripts. Fix: capture an element when possible, reduce asset sizes, disable animations, and process jobs with bounded concurrency.

Cause: the page state includes an overlay. Fix: click the consent action, remove or hide the overlay with page code, or use a capture service that handles these elements before the screenshot.

11. Performance, reliability, and cost considerations

  • Browser startup is expensive compared with reusing an existing browser process. Reuse browsers while isolating jobs in separate contexts.
  • Full-page screenshots consume memory proportional to page dimensions and asset size. Prefer element captures for cards, charts, and thumbnails.
  • Network-dependent pages can vary by time, region, cookies, and third-party scripts. Pin the viewport, browser, fonts, locale, timezone, and test data for stable output.
  • Set application-level timeouts and retry only transient navigation failures. Repeatedly retrying a broken page increases latency and load.
  • Playwright itself has no per-screenshot charge; your costs are compute, browser storage, bandwidth, and any hosted rendering service.
  • WeasyPrint may be lighter for static documents, but workload time depends on the actual HTML and CSS.

12. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One request returns a PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses identify the result with X-Page-Verdict and X-Billed headers. AI agents can use the MCP tools take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page and element capture, dark mode, device presets, custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, async jobs, bulk capture, usage data, and an OpenAPI specification. Plans include 1,000 free shots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. FAQ

Can I create a screenshot without saving a file?

Yes. Omit the Playwright path argument; the method returns image bytes.

Should I use PNG, JPEG, or WebP?

Use PNG for lossless UI and transparency, JPEG for photographs where size matters, and WebP when your consumers support it and you want quality control with smaller files.

Can Playwright capture only part of a page?

Yes. Use a locator screenshot for a visible element. A scrollable element captures the portion currently in view.

Why is my output different on another machine?

Browser versions, fonts, viewport, assets, locale, timezone, and dynamic data can all change pixels. Control those inputs for reproducible images.

When is WeasyPrint a better fit?

Use it for static, document-oriented HTML where pagination matters and JavaScript is unnecessary. Validate support for your CSS and resource URLs.