ScreenshotNeo

BlogHTML to image & PDF

Save a Webpage as a PDF with Python and Playwright

Use Playwright’s Python `page.pdf()` to save a rendered webpage as a PDF. Set print styles, paper size, margins, backgrounds, and output handling.

By the ScreenshotNeo team4 October 20269 min read

Use Playwright’s Python page.pdf() method to render the current page in Chromium and save it as a PDF. Install Playwright and its browser, navigate to a fully qualified URL, then call page.pdf(path="page.pdf"). By default, the PDF uses print CSS, omits background graphics, and uses no margins. Set those choices explicitly when the page needs different output. Playwright Page API.

1. Install Playwright and Chromium

In a virtual environment, install the Python package and the browser binaries Playwright needs. The browser installation is a separate step from installing the Python package. See the official Playwright Python library guide.

python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
python -m playwright install chromium

On Windows PowerShell, activate the environment with .venv\Scripts\Activate.ps1. If your system needs operating-system dependencies for Chromium, consult the Playwright browser installation guide.

2. Save a webpage as a PDF

This complete synchronous script navigates to a URL, waits for the page load event, and writes an A4 PDF with background graphics. Save it as save_page.py and run python save_page.py.

from pathlib import Path
from playwright.sync_api import sync_playwright

URL = "https://example.com"
OUTPUT = Path("page.pdf")

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page()
        response = page.goto(URL, wait_until="load", timeout=60_000)
        if response is not None and response.status >= 400:
            raise RuntimeError(f"Navigation returned HTTP {response.status}: {URL}")

        page.pdf(
            path=str(OUTPUT),
            format="A4",
            print_background=True,
            margin={"top": "12mm", "right": "12mm", "bottom": "12mm", "left": "12mm"},
        )
        print(f"Saved {OUTPUT.resolve()}")
    finally:
        browser.close()

Use Chromium for PDF generation. A website may still produce different output depending on its fonts, print styles, JavaScript, and remote resources, so inspect the resulting file for the specific page you need.

3. Control layout, media, and rendering

PDF output depends on choices about CSS media, paper dimensions, margins, scale, colors, and page ranges. Decide these before tuning the page itself.

page.pdf() renders with print CSS by default. This is usually right for documents designed to print: sites can hide navigation, change colors, or reflow content with @media print. To render the page using screen CSS instead, emulate screen media before creating the PDF:

page.emulate_media(media="screen")
page.pdf(path="screen-layout.pdf", format="A4", print_background=True)

Choose one mode deliberately. Screen mode does not mean the PDF is a screenshot; it means the browser applies screen media rules while laying out the PDF.

Paper size, width, and height

Use format for a named size such as "A4", "Letter", "Legal", or "A3". The API also supports explicit width and height. Dimensions can use px, in, cm, or mm; unlabeled values are interpreted as pixels. If you set both format and width or height, format takes priority.

When a page defines a CSS @page size, set prefer_css_page_size=True to give that declaration priority over API dimensions. Otherwise, the content is scaled to fit the requested paper size.

Margins, background graphics, and color

Margins default to zero. Pass a dictionary of sides and unit-bearing values to control them. print_background defaults to False; set it to True to include CSS backgrounds and background images. Print rendering can adjust colors. If exact CSS colors matter, the page’s styles can use -webkit-print-color-adjust: exact; the PDF option alone does not force every site to preserve exact colors.

Scale, page ranges, headers, and footers

  • scale defaults to 1 and accepts values from 0.1 through 2. Smaller values fit more content but also shrink text.
  • page_ranges selects pages, for example "1-3, 5". An empty value means all pages. Page numbering applies to the generated document.
  • display_header_footer=True enables Chromium’s print header and footer. Supply header_template and footer_template as HTML strings. Template scripts are not evaluated and the page’s styles are not available inside the templates; style the template markup inline.
  • tagged and outline are documented options in current API documentation. Their presence does not by itself guarantee that the generated file has the accessibility or navigation quality your workflow requires; review the output in its intended viewer.

Wait for dynamic pages to become ready

Navigation completion and application readiness are not always the same thing. If content appears after a client-side request, wait for a meaningful page element before calling page.pdf(). Prefer a selector tied to the content you need over an arbitrary long sleep:

page.goto("https://example.com/report", wait_until="domcontentloaded", timeout=60_000)
page.locator("main article").wait_for(state="visible", timeout=30_000)
page.pdf(path="report.pdf", format="Letter", print_background=True)

For pages whose images load as you scroll, a full-page PDF may not include content that the page has not loaded. If the site lazy-loads important material, scroll through the relevant page before producing the PDF, then wait for its content to settle. Avoid assuming that networkidle is appropriate for every site; pages with ongoing network activity may never reach that condition.

4. Return PDF bytes instead of saving directly

Without a path, page.pdf() returns PDF bytes. This is useful when you want to send the result to another service, store it in object storage, or write it with your own file-handling code.

from pathlib import Path
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page()
        page.goto("https://example.com", wait_until="load", timeout=60_000)
        pdf_bytes = page.pdf(format="A4", print_background=True)
        Path("page.pdf").write_bytes(pdf_bytes)
    finally:
        browser.close()

For a service handling many pages, manage browser and context lifetimes explicitly and avoid launching a new browser for every URL when reuse is appropriate. Keep output paths unique if jobs may run concurrently.

5. Async version

Use Playwright’s asynchronous API when the rest of your application is already async. Await each browser operation and close the browser even if navigation or PDF generation raises an exception.

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            page = await browser.new_page()
            response = await page.goto("https://example.com", wait_until="load", timeout=60_000)
            if response is not None and response.status >= 400:
                raise RuntimeError(f"Navigation returned HTTP {response.status}")
            await page.pdf(
                path="page.pdf",
                format="A4",
                print_background=True,
                margin={"top": "12mm", "bottom": "12mm", "left": "12mm", "right": "12mm"},
            )
        finally:
            await browser.close()

asyncio.run(main())

6. Download a PDF attachment instead of printing the page

page.pdf() creates a PDF from the rendered webpage. It is not the right method when a link or button downloads an existing PDF attachment. For an attachment, wait for the download event and save the file before closing the browser context; Playwright deletes context downloads when that context closes. See the official download guide.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page()
        page.goto("https://example.com/files", wait_until="load")
        with page.expect_download(timeout=30_000) as download_info:
            page.get_by_role("link", name="Download PDF").click()
        download = download_info.value
        download.save_as("attachment.pdf")
    finally:
        browser.close()

Replace the link locator with one that matches the actual page. If clicking triggers a navigation to a PDF viewer rather than a download, inspect that behavior and handle the resulting response or page separately.

7. Or skip the browser setup

If you need a PDF from a URL without installing Chromium or managing a browser process, ScreenshotNeo provides a one-request screenshot and PDF API. Its capture handles cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; and its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. PDF paper size, orientation, backgrounds, margins, and page ranges are configurable. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o stripe.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("stripe.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
const { writeFile } = await import('node:fs/promises');
await writeFile('stripe.pdf', Buffer.from(await res.arrayBuffer()));

Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card.

8. Troubleshooting

Symptom Likely cause What to change
Executable doesn't exist or Chromium launch fails The Playwright package is installed but its browser binary is missing. Run python -m playwright install chromium in the same environment.
PDF has different layout from the browser PDF uses print media by default, or print CSS changes the page. Keep print mode if intended; otherwise call page.emulate_media(media="screen") before page.pdf().
Background colors or images are missing print_background defaults to false. Set print_background=True. For precise colors, check the page’s print CSS and -webkit-print-color-adjust.
Content is missing or a spinner appears The page navigated, but client-side content was not ready. Wait for a locator that signals the required content is visible. Set an appropriate timeout and check for page-side errors.
PDF is blank or only the first part appears The target may require authentication, lazy-loads below the fold, or has different print styles. Use the required session state, scroll to trigger relevant lazy content, and inspect the site’s print CSS. Verify the URL and final page state.
Wrong paper size or content unexpectedly shrinks format overrides width and height, or a CSS @page rule is interacting with API dimensions. Choose one sizing strategy. Use prefer_css_page_size=True when the site’s CSS page size should win.
Small text or clipped content Scale, margins, or paper dimensions do not suit the page. Adjust paper size and margins first; change scale within 0.1–2 only if needed. Check the result page by page.
Download wait times out The action did not produce an attachment, the locator did not match, or the site opened a viewer instead. Confirm the click target and inspect whether the response is a download or navigation. Save the download before closing its context.
Script hangs or times out The site is slow, keeps network connections open, or blocks automation. Use a realistic navigation timeout, wait for a specific selector when possible, and capture diagnostic errors. A browser-rendered workflow cannot guarantee access to every site.

9. Performance, reliability, and cost

Playwright’s software is installed as a Python package, and Chromium browser binaries must also be installed. Browser launch, page navigation, fonts, images, scripts, and page complexity all contribute to the time and resources required; no single render time applies to all sites. For repeated work, reuse a browser process where suitable, isolate jobs with contexts, set timeouts, and close resources reliably. Reuse must not leak cookies or other session state between users.

For more predictable output, choose the media mode, viewport, paper size, readiness condition, and margins explicitly. Save files to unique paths, validate the HTTP response where available, and inspect generated PDFs for important workflows. A successful PDF call does not establish that every dynamic widget, external font, or image rendered as intended.

The Playwright library and browser setup are software dependencies rather than per-document API charges. Operational costs depend on the machine and infrastructure used to run browser jobs. If you prefer a hosted capture API, ScreenshotNeo has a free allowance and paid plans; review its documentation for current request options and behavior.

10. FAQ

Can Playwright save the page as PDF without writing a file first?

Yes. Call page.pdf() without path; it returns PDF bytes that Python can store or pass to another service.

Does this work with Firefox or WebKit?

The documented Playwright PDF workflow is Chromium based. Launch p.chromium for this task.

How do I save only selected pages?

Pass a range such as page_ranges="1-3, 5" to page.pdf().

Why does a downloaded PDF differ from page.pdf()?

A downloaded attachment is an existing file supplied by the site. page.pdf() generates a new PDF from the browser’s rendered page.

Can I guarantee identical PDFs across machines?

Not solely by setting PDF options. Browser version, installed fonts, operating system, site content, and external resources can affect rendering. Pin and manage your environment, then review output for your use case.

References