ScreenshotNeo

BlogHow-to

Take a Webpage Screenshot in Python with Playwright and Save It as PDF

Use Playwright for Python to capture a webpage as a screenshot or PDF. Learn full-page capture, PDF styling, setup, options, and fixes for common errors.

By the ScreenshotNeo team4 October 20268 min read

Use Playwright’s Python API to open a page, save a screenshot with page.screenshot(), and create a PDF with page.pdf(). Set full_page=True for the entire scrollable page. PDF output uses print CSS by default; call page.emulate_media(media="screen") first if you want screen styling.

1. Install Playwright and its browser

Install the Python package, then install Chromium. The browser installation is a separate step because Playwright needs a browser binary to render the page.

python -m pip install playwright
python -m playwright install chromium

Save this as capture_page.py. It uses Playwright’s synchronous API and writes both a full-page PNG and a PDF.

from pathlib import Path
from playwright.sync_api import sync_playwright

URL = "https://example.com"

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(URL, wait_until="load", timeout=60_000)

    page.screenshot(path="screenshot.png", full_page=True)
    page.pdf(path="page.pdf", print_background=True)

    browser.close()

Run it with python capture_page.py. Replace URL with the page you want. The PNG is a raster image of the page; the PDF is paginated print output, so the two files need not look identical.

2. Choose viewport or full-page capture

By default, page.screenshot() captures the currently visible viewport. To capture the full scrollable document, pass full_page=True:

page.screenshot(path="viewport.png")
page.screenshot(path="full-page.png", full_page=True)

A full-page screenshot is one tall image, not a set of PDF pages. Very long pages can produce large images or exceed browser and memory limits. For a PDF, use page.pdf(), which lays out content across pages.

You can also capture an element or receive image bytes instead of writing a file, which is useful if you need to post-process the image or send it elsewhere. See the Playwright Python screenshot guide for element and buffer examples.

3. Control PDF styling and layout

page.pdf() renders with print CSS media by default. Sites may hide navigation, change colors, or rearrange content for printing. To use the page’s screen CSS instead, emulate screen media before creating the PDF:

page.emulate_media(media="screen")
page.pdf(path="screen-style.pdf", print_background=True)

PDF backgrounds are omitted by default. Set print_background=True when background colors or graphics matter. Print color adjustment may also change colors; CSS can request color fidelity with -webkit-print-color-adjust: exact.

Other documented PDF controls include paper format or explicit width and height, margins, scale, page ranges, and whether CSS page size takes precedence. For example:

page.pdf(
    path="letter.pdf",
    format="Letter",
    margin={"top": "0.5in", "right": "0.5in", "bottom": "0.5in", "left": "0.5in"},
    print_background=True,
    prefer_css_page_size=True,
)

Choose the paper and margins to suit the document. If the output has unexpected page breaks or clipping, inspect the page’s print styles and try a different paper size or scale. The available settings and their accepted values are documented in the Playwright Page API.

4. Wait for the page to be ready

Navigation completion does not always mean that client-rendered content, images, or fonts are ready. Choose a wait condition that matches the page:

# Wait for the document load event (used in the main example)
page.goto(URL, wait_until="load")

# For a known content element, wait for it explicitly
page.goto(URL, wait_until="domcontentloaded")
page.locator("main article").wait_for(state="visible", timeout=15_000)

# For a page with a short, known rendering delay
page.goto(URL, wait_until="load")
page.wait_for_timeout(1_000)

Use a selector wait when the page has a reliable element that indicates its content is ready. A fixed delay is simple but can waste time or still be too short. Avoid assuming that every page becomes quiet on the network; analytics or streaming requests may continue indefinitely.

5. Use the asynchronous Python API

For async applications, use async_playwright and await navigation and capture calls:

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page()
        await page.goto("https://example.com", wait_until="load", timeout=60_000)
        await page.screenshot(path="screenshot.png", full_page=True)
        await page.pdf(path="page.pdf", print_background=True)
        await browser.close()

asyncio.run(main())

Use the sync form for a straightforward script and the async form when integrating with an asyncio application. Avoid calling asyncio.run() from inside an already-running event loop; in that setting, await the coroutine from the existing loop.

6. Complete runnable alternatives in cURL, Python, and Node.js

Playwright itself is a browser automation library, so its screenshot and PDF methods are called from Python or another supported language rather than cURL. For reference, this is the complete Python workflow from installation through capture:

python -m pip install playwright
python -m playwright install chromium
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto("https://example.com", wait_until="load", timeout=60_000)
    page.screenshot(path="page.png", full_page=True)
    page.pdf(path="page.pdf", print_background=True)
    browser.close()

There is no cURL command that invokes the local Playwright browser API directly. If you need an HTTP request instead of maintaining browser setup, ScreenshotNeo provides a screenshot API; its example calls are below. For the equivalent Playwright workflow in Node.js, install the package and browser, then run:

npm install playwright
npx playwright install chromium
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'load', timeout: 60000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
  await page.pdf({ path: 'page.pdf', printBackground: true });
  await browser.close();
})();

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. A single GET request returns an image or PDF, without installing and managing a browser in your script. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o page.pdf

The API can remove cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month with no card.

8. Troubleshooting

Symptom Likely cause Fix
Browser executable not found The Python package is installed, but its browser binary is not. Run python -m playwright install chromium in the same environment that runs the script.
Navigation timeout The site is slow, unreachable, or keeps background requests open. Set an appropriate navigation timeout, use wait_until="load" or "domcontentloaded", then wait for a specific content selector. Check the URL and network access.
Screenshot is blank or content is missing The page may render content after navigation or require a selector to become visible. Wait for the relevant locator before capture. Check that the target URL loads in the browser and that the page is not showing a bot check.
PDF looks different from the browser PDF uses print media by default, and print styles can alter layout. Call page.emulate_media(media="screen") before page.pdf() if screen styling is desired.
PDF is missing background colors or images Background graphics are omitted by default. Set print_background=True.
PDF content is clipped or split awkwardly Paper size, margins, scaling, and CSS page rules affect pagination. Adjust format, margins, or scale; inspect print CSS and consider prefer_css_page_size=True.
Full-page image is huge or capture fails A very long document creates a tall raster image and can use substantial memory. Use a PDF for paginated output, capture only the needed element, or capture a viewport instead.
Script hangs or browser stays open The browser is not closed after an exception or the process is waiting on page work. Use a try/finally around browser work in long-running scripts, close the browser, and use bounded navigation and selector timeouts.
asyncio.run() reports an active event loop The code is running inside an environment that already owns an event loop. Await the async function from that loop instead of calling asyncio.run().

9. Performance, reliability, and cost

Playwright runs a real browser, so the script needs the browser binary and enough memory for the page and output. A full-page PNG can be much larger than a viewport screenshot, particularly for tall pages. Prefer PDF when the intended deliverable is paginated; restrict capture to an element or viewport when that meets the requirement.

For reliability, use explicit timeouts, wait for a meaningful page condition, and close the browser even if navigation or capture raises an exception. If a page is dynamic, avoid relying on a fixed delay alone. The dossier provides no benchmark or fixed runtime expectation; actual time and memory depend on the target page and environment.

Playwright is software you run in your own environment; account for the compute and maintenance involved in installing and running Chromium. ScreenshotNeo offers a free allowance of 1,000 shots per month, then plans from $5 for 3,000; higher listed plans are Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. Every feature is on every plan. Compare that managed request flow with the browser control and local processing of Playwright based on your deployment needs.

10. FAQ

Does page.screenshot() create a PDF?

No. It writes an image. Use page.pdf() for a paginated PDF.

Can I save a full page as one image?

Yes. Pass full_page=True to page.screenshot().

Why does the PDF not match the visible webpage?

PDF generation uses print CSS by default. Emulate screen media before generating it when you want screen styling.

Can I make the PDF include background graphics?

Yes. Pass print_background=True to page.pdf().

Can I use this workflow with async Python?

Yes. Playwright provides an async Python API; use async_playwright and await each browser and page operation.

Sources