ScreenshotNeo

BlogHow-to

Pyppeteer Tutorial: Automate Screenshots with Headless Chrome

Learn to capture website screenshots with Pyppeteer and headless Chrome, with runnable Python examples, troubleshooting, and deployment guidance.

By the ScreenshotNeo team4 October 20269 min read

Use Pyppeteer to launch headless Chromium, open a page, navigate to a URL, save a screenshot, and close the browser. Pyppeteer is an unofficial Python port of Puppeteer. Its repository currently describes the project as unmaintained and suggests Playwright Python as an alternative, so this tutorial is aimed at readers who specifically need Pyppeteer or are maintaining an existing workflow.

What Pyppeteer does and its maintenance status

Pyppeteer controls Chromium through an asynchronous Python API. The basic screenshot sequence is launch → page → navigation → screenshot → close. The project README documents Python 3.8 or later as its baseline, but that is not a guarantee that every current Python and Chromium combination will work. The repository’s unmaintained status also means you should assess compatibility in your own environment before depending on it for a new production system.

The repository names Playwright Python as an alternative. Playwright’s official Python documentation describes launching Chromium, Firefox, or WebKit and taking screenshots. Choose after considering maintenance status, runtime and browser compatibility in your target environment, browser installation, the API changes needed for an existing script, and deployment constraints. The cited documentation does not establish a benchmark or a feature-by-feature reliability comparison.

Install Pyppeteer and prepare Chromium

Install the package in your active Python environment:

python -m pip install pyppeteer

Pyppeteer may download Chromium on first use if it cannot find a local browser. To perform that provisioning step ahead of running the script, the project documents:

pyppeteer-install

Run the command using the same environment where Pyppeteer is installed. In restricted build or deployment environments, verify that the browser download is permitted and that the resulting browser files are available to the process that runs your script.

Capture a page with a minimal runnable script

Save this as screenshot.py. It accepts a URL and output path, waits for navigation, captures the page, and closes the browser even if navigation or capture raises an error.

import argparse
import asyncio
from pathlib import Path
from urllib.parse import urlparse

from pyppeteer import launch


def validate_url(value: str) -> str:
    parsed = urlparse(value)
    if parsed.scheme not in {"http", "https"} or not parsed.netloc:
        raise argparse.ArgumentTypeError("URL must be an absolute http:// or https:// URL")
    return value


async def capture(url: str, output: Path) -> None:
    browser = await launch(headless=True)
    try:
        page = await browser.newPage()
        response = await page.goto(url, {"waitUntil": "networkidle2", "timeout": 30000})
        if response is not None and response.status >= 400:
            raise RuntimeError(f"Navigation returned HTTP {response.status} for {url}")
        await page.screenshot({"path": str(output), "fullPage": True})
    finally:
        await browser.close()


def main() -> None:
    parser = argparse.ArgumentParser(description="Capture a website screenshot with Pyppeteer")
    parser.add_argument("url", type=validate_url)
    parser.add_argument("output", nargs="?", default="screenshot.png")
    args = parser.parse_args()
    output = Path(args.output)
    output.parent.mkdir(parents=True, exist_ok=True)
    asyncio.run(capture(args.url, output))
    print(f"Saved {output}")


if __name__ == "__main__":
    main()

Run it with:

python screenshot.py https://example.com output/example.png

The Pyppeteer README’s example runs an async main with asyncio.get_event_loop().run_until_complete(main()). This tutorial uses asyncio.run as a straightforward script entry point; embedded applications and notebooks that already manage an event loop need an entry point appropriate to that context.

The navigation call uses networkidle2, which waits for network activity to become quiet. Some sites keep connections open or load analytics continuously, so a network-idle condition may take too long or never occur. In those cases, use domcontentloaded or load, then wait for a specific element or a short deliberate delay before capture.

Choose screenshot size, page state, and capture target

Viewport screenshot or full page

By default, a screenshot covers the visible viewport. Set fullPage to True to capture the full document height:

await page.screenshot({"path": "full-page.png", "fullPage": True})

Long pages can produce large images and consume more memory. A full-page screenshot is also a single tall image; it does not paginate content like a PDF. Lazy-loaded images may not appear if they have not been brought into view before capture. For such pages, scroll through the document in increments and allow content to load before taking the screenshot.

Set a viewport

Set the viewport before navigation when the layout depends on screen size:

await page.setViewport({"width": 1440, "height": 900, "deviceScaleFactor": 1})

A larger viewport can change responsive breakpoints and produce a different layout. deviceScaleFactor controls pixel density; increasing it can improve sharpness but also increases output dimensions and file size.

Capture one element

Wait for a selector, find its element, then capture that element’s bounds:

await page.waitForSelector("main article", {"visible": True, "timeout": 10000})
element = await page.querySelector("main article")
if element is None:
    raise RuntimeError("Article element was not found")
await element.screenshot({"path": "article.png"})

Element screenshots are useful for a chart, card, or article region. If the selector matches multiple elements, querySelector returns the first one. Use a more specific selector when the page has repeated components.

Wait for a meaningful state

Navigation completion does not always mean a single-page application has rendered its final content. Wait for a page-specific element when possible:

await page.goto(url, {"waitUntil": "domcontentloaded", "timeout": 30000})
await page.waitForSelector("[data-report-ready='true']", {"timeout": 15000})
await page.screenshot({"path": "ready.png", "fullPage": True})

Use selectors that represent the state you need, such as the main content container. Fixed sleeps are less reliable because fast pages wait unnecessarily while slow pages may still be incomplete.

Pass options to Chromium and the page

Pyppeteer exposes browser and page configuration through Python dictionaries. These examples show common controls; the exact Chromium arguments needed depend on the host environment.

Headless mode and executable path

browser = await launch(
    headless=True,
    executablePath="/path/to/chromium",  # Optional: use a local browser
    args=["--no-sandbox"],              # Only if required by the environment
)

Omit executablePath to let Pyppeteer locate or provision its browser. A custom executable can help when deployments supply their own Chromium, but the unmaintained project status makes compatibility with arbitrary current browser builds something to validate rather than assume. The --no-sandbox argument weakens browser isolation; use it only when the environment requires it and account for that tradeoff.

page.goto accepts options such as waitUntil and timeout. Common wait conditions include:

  • domcontentloaded: the document has been parsed; later assets may still be loading.
  • load: the page’s load event has fired.
  • networkidle0 or networkidle2: wait for a quiet network period using the corresponding connection threshold.

Choose the least restrictive condition that gives the page enough time to render, and follow it with a selector wait when the content has a clear ready marker.

Other useful page settings

await page.setUserAgent("ExampleCaptureBot/1.0")
await page.setExtraHTTPHeaders({"Accept-Language": "en-US,en;q=0.9"})
await page.setJavaScriptEnabled(True)
await page.setDefaultNavigationTimeout(30000)

Only set a custom user agent or headers when your use case requires them. A site may serve different content based on these values. Disabling JavaScript can prevent client-rendered content from appearing.

Run captures reliably and efficiently

  • Always close the browser. Use try/finally so browser processes do not remain after an exception.
  • Reuse a browser for batches. Launching Chromium for every URL adds startup and provisioning overhead. For a batch, keep one browser open and create a separate page per capture, then close each page and the browser.
  • Limit concurrency. Each open page consumes memory and browser resources. Start with a small number of concurrent pages, then adjust based on the machine and target sites.
  • Use explicit timeouts. Set navigation and selector timeouts to bound stalled work and report which step failed.
  • Keep output manageable. Use viewport captures when a full document is unnecessary. Choose dimensions and device scale based on the downstream use.
  • Provision the browser predictably. Install Chromium during image or environment setup when possible, and check that the runtime user can read and execute the browser files.
  • Record failure context. Log the target URL, wait condition, exception, and elapsed time. Avoid logging credentials or sensitive query parameters.

Pyppeteer and Chromium are software dependencies rather than per-image API charges, so cost depends on the compute, storage, and operations used to run them. This research does not provide performance benchmarks or comparative reliability measurements. Measure the workflow on your own target pages and deployment environment.

Troubleshoot common failures

Symptom Likely cause What to do
Chromium download or launch fails First-run provisioning is blocked, the browser is missing, or the runtime cannot access it. Run pyppeteer-install in the intended environment, confirm browser file permissions, and inspect the underlying launch error. If using a local executable, verify its path and compatibility.
Navigation Timeout Exceeded The site is slow, holds network connections open, or the selected wait condition is too strict. Set an explicit timeout; try domcontentloaded followed by a wait for the content selector you need.
Screenshot is blank or missing content The page is client-rendered, a selector was not ready, or content is below the fold and lazy loaded. Wait for a meaningful selector. For lazy content, scroll through the page and give newly requested assets time to load before capture.
Element screenshot fails or captures the wrong region The selector is absent, hidden, ambiguous, or matches a different repeated element. Wait for the selector with visibility enabled, check that it exists, and narrow the selector.
Browser exits unexpectedly in a container Required system libraries or permissions may be absent, or the browser sandbox conflicts with container settings. Check Chromium’s launch output and the container’s installed dependencies. Use --no-sandbox only if required by the environment and with its isolation tradeoff understood.
Script reports an event-loop error The script is running inside an environment that already owns an asyncio loop. Use that environment’s async execution mechanism instead of calling asyncio.run from inside an active loop.
Output path does not exist The parent directory has not been created. Create it before capture; the sample script does this with mkdir(parents=True, exist_ok=True).
Current Chromium behaves differently from expected Pyppeteer’s repository is unmaintained; Puppeteer’s current browser matrix does not establish Pyppeteer compatibility. Validate the exact Python, Pyppeteer, Chromium, and operating-system combination. Consider evaluating Playwright Python if maintaining Pyppeteer is no longer practical.

Or skip the browser setup

If you need a screenshot without provisioning and maintaining Chromium, ScreenshotNeo provides a website screenshot API and MCP server. Its [documentation](https://screenshotneo.com/docs/) describes the API. A single GET request can return a screenshot or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace YOUR_API_KEY with an API key from your account. The Node.js snippet uses Bun’s file-writing helper; in Node.js, save the returned bytes with await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))) inside an async function.

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Frequently asked questions

Is Pyppeteer the same package as Puppeteer?

No. Pyppeteer is an unofficial Python port of Puppeteer. Puppeteer’s JavaScript examples and browser support details should not be treated automatically as Pyppeteer compatibility guarantees.

Can Pyppeteer capture an element instead of a whole page?

Yes. Wait for the target selector, query the element, and call its screenshot method as shown above.

Should I start a new project with Pyppeteer?

The repository describes Pyppeteer as unmaintained and points to Playwright Python as an alternative. If you are starting fresh, evaluate that maintenance context and verify the browser/runtime support your project needs.

Does a current Puppeteer browser matrix apply to Pyppeteer?

No. The Puppeteer browser support documentation describes Puppeteer releases and their browsers. It does not establish which current Chromium versions work with Pyppeteer.

Sources