ScreenshotNeo

BlogHTML to image & PDF

How to Fix Pyppeteer Generating Blank PDFs Instead of the Full Document

Find why Pyppeteer PDFs are blank, then fix navigation, resources, print CSS, pagination, and readiness with runnable diagnostics.

By the ScreenshotNeo team1 October 20268 min read

A blank Pyppeteer PDF usually means the problem occurred before PDF encoding: the page had no usable content at print time, its resources failed to load, print CSS hid the content, or pagination created an apparently empty page. Start by checking the rendered DOM and failed requests, then verify print media and layout.

Pyppeteer’s page.pdf() uses print CSS media by default. A page that looks correct on screen can therefore print as an empty document. Navigation completion also does not guarantee that an application has finished rendering its data.

1. Identify the kind of blank PDF

Open the file and classify the symptom before changing options:

  • One completely blank page: navigation, application rendering, resources, or print CSS is likely involved.
  • Some content is missing: inspect failed assets, authentication, fonts, images, and @media print rules.
  • Correct content plus an extra blank page: inspect page geometry, margins, overflow, break rules, and document height.

Run diagnostics in the same process that creates the PDF. Log the navigation response, final URL, title, body text, dimensions, console messages, page errors, and failed requests.

2. A complete diagnostic Pyppeteer script

Install Pyppeteer with pip install pyppeteer. The first launch downloads a compatible Chromium revision unless you provide an executable path.

import asyncio
from pathlib import Path
from pyppeteer import launch

URL = "https://example.com"
OUTPUT = Path("debug.pdf")

async def main():
    browser = await launch(
        headless=True,
        args=["--no-sandbox", "--disable-setuid-sandbox"],
    )
    page = await browser.newPage()

    page.on("console", lambda message: print("CONSOLE:", message.type, message.text))
    page.on("pageerror", lambda error: print("PAGE ERROR:", error))
    page.on(
        "requestfailed",
        lambda request: print(
            "REQUEST FAILED:", request.url, request.failure
        ),
    )
    page.on(
        "response",
        lambda response: print(
            "HTTP:", response.status, response.url
        ) if response.status >= 400 else None,
    )

    try:
        response = await page.goto(
            URL,
            {
                "waitUntil": ["domcontentloaded", "networkidle2"],
                "timeout": 60000,
            },
        )
        print("NAVIGATION STATUS:", response.status if response else None)
        print("FINAL URL:", page.url)
        print("TITLE:", await page.title())

        # Replace this selector with your application’s real ready signal.
        try:
            await page.waitForSelector("body", {"visible": True, "timeout": 10000})
        except Exception as error:
            print("READY SELECTOR WARNING:", error)

        body_text = await page.evaluate(
            """() => document.body ? document.body.innerText.slice(0, 1000) : ''"""
        )
        dimensions = await page.evaluate(
            """() => ({
                readyState: document.readyState,
                bodyTextLength: document.body ? document.body.innerText.length : 0,
                scrollWidth: document.documentElement.scrollWidth,
                scrollHeight: document.documentElement.scrollHeight,
                bodyHeight: document.body ? document.body.getBoundingClientRect().height : 0
            })"""
        )
        print("BODY SAMPLE:", repr(body_text))
        print("DIMENSIONS:", dimensions)

        await page.pdf(
            {
                "path": str(OUTPUT),
                "format": "A4",
                "printBackground": True,
                "margin": {"top": "20mm", "right": "15mm", "bottom": "20mm", "left": "15mm"},
            }
        )
        print("WROTE:", OUTPUT)
    finally:
        await browser.close()

asyncio.get_event_loop().run_until_complete(main())

Pyppeteer documents navigation failures for invalid URLs, SSL errors, timeouts, and failed main resources. A None response can also occur for about:blank or same-URL hash navigation. See the Pyppeteer API reference.

3. Wait for application content, not only navigation

load, domcontentloaded, networkidle0, and networkidle2 describe browser navigation conditions. Network idle means that requests were quiet for a period; it does not prove that your framework received data or finished rendering.

Wait for a selector

await page.goto(
    "https://app.example.com/report",
    {"waitUntil": "domcontentloaded", "timeout": 60000},
)
await page.waitForSelector("#report-ready", {"visible": True, "timeout": 30000})
await page.pdf({"path": "report.pdf", "format": "A4"})

Wait for an application flag

await page.waitForFunction(
    """() => window.reportReady === true""",
    {"timeout": 30000},
)

Use a fixed sleep only when the site has no reliable readiness signal. If the body text is empty, investigate redirects, authentication, blocked scripts, failed API requests, and empty API responses before changing PDF options.

4. Check resources and document setup

A successful top-level navigation does not mean that stylesheets, images, fonts, scripts, or API calls succeeded. Missing resources can leave an otherwise valid HTML document visually empty.

  • For setContent() and data: documents, verify that stylesheet, image, font, and script URLs are absolute or use a reachable base URL.
  • For local documents, confirm file permissions and that every asset path resolves from the file URL’s directory.
  • For remote assets, inspect HTTP status, certificates, redirects, authentication, and cross-origin requirements.
  • Listen for requestfailed, console errors, and page errors as shown in the diagnostic script.
  • Check that an API response contains the expected records instead of an empty state.

Use a base URL with local HTML

from pathlib import Path

html = Path("report.html").read_text(encoding="utf-8")
await page.setContent(html)
# Prefer absolute asset URLs in report.html. For a local page, navigate to its file URL instead:
await page.goto(Path("report.html").resolve().as_uri(), {"waitUntil": "networkidle0"})

An issue report describes blank output alongside missing files and unapplied CSS. Treat that report as evidence of the resource-problem class, not as proof that one path correction fixes every document. See Puppeteer issue #6417.

5. Inspect print CSS and media emulation

The PDF API renders with print media by default. Rules inside @media print can hide content, change colors, reposition elements, or apply page breaks.

Capture the intended print layout

await page.emulateMedia("print")
await page.pdf({
    "path": "print-layout.pdf",
    "format": "A4",
    "printBackground": True,
})

Capture the screen layout intentionally

await page.emulateMedia("screen")
await page.pdf({
    "path": "screen-layout.pdf",
    "format": "A4",
    "printBackground": True,
})

Use screen emulation only when the desired output is the screen design. A durable fix is usually to correct the print stylesheet.

  • Search for display: none, visibility: hidden, zero opacity, and white text on a white background.
  • Check absolutely positioned elements whose containing block changes under print.
  • Inspect break-before, break-after, break-inside, and older page-break-* rules.
  • Ensure essential information is not supplied only through a CSS background image.
  • Compare computed styles under print and screen media with browser DevTools or page.evaluate().

6. Verify PDF options

Layout options affect what is visible and how pages are split, but they cannot restore DOM content that is absent or hidden.

Option What to check
format Use a known paper size such as A4 or Letter.
width/height Do not combine conflicting dimensions without understanding which setting wins.
margin Large margins can move a small element to another page.
scale Extreme values can make content appear clipped or unexpectedly small.
printBackground Defaults to false; enable it when backgrounds carry required visual information.
pageRanges An empty value means all pages. A wrong range can select no useful page.
landscape Use it when wide tables or positioned content otherwise overflow.

7. Diagnose an extra blank trailing page

If the document has the expected pages followed by an empty page, inspect geometry rather than navigation:

  1. Remove custom margins and page-break rules temporarily.
  2. Inspect the total scroll height and the height of the last element.
  3. Check horizontal overflow that may increase the printable area.
  4. Test html, body { height: 100%; } by removing it in a minimal reproduction. One older report identified that rule as a suspected cause in its specific reproduction; it is not a universal diagnosis.
  5. Compare Chromium and Pyppeteer versions after reducing the document to the smallest failing HTML.

See issue #589 and the version-specific reproduction in issue #6704.

8. Minimal known-good example

import asyncio
from pyppeteer import launch

async def main():
    browser = await launch()
    page = await browser.newPage()
    await page.goto(
        "https://example.com",
        {"waitUntil": "networkidle2", "timeout": 60000},
    )
    await page.emulateMedia("print")
    await page.pdf({
        "path": "example.pdf",
        "format": "A4",
        "printBackground": True,
        "margin": {"top": "16mm", "right": "16mm", "bottom": "16mm", "left": "16mm"},
    })
    await browser.close()

asyncio.get_event_loop().run_until_complete(main())

If this works but your document does not, add your application readiness wait, resources, styles, and PDF options back one at a time.

9. Version and environment checklist

  • Record the Pyppeteer version, Chromium revision or executable path, operating system, launch arguments, URL, and PDF options.
  • Reproduce with a minimal HTML file and one stylesheet.
  • Check whether the same URL prints correctly in a manually opened Chromium instance.
  • Compare screen and print output separately.
  • Change one variable at a time so a successful run identifies the actual fix.

10. Or skip the browser setup

ScreenshotNeo provides a website screenshot and PDF API when maintaining Chromium code is unnecessary. Its capture pipeline accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture, lazy-image loading, CSS selector capture, custom CSS and JavaScript, waits for selectors or network idle, cookies and headers, PDF controls, caching, signed links, async jobs, bulk capture, and a usage API. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and get 1,000 screenshots each month with no card.

11. Performance, reliability, and cost notes

  • Performance: Prefer an application-ready selector over a long fixed delay. Reuse a browser process when generating many PDFs, while creating isolated pages for separate jobs.
  • Reliability: Set navigation and readiness timeouts, capture diagnostics, and close pages and browsers in finally blocks. Do not treat network idle as proof of rendered data.
  • Resource cost: Fonts, images, JavaScript, and repeated browser launches increase work. Block unnecessary resources only when they cannot affect the document.
  • Repeatability: Pin Pyppeteer and Chromium versions for production and keep a minimal reproduction for regressions.
  • API alternative: ScreenshotNeo’s cache can reduce repeated captures, and failed loads, blank pages, bot checks, timeouts, and cache hits are not billed.

12. Troubleshooting table

Symptom Likely cause Fix
Navigation throws a timeout Slow page, blocked request, or unsuitable wait condition Inspect failed requests, raise the timeout when justified, and wait for a real ready selector.
Body text is empty Redirect, authentication, failed script, or empty API response Log the final URL, response status, console errors, and API calls.
Screen works, PDF is blank Print CSS hides or relocates content Inspect @media print; emulate screen only when screen output is intended.
Images or styles are absent Unreachable relative paths, failed requests, certificates, or permissions Use reachable absolute paths and inspect requestfailed and HTTP errors.
Background information is missing printBackground is false Set printBackground: True; this does not restore hidden DOM content.
Only the last page is blank Height, overflow, margins, or page-break interaction Reduce to minimal HTML and test geometry rules one at a time.
Fonts change the page count Font loading or fallback changes text metrics Wait for the page’s font and content readiness before calling pdf().

FAQ

Does networkidle0 guarantee a complete PDF?

No. It only describes network quiet; application data may still be absent or rendering may depend on a separate readiness signal.

Why does printBackground not fix a blank PDF?

It prints CSS backgrounds. It cannot restore text hidden by print CSS, missing resources, or an empty application state.

Should I always emulate screen media?

No. Use print media for a print stylesheet. Emulate screen when the requirement is specifically to reproduce the on-screen layout.

Can a valid HTTP 200 still produce a blank PDF?

Yes. Dependent assets, scripts, fonts, or API calls can fail after the main document responds.

What is the fastest isolation method?

Print a minimal static page, then add navigation waits, application data, resources, print CSS, and layout rules one layer at a time.