ScreenshotNeo

BlogHow-to

How to Fix Blank HTML After a Page Loads in Pyppeteer

Diagnose blank Pyppeteer HTML by checking navigation, DOM timing, selectors, JavaScript errors, and failed requests with runnable fixes.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: a completed page.goto() only means that the selected navigation milestone finished. It does not prove that a client-rendered application has inserted the content you need. Log the navigation response and current URL, inspect page.content() and the body text, then wait for a meaningful application selector before extracting HTML. If the DOM is still empty, investigate the URL, HTTP response, JavaScript errors, failed requests, and your evaluation syntax.

1. Run a diagnostic capture first

Use this small script against the affected URL. It records the evidence needed to distinguish navigation failure, timing problems, and rendering errors.

import asyncio
import pyppeteer

TARGET = "https://example.com"

async def main():
    browser = await pyppeteer.launch(headless=True)
    page = await browser.newPage()

    page.on("console", lambda msg: print("CONSOLE:", msg.type, msg.text))
    page.on("pageerror", lambda exc: print("PAGE ERROR:", exc))
    page.on(
        "requestfailed",
        lambda req: print("REQUEST FAILED:", req.url, req.failure),
    )

    try:
        response = await page.goto(
            TARGET,
            {"waitUntil": "load", "timeout": 30000},
        )
        print("current URL:", page.url)
        print("response object:", response)
        if response is not None:
            print("response URL:", response.url)
            print("status:", response.status)

        print("title:", await page.title())
        html = await page.content()
        print("HTML length:", len(html))
        print("HTML preview:", html[:1000])

        body_text = await page.evaluate(
            "document.body ? document.body.innerText : ''",
            force_expr=True,
        )
        print("body text preview:", body_text[:1000])
    except Exception as exc:
        print("NAVIGATION ERROR:", repr(exc))
    finally:
        await browser.close()

asyncio.run(main())

Pyppeteer normally returns the main-resource response from goto(). A None response is expected for about:blank and for a same-URL navigation that changes only the hash. The API can also raise for an invalid URL, SSL failure, timeout, or main-resource failure. See the Pyppeteer API reference.

2. Confirm that you navigated to the intended document

A surprisingly common cause of “blank HTML” is an incomplete or unintended URL. Pass a complete URL with https:// or http://, and print page.url after navigation.

response = await page.goto(
    "https://example.com/products",
    {"waitUntil": "domcontentloaded", "timeout": 30000},
)
print("landed on:", page.url)
print("status:", None if response is None else response.status)

Check redirects, authentication pages, consent interstitials, and bot checks. A successful transport request can still land on an error or challenge document. Preserve the final URL and status in your diagnostic output.

3. Separate an empty document from extraction timing

page.content() serializes the current DOM, including the doctype. It tells you what exists at the exact moment you call it; it does not wait for application data. Compare the serialized HTML, title, and body text:

html = await page.content()
title = await page.title()
text = await page.evaluate(
    "document.body ? document.body.innerText : ''",
    force_expr=True,
)
print({"title": title, "html_chars": len(html), "text": text[:500]})

If you see an application shell such as a root element but not the expected records, navigation worked and rendering is incomplete. If the expected element exists but has no text, inspect its descendants and computed styles. If the serialized document itself is unexpectedly small, revisit the URL, response, redirects, and page errors.

4. Wait for the page’s actual readiness condition

Pyppeteer’s goto() default is load. Its documented alternatives are:

Option What it establishes When to use it Limitation
domcontentloaded Initial HTML has been parsed Fast extraction from server-rendered markup Scripts, images, and data may still be pending
load Load event fired Default when page resources should be loaded Single-page apps may render later
networkidle0 No more than zero connections for 500 ms Pages that become truly quiet Analytics, sockets, or polling can prevent completion
networkidle2 No more than two connections for 500 ms Pages with a small amount of background traffic Network quiet does not prove your content exists

For client-rendered pages, wait for a selector that represents the content you actually need. The selector wait resolves when a matching element appears and raises on timeout.

await page.goto(
    "https://example.com/search?q=pyppeteer",
    {"waitUntil": "domcontentloaded", "timeout": 30000},
)
await page.waitForSelector(
    "#results .result-card",
    {"visible": True, "timeout": 20000},
)
html = await page.content()

Choose a real application selector, such as a results container, article heading, table row, or logged-in dashboard element. Do not use a generic selector like body as proof that the desired data is ready.

5. Combine network idle with a selector when needed

Some pages need both resource settling and an application-specific condition. Network idle is useful as a milestone, while the selector confirms that the requested content exists.

await page.goto(
    "https://example.com/app",
    {"waitUntil": "networkidle2", "timeout": 60000},
)
await page.waitForSelector(
    "main[data-ready='true']",
    {"timeout": 30000},
)
content = await page.content()

Lazy-loaded pages may need scrolling or an additional, bounded delay after the selector appears. Prefer a selector or state attribute over an arbitrary sleep because it adapts to variable load times.

6. Check rendering evidence when markup exists but looks blank

HTML can contain the application while the visible page remains empty. Inspect dimensions, visibility, text, and ancestors:

evidence = await page.evaluate("""() => {
  const root = document.querySelector('#app');
  if (!root) return {exists: false};
  const style = getComputedStyle(root);
  const rect = root.getBoundingClientRect();
  return {
    exists: true,
    text: root.innerText,
    width: rect.width,
    height: rect.height,
    display: style.display,
    visibility: style.visibility,
    opacity: style.opacity,
    html: root.outerHTML.slice(0, 1000)
  };
}""")
print(evidence)

Also capture browser console messages, uncaught page errors, and failed document, script, stylesheet, and data requests. These logs identify concrete failures such as a missing bundle, a rejected API call, or a JavaScript exception instead of guessing at navigation settings.

7. Use Pyppeteer evaluation syntax correctly

Pyppeteer tries to infer whether a string passed to evaluate() is a function or an expression. The project README documents that this inference can fail. For a plain expression, pass force_expr=True:

text = await page.evaluate(
    "document.body.textContent",
    force_expr=True,
)

For more complex work, pass a JavaScript function and arguments explicitly:

heading = await page.evaluate("""() => {
  const node = document.querySelector('h1');
  return node ? node.textContent.trim() : null;
}""")

If this diagnostic call fails, the navigation may have succeeded; the extraction expression itself is the problem.

8. Complete extraction example with retries and evidence

import asyncio
import json
import pyppeteer

URL = "https://example.com/app"
READY = "main article"

async def extract():
    browser = await pyppeteer.launch(
        headless=True,
        args=["--no-sandbox"],
    )
    page = await browser.newPage()
    await page.setViewport({"width": 1440, "height": 900})

    errors = []
    page.on("pageerror", lambda exc: errors.append(f"pageerror: {exc}"))
    page.on(
        "requestfailed",
        lambda req: errors.append(f"requestfailed: {req.url} {req.failure}"),
    )

    try:
        response = await page.goto(
            URL,
            {"waitUntil": "domcontentloaded", "timeout": 45000},
        )
        await page.waitForSelector(READY, {"timeout": 30000})
        result = {
            "final_url": page.url,
            "status": None if response is None else response.status,
            "title": await page.title(),
            "html": await page.content(),
            "errors": errors,
        }
        return result
    finally:
        await browser.close()

result = asyncio.run(extract())
with open("page.json", "w", encoding="utf-8") as file:
    json.dump(result, file, ensure_ascii=False, indent=2)

Use a bounded timeout for every wait. Save the final URL, status, HTML length, and error list with each failed reproduction so a later fix can be compared with evidence.

9. Troubleshooting common errors

Symptom or error Likely cause Fix
goto() raises “Invalid URL” Missing scheme or malformed URL Use a complete https:// or http:// URL and log the final URL.
SSL or certificate error Certificate or TLS problem at the target Fix the target certificate where possible; record the exact exception before considering environment-specific launch settings.
Navigation timeout Slow resources, a never-ending connection, or an unsuitable milestone Set a realistic timeout, try domcontentloaded, and separately wait for the required selector.
Response is None about:blank or hash-only same-URL navigation Print page.url and verify that a real document navigation occurred.
HTML contains only a root element Client-side app has not rendered data Wait for the content selector, inspect API requests and page errors, and verify the app’s loading state.
Network idle never completes Polling, analytics, WebSockets, or streaming requests Use domcontentloaded or load, then wait for a meaningful selector.
Selector wait times out Wrong selector, failed render, redirect, authentication, or content not present Print page.url, inspect page.content(), and confirm the selector in a normal browser.
Body text is empty but HTML exists Hidden content, empty shell, or data rendered elsewhere Inspect descendants, computed styles, dimensions, shadow roots, and the requests that supply data.
evaluate() throws or returns an unexpected value Expression/function inference failed Use force_expr=True for expressions or pass an explicit JavaScript function.
Scripts or API calls fail Blocked request, credentials, CORS, bot check, or runtime exception Collect console, page-error, and request-failure logs; reproduce with the same headers and authentication context.

10. Performance, reliability, and cost considerations

  • Use the earliest reliable milestone. Start with domcontentloaded for server-rendered pages. Add a selector wait for JavaScript-rendered content instead of making every page wait for global network idle.
  • Keep waits bounded. Separate navigation and selector timeouts so a stalled background connection does not hide the real failure.
  • Reuse a browser when processing many URLs. Launching Chromium is expensive; create new pages or contexts per job and close them deterministically. Watch memory when pages contain large images or long-lived connections.
  • Capture evidence on failure. Store the final URL, response status, a short HTML preview, console errors, page errors, and failed-request URLs. This makes retries targeted rather than blind.
  • Retry only transient failures. A short retry can help with temporary network errors, but repeated selector timeouts usually indicate a wrong readiness condition or a deterministic application failure.
  • Plan for version drift. The Pyppeteer repository describes the project as unmaintained and points readers toward Puppeteer documentation and troubleshooting. Record your Pyppeteer, Python, Chromium, and launch-argument versions, and verify current upstream advice against the installed package.
  • Installation context. The project README estimates an approximately 150 MB first-run Chromium download when Chromium is not already available. Cache the browser in build environments where appropriate.

11. A practical investigation checklist

  1. Confirm the URL includes a scheme.
  2. Log page.url, the navigation response, and status.
  3. Catch and preserve the exact goto() exception.
  4. Print page.title(), page.content(), and body text.
  5. Choose a selector that proves the requested content exists.
  6. Wait for that selector with a bounded timeout.
  7. Capture console, page-error, and request-failure events.
  8. Inspect visibility, dimensions, and ancestor styles if markup exists but looks blank.
  9. Verify evaluate() syntax, using force_expr=True for expressions.
  10. Record Pyppeteer, Python, Chromium, launch settings, URL, status, and logs.

Or skip the browser setup

If you only need a rendered screenshot or PDF, ScreenshotNeo handles the browser capture through one request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients take screenshots with take_screenshot, inspect pages with get_page_info, and create PDFs with capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Does networkidle0 guarantee that HTML is ready?

No. It describes connection activity for 500 ms. A selector for the content you need is a stronger readiness check.

Why can goto() succeed while the page is blank?

Navigation can succeed before a single-page application fetches data, or the app can fail during JavaScript execution. Inspect the DOM, console, page errors, and failed requests.

Should I add a fixed sleep?

Use a bounded delay only when the site has a known animation or lazy-load gap. Prefer a meaningful selector or state attribute because fixed sleeps are slower and less reliable.

What details should I include when asking for help?

Provide the target URL, minimal script, Pyppeteer and Chromium versions, launch arguments, final URL, response status, extracted HTML preview, and concise console and network errors.