How to Fix Blank HTML After a Page Loads in Pyppeteer
Diagnose blank Pyppeteer HTML by checking navigation, DOM timing, selectors, JavaScript errors, and failed requests with runnable fixes.
Direct answer: a completed page.goto() only means that the selected navigation milestone finished. It does not prove that a client-rendered application has inserted the content you need. Log the navigation response and current URL, inspect page.content() and the body text, then wait for a meaningful application selector before extracting HTML. If the DOM is still empty, investigate the URL, HTTP response, JavaScript errors, failed requests, and your evaluation syntax.
1. Run a diagnostic capture first
Use this small script against the affected URL. It records the evidence needed to distinguish navigation failure, timing problems, and rendering errors.
import asyncio
import pyppeteer
TARGET = "https://example.com"
async def main():
browser = await pyppeteer.launch(headless=True)
page = await browser.newPage()
page.on("console", lambda msg: print("CONSOLE:", msg.type, msg.text))
page.on("pageerror", lambda exc: print("PAGE ERROR:", exc))
page.on(
"requestfailed",
lambda req: print("REQUEST FAILED:", req.url, req.failure),
)
try:
response = await page.goto(
TARGET,
{"waitUntil": "load", "timeout": 30000},
)
print("current URL:", page.url)
print("response object:", response)
if response is not None:
print("response URL:", response.url)
print("status:", response.status)
print("title:", await page.title())
html = await page.content()
print("HTML length:", len(html))
print("HTML preview:", html[:1000])
body_text = await page.evaluate(
"document.body ? document.body.innerText : ''",
force_expr=True,
)
print("body text preview:", body_text[:1000])
except Exception as exc:
print("NAVIGATION ERROR:", repr(exc))
finally:
await browser.close()
asyncio.run(main())
Pyppeteer normally returns the main-resource response from goto(). A None response is expected for about:blank and for a same-URL navigation that changes only the hash. The API can also raise for an invalid URL, SSL failure, timeout, or main-resource failure. See the Pyppeteer API reference.
2. Confirm that you navigated to the intended document
A surprisingly common cause of “blank HTML” is an incomplete or unintended URL. Pass a complete URL with https:// or http://, and print page.url after navigation.
response = await page.goto(
"https://example.com/products",
{"waitUntil": "domcontentloaded", "timeout": 30000},
)
print("landed on:", page.url)
print("status:", None if response is None else response.status)
Check redirects, authentication pages, consent interstitials, and bot checks. A successful transport request can still land on an error or challenge document. Preserve the final URL and status in your diagnostic output.
3. Separate an empty document from extraction timing
page.content() serializes the current DOM, including the doctype. It tells you what exists at the exact moment you call it; it does not wait for application data. Compare the serialized HTML, title, and body text:
html = await page.content()
title = await page.title()
text = await page.evaluate(
"document.body ? document.body.innerText : ''",
force_expr=True,
)
print({"title": title, "html_chars": len(html), "text": text[:500]})
If you see an application shell such as a root element but not the expected records, navigation worked and rendering is incomplete. If the expected element exists but has no text, inspect its descendants and computed styles. If the serialized document itself is unexpectedly small, revisit the URL, response, redirects, and page errors.
4. Wait for the page’s actual readiness condition
Pyppeteer’s goto() default is load. Its documented alternatives are:
| Option | What it establishes | When to use it | Limitation |
|---|---|---|---|
domcontentloaded |
Initial HTML has been parsed | Fast extraction from server-rendered markup | Scripts, images, and data may still be pending |
load |
Load event fired | Default when page resources should be loaded | Single-page apps may render later |
networkidle0 |
No more than zero connections for 500 ms | Pages that become truly quiet | Analytics, sockets, or polling can prevent completion |
networkidle2 |
No more than two connections for 500 ms | Pages with a small amount of background traffic | Network quiet does not prove your content exists |
For client-rendered pages, wait for a selector that represents the content you actually need. The selector wait resolves when a matching element appears and raises on timeout.
await page.goto(
"https://example.com/search?q=pyppeteer",
{"waitUntil": "domcontentloaded", "timeout": 30000},
)
await page.waitForSelector(
"#results .result-card",
{"visible": True, "timeout": 20000},
)
html = await page.content()
Choose a real application selector, such as a results container, article heading, table row, or logged-in dashboard element. Do not use a generic selector like body as proof that the desired data is ready.
5. Combine network idle with a selector when needed
Some pages need both resource settling and an application-specific condition. Network idle is useful as a milestone, while the selector confirms that the requested content exists.
await page.goto(
"https://example.com/app",
{"waitUntil": "networkidle2", "timeout": 60000},
)
await page.waitForSelector(
"main[data-ready='true']",
{"timeout": 30000},
)
content = await page.content()
Lazy-loaded pages may need scrolling or an additional, bounded delay after the selector appears. Prefer a selector or state attribute over an arbitrary sleep because it adapts to variable load times.
6. Check rendering evidence when markup exists but looks blank
HTML can contain the application while the visible page remains empty. Inspect dimensions, visibility, text, and ancestors:
evidence = await page.evaluate("""() => {
const root = document.querySelector('#app');
if (!root) return {exists: false};
const style = getComputedStyle(root);
const rect = root.getBoundingClientRect();
return {
exists: true,
text: root.innerText,
width: rect.width,
height: rect.height,
display: style.display,
visibility: style.visibility,
opacity: style.opacity,
html: root.outerHTML.slice(0, 1000)
};
}""")
print(evidence)
Also capture browser console messages, uncaught page errors, and failed document, script, stylesheet, and data requests. These logs identify concrete failures such as a missing bundle, a rejected API call, or a JavaScript exception instead of guessing at navigation settings.
7. Use Pyppeteer evaluation syntax correctly
Pyppeteer tries to infer whether a string passed to evaluate() is a function or an expression. The project README documents that this inference can fail. For a plain expression, pass force_expr=True:
text = await page.evaluate(
"document.body.textContent",
force_expr=True,
)
For more complex work, pass a JavaScript function and arguments explicitly:
heading = await page.evaluate("""() => {
const node = document.querySelector('h1');
return node ? node.textContent.trim() : null;
}""")
If this diagnostic call fails, the navigation may have succeeded; the extraction expression itself is the problem.
8. Complete extraction example with retries and evidence
import asyncio
import json
import pyppeteer
URL = "https://example.com/app"
READY = "main article"
async def extract():
browser = await pyppeteer.launch(
headless=True,
args=["--no-sandbox"],
)
page = await browser.newPage()
await page.setViewport({"width": 1440, "height": 900})
errors = []
page.on("pageerror", lambda exc: errors.append(f"pageerror: {exc}"))
page.on(
"requestfailed",
lambda req: errors.append(f"requestfailed: {req.url} {req.failure}"),
)
try:
response = await page.goto(
URL,
{"waitUntil": "domcontentloaded", "timeout": 45000},
)
await page.waitForSelector(READY, {"timeout": 30000})
result = {
"final_url": page.url,
"status": None if response is None else response.status,
"title": await page.title(),
"html": await page.content(),
"errors": errors,
}
return result
finally:
await browser.close()
result = asyncio.run(extract())
with open("page.json", "w", encoding="utf-8") as file:
json.dump(result, file, ensure_ascii=False, indent=2)
Use a bounded timeout for every wait. Save the final URL, status, HTML length, and error list with each failed reproduction so a later fix can be compared with evidence.
9. Troubleshooting common errors
| Symptom or error | Likely cause | Fix |
|---|---|---|
goto() raises “Invalid URL” |
Missing scheme or malformed URL | Use a complete https:// or http:// URL and log the final URL. |
| SSL or certificate error | Certificate or TLS problem at the target | Fix the target certificate where possible; record the exact exception before considering environment-specific launch settings. |
| Navigation timeout | Slow resources, a never-ending connection, or an unsuitable milestone | Set a realistic timeout, try domcontentloaded, and separately wait for the required selector. |
Response is None |
about:blank or hash-only same-URL navigation |
Print page.url and verify that a real document navigation occurred. |
| HTML contains only a root element | Client-side app has not rendered data | Wait for the content selector, inspect API requests and page errors, and verify the app’s loading state. |
| Network idle never completes | Polling, analytics, WebSockets, or streaming requests | Use domcontentloaded or load, then wait for a meaningful selector. |
| Selector wait times out | Wrong selector, failed render, redirect, authentication, or content not present | Print page.url, inspect page.content(), and confirm the selector in a normal browser. |
| Body text is empty but HTML exists | Hidden content, empty shell, or data rendered elsewhere | Inspect descendants, computed styles, dimensions, shadow roots, and the requests that supply data. |
evaluate() throws or returns an unexpected value |
Expression/function inference failed | Use force_expr=True for expressions or pass an explicit JavaScript function. |
| Scripts or API calls fail | Blocked request, credentials, CORS, bot check, or runtime exception | Collect console, page-error, and request-failure logs; reproduce with the same headers and authentication context. |
10. Performance, reliability, and cost considerations
- Use the earliest reliable milestone. Start with
domcontentloadedfor server-rendered pages. Add a selector wait for JavaScript-rendered content instead of making every page wait for global network idle. - Keep waits bounded. Separate navigation and selector timeouts so a stalled background connection does not hide the real failure.
- Reuse a browser when processing many URLs. Launching Chromium is expensive; create new pages or contexts per job and close them deterministically. Watch memory when pages contain large images or long-lived connections.
- Capture evidence on failure. Store the final URL, response status, a short HTML preview, console errors, page errors, and failed-request URLs. This makes retries targeted rather than blind.
- Retry only transient failures. A short retry can help with temporary network errors, but repeated selector timeouts usually indicate a wrong readiness condition or a deterministic application failure.
- Plan for version drift. The Pyppeteer repository describes the project as unmaintained and points readers toward Puppeteer documentation and troubleshooting. Record your Pyppeteer, Python, Chromium, and launch-argument versions, and verify current upstream advice against the installed package.
- Installation context. The project README estimates an approximately 150 MB first-run Chromium download when Chromium is not already available. Cache the browser in build environments where appropriate.
11. A practical investigation checklist
- Confirm the URL includes a scheme.
- Log
page.url, the navigation response, and status. - Catch and preserve the exact
goto()exception. - Print
page.title(),page.content(), and body text. - Choose a selector that proves the requested content exists.
- Wait for that selector with a bounded timeout.
- Capture console, page-error, and request-failure events.
- Inspect visibility, dimensions, and ancestor styles if markup exists but looks blank.
- Verify
evaluate()syntax, usingforce_expr=Truefor expressions. - Record Pyppeteer, Python, Chromium, launch settings, URL, status, and logs.
Or skip the browser setup
If you only need a rendered screenshot or PDF, ScreenshotNeo handles the browser capture through one request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients take screenshots with take_screenshot, inspect pages with get_page_info, and create PDFs with capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Does networkidle0 guarantee that HTML is ready?
No. It describes connection activity for 500 ms. A selector for the content you need is a stronger readiness check.
Why can goto() succeed while the page is blank?
Navigation can succeed before a single-page application fetches data, or the app can fail during JavaScript execution. Inspect the DOM, console, page errors, and failed requests.
Should I add a fixed sleep?
Use a bounded delay only when the site has a known animation or lazy-load gap. Prefer a meaningful selector or state attribute because fixed sleeps are slower and less reliable.
What details should I include when asking for help?
Provide the target URL, minimal script, Pyppeteer and Chromium versions, launch arguments, final URL, response status, extracted HTML preview, and concise console and network errors.


