Why Pyppeteer page.goto Hangs Despite a 1000 ms Timeout
A 1000 ms Pyppeteer timeout does not guarantee a 1000 ms return. Learn why networkidle0 hangs and how to enforce a hard deadline.
page.goto() can appear to ignore timeout=1000 when the selected readiness condition never becomes true. The most common cause is waitUntil: 'networkidle0': Pyppeteer waits for the page to have zero active network connections for at least 500 milliseconds. Pages with analytics, polling, streaming, ads, chat, or other long-lived requests may never reach that state during the interval you expected.
The navigation timeout is a deadline used by Pyppeteer’s navigation watcher. It does not change what counts as a successful navigation, and it is not automatically a deadline for every coroutine involved in browser startup, page creation, redirects, or cleanup.
What the 1000 ms timeout actually controls
In this call:
await page.goto(url, {'waitUntil': 'networkidle0', 'timeout': 1000})
timeout is measured in milliseconds. Pyppeteer creates a navigation watcher, waits for the requested lifecycle event, and raises when the navigation deadline is exceeded. The documented failure cases include an SSL error, an invalid URL, a timeout during navigation, or failure of the main resource. See the Pyppeteer Page.goto reference.
waitUntil defines the readiness signal:
| Value | Meaning | Typical use |
|---|---|---|
load |
Wait for the page’s load event. | Basic documents where subresources should be loaded. |
domcontentloaded |
Wait for the DOMContentLoaded event. | Fast scraping or screenshots when the required content is available in the initial DOM. |
networkidle0 |
Wait until there are no more than zero active connections for at least 500 ms. | Pages that truly become quiet after loading. |
networkidle2 |
Wait until there are no more than two active connections for at least 500 ms. | Pages with a small amount of expected background traffic. |
You can select one value or combine lifecycle values. Network-idle conditions are global page conditions; they do not know which element your application needs.
The reliable fix: choose a readiness condition you can observe
For most applications, navigate to the document, then wait for the specific selector that proves the page is usable.
import asyncio
from pyppeteer import launch
async def capture(url: str) -> None:
browser = await launch(headless=True)
page = await browser.newPage()
try:
await page.goto(
url,
{
"waitUntil": "domcontentloaded",
"timeout": 10_000,
},
)
await page.waitForSelector(
".content",
{"timeout": 10_000},
)
await page.screenshot({"path": "page.png", "fullPage": True})
finally:
await page.close()
await browser.close()
asyncio.run(capture("https://example.com"))
This separates two questions: has the document started successfully, and is the content your code needs present? A selector or application-state check is usually more deterministic than waiting for every third-party request to stop.
Guarantee an end-to-end deadline with asyncio
If the entire operation must finish within a hard wall-clock budget, wrap the navigation task in an outer asyncio deadline. This protects the caller even when work around navigation takes longer than the page timeout.
import asyncio
from pyppeteer import launch
async def navigate(page, url: str) -> None:
await page.goto(
url,
{
"waitUntil": "domcontentloaded",
"timeout": 10_000,
},
)
await page.waitForSelector(".content", {"timeout": 5_000})
async def capture_with_budget(url: str, budget_seconds: float = 20.0) -> None:
browser = await launch(headless=True)
page = await browser.newPage()
try:
await asyncio.wait_for(navigate(page, url), timeout=budget_seconds)
await page.screenshot({"path": "page.png"})
except asyncio.TimeoutError:
print(f"end-to-end budget exceeded: {budget_seconds}s")
finally:
await page.close()
await browser.close()
asyncio.run(capture_with_budget("https://example.com", 20.0))
The outer timeout is an engineering boundary around your coroutine. It does not claim that Pyppeteer’s own navigation timeout covers browser launch, newPage(), screenshot encoding, or cleanup.
Set a default navigation timeout
Use a per-call timeout when different URLs need different budgets, or set a default for a page:
page.setDefaultNavigationTimeout(10_000)
await page.goto(
"https://example.com",
{"waitUntil": "domcontentloaded"},
)
Passing timeout: 0 disables Pyppeteer’s navigation timeout. That can be useful for controlled experiments, but it removes an important safety boundary; keep an outer asyncio deadline if you do this.
Why networkidle0 can wait forever in practice
- Polling: JavaScript periodically requests updates, so the connection count never stays at zero for 500 ms.
- Streaming or sockets: Server-sent events, WebSockets, and long-polling keep a connection open by design.
- Third-party scripts: Analytics, advertising, consent tools, chat widgets, and embeds can continue making requests after the useful content is rendered.
- Slow resources: A single image, font, or script can keep the network busy beyond a short navigation budget.
- Redirects: The URL may pass through several responses before the final document is reached.
A report involving https://ig.com.br/ illustrates the site-specific nature of this behavior: other URLs timed out normally while that page continued to exercise a different lifecycle. Treat such a case as a readiness-condition mismatch, not proof that the timeout option is ignored.
A complete diagnostic workflow
- Record the inputs. Log the URL, selected
waitUntilvalue, configured timeout, start time, and final elapsed time. - Attach listeners before navigation. This shows redirects, failed requests, and responses that continue after the document is usable.
- Try
domcontentloaded. If it succeeds quickly, persistent background traffic is the likely cause. - Wait for the application signal. Use
waitForSelectoror a page-state predicate for the content you actually require. - Inspect the main response. Check URL validity, SSL errors, redirects, and failures of the main resource.
- Test the browser environment separately. If no request is emitted and even
browser.newPage()is slow, investigate the Chrome executable, Python/Chrome combination, and sandbox configuration.
import asyncio
from pyppeteer import launch
async def debug_navigation(url: str) -> None:
browser = await launch(headless=True)
page = await browser.newPage()
page.on("request", lambda req: print("REQUEST", req.method, req.url))
page.on("response", lambda res: print("RESPONSE", res.status, res.url))
page.on("requestfailed", lambda req: print("FAILED", req.url, req.failure))
try:
await page.goto(
url,
{"waitUntil": "domcontentloaded", "timeout": 10_000},
)
finally:
await page.close()
await browser.close()
asyncio.run(debug_navigation("https://example.com"))
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
TimeoutError with networkidle0 |
Requests never become completely idle. | Use domcontentloaded plus a selector, or use networkidle2 when limited background traffic is acceptable. |
| Timeout despite visible content | The page is usable before the chosen lifecycle condition completes. | Define readiness around the required element or application state. |
| Invalid URL error | The URL is missing a scheme or is malformed. | Pass a complete URL such as https://example.com. |
| SSL error | The target certificate is invalid or self-signed. | Fix the certificate or deliberately configure browser SSL handling only in an environment where that is acceptable. |
| Main resource failed | The server, DNS, connection, or redirect chain failed. | Inspect request failures and the main response; retry only when the failure is transient. |
browser.newPage() hangs |
Browser/protocol startup or the local Chrome environment is unhealthy. | Verify the Chrome executable, Python/Chrome compatibility, and sandbox settings. This is outside the navigation timeout. |
| Outer deadline fires but process remains | Pages, browser processes, or tasks were not cleaned up. | Close the page and browser in finally; cancel and await child tasks where applicable. |
Performance and reliability guidance
- Prefer the earliest lifecycle event that provides the data you need. Waiting for global network idle adds latency and couples reliability to third-party traffic.
- Use separate budgets for navigation, selector readiness, and the complete job. A single 1000 ms value rarely fits every stage.
- Keep cleanup unconditional. Close pages and browsers after success, timeout, and cancellation.
- Record the lifecycle mode and elapsed time with each result so a slow page can be distinguished from a browser startup problem.
- Retry only failures that are plausibly transient. Repeating a permanently open streaming page with
networkidle0will reproduce the same wait. - When taking screenshots at scale, reuse a healthy browser process while isolating pages, and enforce a per-job outer deadline so one URL cannot occupy the worker indefinitely.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Does timeout=1000 mean the function must return in one second?
No. It is Pyppeteer’s navigation deadline. Use an outer asyncio timeout for a hard end-to-end limit.
Should I always replace networkidle0?
No. Keep it when the page is known to become quiet and that condition matters. Otherwise, use a lifecycle event plus a selector or application-state check.
What is the difference between networkidle0 and networkidle2?
networkidle0 requires zero active connections; networkidle2 allows up to two during the 500 ms quiet window.
Can I disable the navigation timeout?
Yes. Pass timeout: 0 or change the default, but retain an outer deadline so a job cannot run forever.
Why does changing the timeout not fix a single troublesome domain?
The domain may have a persistent connection, a failed main resource, an SSL problem, or a browser-environment issue. Identify which layer is failing before changing the number.


