How to Navigate to the Next Page With Pyppeteer
Click site pagination safely with Pyppeteer, wait for the right state, handle infinite scroll and failures, and know when browser automation is unnecessary.
For a normal pagination link or button, start the navigation wait and the click at the same time:
await asyncio.gather(
page.waitForNavigation(),
page.click('YOUR_NEXT_SELECTOR'),
)
Replace YOUR_NEXT_SELECTOR with the selector for the target site’s actual Next control. Starting waitForNavigation() only after clicking can miss a fast navigation and cause a timeout. Pyppeteer treats a new document, reload, and History API URL change as navigation; an anchor or History API transition can return None even though the URL changed. See the Pyppeteer API reference.
Choose the right meaning of “next page”
Before writing a loop, inspect the page and decide which event you need to perform:
| What “next” means | Use | Success signal |
|---|---|---|
| The site’s Next link or button loads another document | asyncio.gather(page.waitForNavigation(), page.click(...)) |
URL or document changes |
| The browser’s next history entry | await page.goForward() |
History entry becomes active |
| The site replaces results in place | Click, then waitForSelector() or waitForFunction() |
New item, page marker, or results container state |
goForward() is for browser history. It does not click a site’s pagination control and returns None when there is no forward entry.
Inspect the pagination control first
- Open the target page in a browser and inspect the Next control.
- Record a stable selector, such as an accessible label, pagination class, or link relation. Treat selectors in examples as placeholders.
- Check how the final page is represented: missing control,
disabledattribute,aria-disabled="true", a disabled class, or a URL that no longer changes. - Identify a post-click condition that was false before the click. A selector that already exists can resolve immediately and give you a false success.
For a link, useful markup might look like <a rel="next" href="/items?page=2">Next</a>. For a button, inspect its attributes and the results container. Do not assume that button, visible text, or .next is universal.
Complete Pyppeteer example for document navigation
import asyncio
from pyppeteer import launch
from pyppeteer.errors import TimeoutError
START_URL = "https://example.com/items?page=1"
NEXT_SELECTOR = 'a[rel="next"]' # Replace with the real selector
async def click_next(page):
# Fail clearly when the control is absent or no longer usable.
next_control = await page.querySelector(NEXT_SELECTOR)
if next_control is None:
return False
disabled = await page.evaluate(
"""(element) => {
const aria = element.getAttribute('aria-disabled');
return element.hasAttribute('disabled') || aria === 'true' ||
element.classList.contains('disabled');
}""",
next_control,
)
if disabled:
return False
try:
await asyncio.gather(
page.waitForNavigation({
"waitUntil": "networkidle2",
"timeout": 30000,
}),
page.click(NEXT_SELECTOR),
)
except TimeoutError:
raise RuntimeError(
f"Next control was clicked, but navigation did not finish: {page.url}"
)
return True
async def main():
browser = await launch({"headless": True})
try:
page = await browser.newPage()
await page.setViewport({"width": 1366, "height": 900})
await page.goto(START_URL, {"waitUntil": "networkidle2", "timeout": 30000})
for page_number in range(1, 101):
print(f"Processing page {page_number}: {page.url}")
# Replace this with extraction for the current page.
# items = await page.querySelectorAll('.result')
moved = await click_next(page)
if not moved:
print("Reached the final page")
break
finally:
await browser.close()
if __name__ == "__main__":
asyncio.get_event_loop().run_until_complete(main())
Install Pyppeteer with pip install pyppeteer. The project describes itself as an unofficial Python port of Puppeteer and currently documents Python 3.8 or newer. On first use it may download a Chromium build of roughly 150 MB, unless a suitable browser executable is already available; verify current details in the project README.
Use a URL change as the success condition
If the control navigates through a predictable query parameter, record the old URL and verify that it changed:
old_url = page.url
await asyncio.gather(
page.waitForNavigation({"waitUntil": "domcontentloaded"}),
page.click(NEXT_SELECTOR),
)
if page.url == old_url:
raise RuntimeError("The click completed without changing the URL")
This catches controls that were clicked but did not activate because an overlay, disabled state, or event handler blocked them.
Pagination that updates the existing document
Many React, Vue, and other single-page interfaces keep the URL and document while replacing the results. In that case, waiting for navigation is the wrong signal. Capture a value before clicking, then wait for it to change.
old_first_item = await page.evaluate("""() => {
const item = document.querySelector('.result');
return item ? item.textContent.trim() : null;
}""", force_expr=True)
await page.click('button[aria-label="Next"]') # Replace selector
await page.waitForFunction(
"""oldValue => {
const item = document.querySelector('.result');
return item && item.textContent.trim() !== oldValue;
}""",
{"timeout": 30000},
old_first_item,
)
You can instead wait for a page marker or newly inserted item:
await page.click('YOUR_NEXT_SELECTOR')
await page.waitForSelector(
'.results-page[data-page="2"]',
{"timeout": 30000},
)
Do not wait for a selector that was already present before the click. If the application reuses the same node, wait for a changed attribute, text value, item count, or a loading indicator to disappear.
Wait for loading to finish
If the interface shows a spinner, combine a state change with spinner removal:
await page.click('YOUR_NEXT_SELECTOR')
await page.waitForFunction(
"""() => !document.querySelector('.loading[aria-busy="true"]')""",
{"timeout": 30000},
)
Use the signal that represents visible, usable results. A fixed asyncio.sleep() can be too short on a slow run and wasteful on a fast run.
Use browser history with goForward()
When your script deliberately called goBack() and wants to restore the next history entry, use:
result = await page.goForward({
"waitUntil": "networkidle2",
"timeout": 30000,
})
if result is None:
print("There is no forward history entry")
This does not discover or activate a website pagination control. It only moves through the browser’s existing history.
Loop through pages safely
A robust loop needs a stop condition, duplicate detection, and cleanup:
seen_urls = set()
for _ in range(100):
if page.url in seen_urls:
raise RuntimeError(f"Pagination loop detected at {page.url}")
seen_urls.add(page.url)
# Extract the current page before moving on.
# await save_items(await page.querySelectorAll('.result'))
if not await click_next(page):
break
else:
raise RuntimeError("Stopped after 100 pages; pagination may be looping")
For in-place pagination, replace URL tracking with a set of page numbers or a hash of the visible item IDs. Also stop when the Next control disappears or is disabled.
Selectors, frames, and click details
Prefer stable selectors
- Prefer
rel="next", an accessible label, a data attribute, or a stable pagination container. - Avoid generated CSS module names and deeply nested paths that change with every build.
- If visible text is the only stable clue, locate the element and verify its text before clicking.
Handle an iframe
If pagination is inside an iframe, query the frame rather than the top-level page:
frame = next(
frame for frame in page.frames
if frame.url.startswith("https://example.com/widget")
)
await frame.click('YOUR_NEXT_SELECTOR')
await frame.waitForSelector('.new-result', {"timeout": 30000})
Use the frame’s URL or another identifying property; frame order is not a stable identifier.
Scroll the control into view
await page.evaluate(
"""selector => document.querySelector(selector)?.scrollIntoView({block: 'center'})""",
NEXT_SELECTOR,
)
await page.click(NEXT_SELECTOR)
Scrolling helps when a sticky header or viewport boundary intercepts the click, but it does not fix a wrong selector or disabled control.
Debugging and error handling
| Symptom | Likely cause | Fix |
|---|---|---|
TimeoutError from waitForNavigation() |
The click did not navigate, the selector is wrong, or the site updates in place. | Confirm the click manually, then use a results-state wait for asynchronous pagination. |
| Navigation succeeds but results are old | The document loaded before the data request completed. | Wait for a result-specific selector, page marker, or network-idle state appropriate to the site. |
| Selector timeout | The control is absent, rendered later, inside an iframe, or changed markup. | Inspect the current HTML, wait for the pagination container, and query the correct frame. |
| Click intercepted or no action | Cookie banner, popup, overlay, or sticky element covers the control. | Dismiss the overlay, scroll the control into view, or click only after the blocking element is gone. |
| Infinite loop | The site keeps the control visible on the final page or returns duplicate content. | Check disabled state, track URLs or item IDs, and enforce a maximum page count. |
goForward() returns None |
No forward history entry exists. | Use the site’s Next control or navigate to a known pagination URL. |
| Chromium launch failure | Browser download, executable path, sandbox, or system dependency issue. | Review the current Pyppeteer README, configure an available Chromium executable, and capture launch stderr. |
For diagnostics, log the current URL, selector, page number, and a short HTML snippet around the pagination container. Keep the browser close in a finally block so a failed navigation does not leave Chromium processes running.
Performance and reliability
- Reuse one browser: launch once and create a page or context per job instead of starting Chromium for every URL.
- Choose the narrowest wait:
domcontentloadedis often faster thannetworkidle2, but only use it when your extraction does not depend on later requests. - Keep timeouts explicit: set navigation and selector timeouts, then report the URL and page number when they expire.
- Limit concurrency: too many simultaneous pages can exhaust CPU, memory, sockets, or the target site’s rate limits.
- Retry carefully: retry transient network failures with backoff, but do not blindly repeat a click that may already have succeeded. Re-check the URL or result marker first.
- Make extraction idempotent: store stable item IDs or canonical URLs so a retry cannot duplicate records.
- Respect the target: follow its terms, robots guidance where applicable, authentication rules, and rate limits.
Or skip the browser setup
If you need an image or PDF of each page rather than DOM interaction, ScreenshotNeo provides a single HTTP request. Its API can wait for selectors or network idle, click elements, use custom headers and cookies, capture full pages, and return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/items?page=2 -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/items?page=2"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/items?page=2' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing result. ScreenshotNeo also has an MCP server so Claude, Cursor, and other MCP clients can take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Cost considerations
Self-hosted Pyppeteer costs depend on your compute, browser downloads, proxy or bandwidth use, and maintenance time. Waiting for every request to become idle can increase runtime on pages with analytics or long polling. Cache pages when permitted, block unnecessary resources, and avoid reopening Chromium for each page.
With ScreenshotNeo, only clean shots are billed; failed loads and cache hits are not. Its plans include Free (1,000 shots/month), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000). Yearly billing gives two months free, and every feature is available on every plan.
FAQ
What selector should I use for Next?
There is no universal selector. Inspect the target site and prefer a stable relation, accessible label, or data attribute. Keep the selector configurable.
Why does waitForNavigation() return None?
Pyppeteer can consider a History API or anchor URL change navigation without returning a new response object. Verify the URL or page state instead.
Can I use page.goForward() for pagination?
Only when the desired page is already the next browser history entry. It does not replace clicking a site’s Next control.
How do I know pagination is complete?
Check the site’s actual disabled or absent state, and add a duplicate URL or item-ID check plus a maximum page limit.
Does Pyppeteer support infinite scroll?
Yes, but infinite scroll is not a Next click. Scroll, wait for a new item count or marker, and stop when no new content appears after a bounded number of attempts.


