How to Click the Next Page Button While Scraping with Pyppeteer
Use Pyppeteer to click pagination safely, wait for the right update, handle final pages, and avoid navigation races.
Direct answer: find the site’s real next-page control, wait until it is visible, then click it. If the click loads a new document, start waitForNavigation() and click() together with asyncio.gather(). If the page updates in place, wait for a page-specific change such as a new heading, first result, or loading marker. Do not rely on a fixed sleep as proof that pagination finished.
The examples below use the Pyppeteer API documented for version 0.0.25. Pyppeteer describes itself as an unofficial Puppeteer port, and the sources consulted do not establish current package maintenance or compatibility with current Chromium. Treat selectors, browser versions, and timeout values as site-specific configuration.
1. Identify how the site paginates
A “Next” control can produce several different behaviors:
| Behavior | What changes | What to wait for |
|---|---|---|
| Full document navigation | A new HTML document loads, often with a new URL | waitForNavigation() started concurrently with the click |
| History API or anchor update | The URL or results change without a conventional document load | A changed result, heading, or URL; navigation may return None |
| In-place pagination | Rows are replaced inside the current document | waitForFunction() or a selector tied to the new content |
| Load more or infinite scroll | Additional rows are appended | A row count increase, loading marker disappearance, or new item |
Inspect the page in browser developer tools or save a snapshot before choosing a selector. Prefer stable attributes such as aria-label, a data attribute, or a semantic relationship. Avoid assuming every site uses the same text, tag, or class name.
2. Install Pyppeteer and open the page
python -m pip install pyppeteer
import asyncio
from pyppeteer import launch
async def main():
browser = await launch(headless=True)
page = await browser.newPage()
await page.goto('https://example.com/results', {
'waitUntil': 'networkidle2',
'timeout': 30000,
})
print(await page.title())
await browser.close()
asyncio.run(main())
The documented default navigation timeout is 30 seconds and the default waitUntil condition is load. Set these explicitly when the target site needs a different policy.
3. Click when Next causes full document navigation
Start both coroutines at the same time. Waiting for navigation only after the click can miss the event. The Pyppeteer API warns that a separate wait can race with the click.
import asyncio
from pyppeteer import launch
URL = 'https://example.com/results?page=1'
NEXT_SELECTOR = 'a[rel="next"]'
async def main():
browser = await launch(headless=True)
page = await browser.newPage()
await page.goto(URL, {'waitUntil': 'networkidle2', 'timeout': 30000})
await page.waitForSelector(NEXT_SELECTOR, {'visible': True, 'timeout': 30000})
response = await asyncio.gather(
page.waitForNavigation({
'waitUntil': 'networkidle2',
'timeout': 30000,
}),
page.click(NEXT_SELECTOR),
)
navigation_response = response[0]
print('URL:', page.url)
print('Navigation response:', navigation_response)
await browser.close()
asyncio.run(main())
This is the documented race-free pattern:
await asyncio.gather(
page.waitForNavigation(),
page.click(next_selector),
)
A navigation response is not a universal success signal. Anchor changes and History API updates can count as navigation while returning None, so also verify that the expected content changed.
4. Handle in-place updates and History API pagination
When the document remains loaded, wait for an observable change. Capture the first result before clicking, then wait until its value differs.
import asyncio
from pyppeteer import launch
URL = 'https://example.com/results'
NEXT_SELECTOR = 'button.next'
FIRST_RESULT_SELECTOR = '.result-card h2'
async def main():
browser = await launch(headless=True)
page = await browser.newPage()
await page.goto(URL, {'waitUntil': 'networkidle2', 'timeout': 30000})
await page.waitForSelector(FIRST_RESULT_SELECTOR, {'visible': True})
old_first = await page.Jeval(FIRST_RESULT_SELECTOR, '(el) => el.textContent')
await page.click(NEXT_SELECTOR)
await page.waitForFunction(
'''(selector, oldValue) => {
const el = document.querySelector(selector);
return el && el.textContent.trim() !== oldValue.trim();
}''',
{'timeout': 30000},
FIRST_RESULT_SELECTOR,
old_first,
)
new_first = await page.Jeval(FIRST_RESULT_SELECTOR, '(el) => el.textContent')
print('Changed from:', old_first.strip())
print('Changed to:', new_first.strip())
await browser.close()
asyncio.run(main())
Other useful conditions include waiting for a page-number element to change, waiting for a loading element to disappear, or waiting for a new result identifier. The exact expression must match the target site’s markup.
5. Build a loop that stops safely
A production scraper should stop when the control is missing, hidden, disabled, or reaches the final page. Check the control before every click and collect the current page before moving on.
import asyncio
from pyppeteer import launch
URL = 'https://example.com/results'
NEXT_SELECTOR = 'a[rel="next"]'
ROW_SELECTOR = '.result-card'
MAX_PAGES = 20
async def is_next_available(page):
return await page.evaluate('''(selector) => {
const el = document.querySelector(selector);
if (!el) return false;
const style = getComputedStyle(el);
const disabled = el.matches(':disabled') ||
el.getAttribute('aria-disabled') === 'true' ||
el.classList.contains('disabled');
return !disabled && style.display !== 'none' &&
style.visibility !== 'hidden';
}''', NEXT_SELECTOR)
async def read_rows(page):
return await page.evaluate('''(selector) =>
Array.from(document.querySelectorAll(selector)).map(el => ({
text: el.innerText.trim(),
href: el.querySelector('a')?.href || null
}))''', ROW_SELECTOR)
async def main():
browser = await launch(headless=True)
page = await browser.newPage()
await page.goto(URL, {'waitUntil': 'networkidle2', 'timeout': 30000})
all_rows = []
for page_number in range(1, MAX_PAGES + 1):
await page.waitForSelector(ROW_SELECTOR, {'visible': True, 'timeout': 30000})
rows = await read_rows(page)
all_rows.extend(rows)
if not await is_next_available(page):
break
before = await page.Jeval(ROW_SELECTOR, '(el) => el.innerText')
await page.click(NEXT_SELECTOR)
await page.waitForFunction(
'''(selector, oldText) => {
const el = document.querySelector(selector);
return el && el.innerText !== oldText;
}''',
{'timeout': 30000},
ROW_SELECTOR,
before,
)
print('Collected rows:', len(all_rows))
await browser.close()
asyncio.run(main())
The loop includes a maximum page limit so a broken selector or repeating page cannot run forever. Add deduplication using a stable result ID or URL when pages can overlap.
6. Choose and validate selectors
- Prefer semantic attributes:
a[rel="next"],button[aria-label="Next page"], or a site-specific data attribute. - Scope the selector: target the pagination container if the page has several buttons with similar labels.
- Check visibility: use
waitForSelector(selector, {'visible': True})before clicking. - Check state: inspect
disabled,aria-disabled, and classes used by the site for the final page. - Use XPath only when needed: Pyppeteer exposes
page.xpath(); text-based XPath can be fragile when labels are localized.
Pyppeteer provides Python-named equivalents for Puppeteer’s query methods, including querySelector(), querySelectorAll(), and xpath(). Its evaluate() method accepts a JavaScript expression or function string; if an expression is detected incorrectly, the project README documents the force_expr=True option.
7. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
PageError or no matching element |
The selector is wrong, the control has not rendered, or it is inside a frame | Inspect the live DOM, wait for the correct container, and switch to the relevant frame if necessary. |
waitForSelector times out |
The selector never appears or is hidden | Verify the selector and use visible: True only when visibility is required. Increase the timeout only after confirming the page is genuinely slow. |
| Navigation timeout after clicking | The click does not cause a full document navigation, or the page is slow | Use a content-based waitForFunction() condition for in-place updates, or adjust waitUntil and timeout for a real navigation. |
Navigation response is None |
The site used an anchor or History API update | Check page.url and wait for changed result content. |
| Click succeeds but rows are unchanged | The wait condition is too weak or the click hit the wrong control | Compare a stable first-row ID or heading before and after the click. |
| Click is intercepted or has no effect | An overlay, cookie banner, or sticky element covers the control | Wait for the overlay to disappear, scroll the control into view, or handle the site’s consent flow before clicking. |
| Repeated pages or an infinite loop | The site returned the same results or the final-state check is incomplete | Track page URLs or stable row IDs and enforce a maximum page count. |
| Browser fails to launch | Chromium is unavailable or incompatible with the installed package | Review the Pyppeteer launch error and package/browser version pairing. The research sources do not establish current Chromium compatibility. |
8. Performance and reliability practices
- Use the narrowest stable selector and avoid repeated full-page DOM scans.
- Collect only fields needed for the job instead of serializing entire HTML documents.
- Use a page-specific readiness condition rather than a long fixed delay.
- Keep navigation and click waits concurrent for document navigations.
- Set explicit timeouts and catch failures per page so one broken page does not discard earlier results.
- Deduplicate by canonical URL or result ID when pagination overlaps.
- Log the page number, URL, selector, and wait condition when a click fails.
- Close the browser in a
finallyblock in long-running jobs to avoid orphaned Chromium processes.
Pyppeteer’s documented defaults are useful starting points, but the correct timeout depends on the target site and network. A larger timeout does not fix a selector that can never match.
9. Or skip the browser setup
If you only need a clean screenshot of each result page, ScreenshotNeo provides a single HTTP request instead of maintaining a browser script. See the ScreenshotNeo API docs for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account to get started.
10. FAQ
Should I use waitForNavigation() after every click?
No. Use it for a click that causes a document navigation. For in-place updates, wait for changed content or another condition that proves the new results are ready.
Why does the click/navigation race happen?
The navigation event can begin immediately when the click fires. Starting the wait and click together ensures the listener is active before the event occurs.
Can I scrape a site with a disabled Next button?
Yes, if you detect the disabled state and stop. Inspect the site’s actual disabled, aria-disabled, or class convention instead of assuming one universal markup pattern.
Is Pyppeteer the same as Puppeteer?
No. Pyppeteer is an unofficial Python port. The API behavior described here comes from its 0.0.25 documentation and should not be treated as proof of compatibility with every current Puppeteer feature.
What if pagination is infinite scroll?
Model it as an interaction followed by a condition such as an increased row count, a new item ID, or disappearance of a loading marker. Numbered-page navigation assumptions do not apply.


