How to Get Search Result URLs With Pyppeteer
Use Pyppeteer to wait for rendered results and extract resolved hrefs, with selectors, XPath, troubleshooting, and a ScreenshotNeo shortcut.
Direct answer: launch Chromium, navigate to the rendered search URL, wait for the result-link selector, then read each anchor’s resolved href. Pyppeteer’s querySelectorAllEval is the concise way to run a browser-side function over every matching link.
1. Install Pyppeteer and choose a selector
Pyppeteer is a Python port of Puppeteer. Its documented page methods include querySelector, querySelectorAll, waitForSelector, and XPath support (documentation). A selector is specific to the search page you automate. Inspect the live DOM and identify the anchors that represent organic results; a.result-link below is only a placeholder.
python -m pip install pyppeteer
The first launch can download Chromium. Check the installed Pyppeteer version and browser requirements on your operating system. The reviewed documentation describes version 0.0.25 and may be stale for newer Python or Chromium combinations.
2. Minimal extraction script
import asyncio
from pyppeteer import launch
async def get_result_urls(search_url, selector):
browser = await launch(headless=True)
try:
page = await browser.newPage()
await page.goto(search_url, {'waitUntil': 'domcontentloaded'})
await page.waitForSelector(selector, {'timeout': 10000})
return await page.querySelectorAllEval(
selector,
'(links) => links.map(link => link.href)',
)
finally:
await browser.close()
if __name__ == '__main__':
urls = asyncio.get_event_loop().run_until_complete(
get_result_urls(
'https://example.com/search?q=pyppeteer',
'a.result-link', # Replace after inspecting this page.
)
)
for url in urls:
print(url)
page.goto loads the page, waitForSelector waits for a matching element, and querySelectorAllEval evaluates the JavaScript function against all matches. Reading href returns the browser-resolved URL, so relative links become absolute.
3. A production-ready variant
import argparse
import asyncio
from pyppeteer import launch
from pyppeteer.errors import TimeoutError
async def collect(search_url, selector, timeout=15000):
browser = await launch(headless=True)
try:
page = await browser.newPage()
await page.setDefaultNavigationTimeout(timeout)
await page.goto(search_url, {
'waitUntil': 'domcontentloaded',
'timeout': timeout,
})
await page.waitForSelector(selector, {'timeout': timeout})
rows = await page.querySelectorAllEval(
selector,
"""(links) => links.map(a => ({
href: a.href,
text: (a.textContent || '').trim()
}))""",
)
seen = set()
output = []
for row in rows:
href = row['href']
if href.startswith(('http://', 'https://')) and href not in seen:
seen.add(href)
output.append(row)
return output
finally:
await browser.close()
async def main():
parser = argparse.ArgumentParser()
parser.add_argument('search_url')
parser.add_argument('selector')
args = parser.parse_args()
try:
rows = await collect(args.search_url, args.selector)
except TimeoutError:
raise SystemExit('No matching results appeared before the timeout')
for row in rows:
print(row['href'])
if __name__ == '__main__':
asyncio.get_event_loop().run_until_complete(main())
The deduplication and HTTP(S) filter are application choices, not Pyppeteer requirements. Use --no-sandbox only when your container policy permits it; omit it on a normal desktop session.
4. Wait for JavaScript-rendered results
domcontentloaded means the initial document is parsed. Scripts may insert result links later, so wait for the result selector:
await page.goto(search_url, {'waitUntil': 'domcontentloaded'})
await page.waitForSelector('a.result-link', {'timeout': 10000})
urls = await page.querySelectorAllEval(
'a.result-link',
'links => links.map(link => link.href)',
)
A fixed sleep can work for a quick experiment, but waiting for the element your extraction needs is more reliable. Increase the timeout only after confirming that the selector is correct; a large timeout can hide a wrong selector.
5. CSS selectors, XPath, and evaluate
CSS with querySelectorAllEval
urls = await page.querySelectorAllEval(
'main a.result-link',
'links => links.map(link => link.href)',
)
Element handles with querySelectorAll
handles = await page.querySelectorAll('a.result-link')
urls = []
for handle in handles:
value = await handle.getProperty('href')
urls.append(await value.jsonValue())
XPath
handles = await page.xpath("//a[contains(@class, 'result-link')]")
urls = []
for handle in handles:
value = await handle.getProperty('href')
urls.append(await value.jsonValue())
When to use page.evaluate
Use page.evaluate for extraction logic that does not fit one selector. Pyppeteer accepts a string representation of a JavaScript expression or function. If a bare expression is misclassified, pass force_expr=True:
body_text = await page.evaluate(
'document.body.textContent',
force_expr=True,
)
A function string runs in the page, so it cannot directly access Python variables unless you pass serializable arguments.
6. Inspect and validate the selector
- Open the exact search URL in a regular browser.
- Inspect an organic result and find the anchor owning the destination URL.
- Confirm that the selector excludes navigation, ads, pagination, and related searches.
- Print anchor text and
hreffor a small sample. - Check empty-result and interstitial states before processing the list.
debug = await page.querySelectorAllEval(
'a.result-link',
"links => links.map(a => ({text: (a.textContent || '').trim(), href: a.href}))",
)
print(debug)
Selectors vary by engine, locale, experiment, and page state. The Pyppeteer references do not define a universal search-engine selector, so keep the selector configurable.
7. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
TimeoutError from waitForSelector |
Wrong selector, slow results, or an interstitial. | Inspect the rendered DOM, verify the selector, and adjust the timeout only after verification. |
| Empty list | No elements match at evaluation time. | Wait for the actual result container and test with a broader temporary selector. |
| Redirect or tracking URLs | The anchor itself uses a redirect endpoint. | Read href first; follow redirects separately only when allowed. |
| Only some results | Infinite scroll or pagination has not loaded more items. | Automate the page-specific load action or navigate pages, waiting for new nodes. |
evaluate parsing error |
Expression/function detection failed. | Pass a function string or use force_expr=True for an expression. |
| Browser launch failure | Missing or incompatible Chromium, or host restrictions. | Install browser dependencies, verify the Pyppeteer version, and review launch arguments. |
| Consent, CAPTCHA, or blank page | The result state was never reached. | Detect that state and handle it according to the site’s rules; do not treat it as results. |
8. Performance, reliability, and cost
- Reuse a browser: launch once and create pages for multiple searches.
- Bound waits: set navigation and selector timeouts so a stuck page cannot block a worker forever.
- Close resources: keep shutdown in
finally. - Limit concurrency: parallel pages consume CPU and memory and can trigger site defenses.
- Cache carefully: search pages change; cache only when stale results are acceptable.
- Measure your workload: the cited Pyppeteer sources describe API behavior, not performance benchmarks or compatibility guarantees.
- Respect access rules: check terms, robots guidance, authentication requirements, and applicable law.
Pyppeteer has no per-request price established by the cited sources. Your costs are the machine, browser runtime, bandwidth, and infrastructure you use.
9. Or skip the browser setup
If you need a clean page image rather than DOM links, ScreenshotNeo provides a website screenshot API. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, cookie banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card.
10. FAQ
Can I return the literal HTML attribute?
Yes. Evaluate link.getAttribute('href') instead of link.href. Use href when you want the browser-resolved absolute URL.
Why does a selector work locally but fail in CI?
The page may use a different locale, experiment, consent state, or viewport. Save rendered HTML or a screenshot in CI and compare the state.
Should I use XPath or CSS?
Use whichever expresses the page structure clearly. Both are documented approaches; neither is universally more reliable.
Is Pyppeteer the same as current Puppeteer?
No. Current Puppeteer documentation is related context, not proof of Pyppeteer feature parity. Verify behavior against your installed version.


