ScreenshotNeo

BlogHow-to

How to Scrape Website Content with Pyppeteer and Asyncio

Learn to scrape JavaScript-rendered pages with Pyppeteer and asyncio, extract text or HTML, control concurrency, and troubleshoot browser automation.

By the ScreenshotNeo team1 October 202610 min read

Direct answer: use Pyppeteer to launch Chromium, open a page with page.goto(), wait for the rendered content you need, then extract it with page.content() or page.evaluate(). Put the workflow in async def main() and start it with asyncio.run(main()). Pyppeteer is an unofficial Python port of Puppeteer for headless Chrome/Chromium automation; asyncio supplies Python’s async/await and task scheduling.

Pyppeteer follows Puppeteer’s browser model but is not an official Google or Python project, and its Python method names differ from JavaScript Puppeteer. Read the Pyppeteer documentation and API reference alongside the version you install.

1. Install Pyppeteer and prepare Chromium

Install the package in a virtual environment:

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install pyppeteer

On its first launch, Pyppeteer may download its bundled Chromium. The project documentation says the bundled browser is the best-supported choice and does not guarantee compatibility with every separately installed Chrome or Chromium version. Python requirements vary by release: older versioned documentation mentions Python 3.6+, while the current development README says Python >=3.8. Check the README and package metadata for the exact version you install.

2. Minimal scraper for rendered HTML

This complete program opens a URL, waits for navigation to finish, saves the rendered document, and always closes the browser:

import asyncio
from pathlib import Path

from pyppeteer import launch

URL = 'https://example.com'

async def main():
    browser = await launch(headless=True)
    try:
        page = await browser.newPage()
        await page.goto(URL, {'waitUntil': 'networkidle2', 'timeout': 60000})
        html = await page.content()
        Path('page.html').write_text(html, encoding='utf-8')
        print(f'Saved {len(html)} characters')
    finally:
        await browser.close()

if __name__ == '__main__':
    asyncio.run(main())

page.content() returns the complete HTML contents of the page, including the doctype. It represents the DOM after scripts have modified it, rather than the original bytes returned by an HTTP request.

3. Extract rendered text with page.evaluate()

When the target is a value in the rendered DOM, evaluate a JavaScript expression in the page:

import asyncio
from pyppeteer import launch

async def main():
    browser = await launch(headless=True)
    try:
        page = await browser.newPage()
        await page.goto('https://example.com', {'waitUntil': 'domcontentloaded'})
        text = await page.evaluate('document.body.textContent', force_expr=True)
        print(text.strip())
    finally:
        await browser.close()

asyncio.run(main())

force_expr=True tells Pyppeteer that the string is an expression. For a function, return a value explicitly:

title = await page.evaluate('''() => document.querySelector('h1')?.textContent || '' ''')
links = await page.evaluate('''() => Array.from(document.querySelectorAll('a')).map(a => ({
    text: a.textContent.trim(),
    href: a.href
}))''')

Use targeted extraction when you only need a field. It avoids parsing a large document later and makes missing data easier to detect.

4. Select elements with Pyppeteer’s Python API

JavaScript Puppeteer uses methods such as $; that character is not a valid Python identifier. Pyppeteer exposes Python-friendly methods such as querySelector() and querySelectorAll():

element = await page.querySelector('article h1')
if element is None:
    raise RuntimeError('article heading was not found')
heading = await page.evaluate('(el) => el.textContent', element)
print(heading.strip())

items = await page.querySelectorAll('article li')
for item in items:
    value = await page.evaluate('(el) => el.textContent', item)
    print(value.strip())

You can also extract attributes:

href = await page.evaluate("el => el.getAttribute('href')", element)

Selectors are evaluated against the current DOM. If a framework replaces the node, select it again after the update instead of retaining a stale element handle.

5. Wait for JavaScript-rendered data

Choose a wait that matches the page:

Wait Use it when Example
domcontentloaded Initial HTML is enough and background requests are irrelevant. {'waitUntil': 'domcontentloaded'}
load Images and subresources should finish loading. {'waitUntil': 'load'}
networkidle0 The page should have no active network connections. {'waitUntil': 'networkidle0'}
networkidle2 A practical choice for pages with a small amount of continuing traffic. {'waitUntil': 'networkidle2'}
Selector wait A known component marks readiness. await page.waitForSelector('.results')
Fixed delay A short animation or delayed request has no reliable selector. await asyncio.sleep(2)
await page.goto('https://example.com/app', {
    'waitUntil': 'domcontentloaded',
    'timeout': 60000,
})
await page.waitForSelector('.product-card', {'timeout': 30000})

Prefer a selector or a page-specific readiness condition over a long fixed sleep. A network-idle condition can never arrive on pages with analytics, polling, or streaming connections.

6. Interact before extracting

Click buttons, fill forms, or scroll lazy content before reading the DOM:

await page.click('button.load-more')
await page.waitForSelector('.new-results')

await page.type('input[name="q"]', 'pyppeteer')
await page.click('button[type="submit"]')
await page.waitForSelector('.search-results')

When a click causes navigation, register the navigation wait and click together. This avoids a race in which navigation starts before your code begins waiting:

await asyncio.gather(
    page.waitForNavigation({'waitUntil': 'networkidle2'}),
    page.click('a.next-page'),
)

For infinite-scroll pages, scroll incrementally and stop when the expected selector appears or the document height stops changing:

previous_height = 0
for _ in range(20):
    height = await page.evaluate('document.body.scrollHeight', force_expr=True)
    if height == previous_height:
        break
    previous_height = height
    await page.evaluate('window.scrollTo(0, document.body.scrollHeight)', force_expr=True)
    await asyncio.sleep(1)

7. A reusable scraper with structured output

import asyncio
import json
from pyppeteer import launch

async def scrape(url: str) -> dict:
    browser = await launch(headless=True)
    try:
        page = await browser.newPage()
        await page.setUserAgent(
            'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 '
            '(KHTML, like Gecko) Chrome Safari/537.36'
        )
        response = await page.goto(
            url,
            {'waitUntil': 'networkidle2', 'timeout': 60000},
        )
        await page.waitForSelector('body', {'timeout': 10000})
        result = await page.evaluate('''() => ({
            title: document.title,
            text: document.body.innerText,
            links: Array.from(document.querySelectorAll('a')).map(a => ({
                text: a.innerText.trim(),
                href: a.href
            }))
        })''')
        return {
            'url': page.url,
            'status': response.status if response else None,
            **result,
        }
    finally:
        await browser.close()

async def main():
    data = await scrape('https://example.com')
    print(json.dumps(data, ensure_ascii=False, indent=2))

if __name__ == '__main__':
    asyncio.run(main())

Keep the browser lifetime outside the per-field extraction logic. If you scrape several pages, reuse one browser process and create or close pages deliberately.

8. Scrape many URLs with bounded asyncio concurrency

Asyncio tasks can overlap browser I/O, but launching an unlimited number of tabs can exhaust memory, file descriptors, CPU, or the target site’s capacity. Use an asyncio.Semaphore or a fixed worker pool. Python documents a semaphore as a counter that blocks when its value reaches zero.

import asyncio
from pyppeteer import launch

URLS = [
    'https://example.com/one',
    'https://example.com/two',
    'https://example.com/three',
]

async def fetch(browser, url, limit):
    async with limit:
        page = await browser.newPage()
        try:
            await page.goto(url, {'waitUntil': 'domcontentloaded', 'timeout': 60000})
            await page.waitForSelector('body', {'timeout': 15000})
            return url, await page.evaluate('document.body.innerText', force_expr=True)
        except Exception as exc:
            return url, f'ERROR: {exc}'
        finally:
            await page.close()

async def main():
    browser = await launch(headless=True)
    try:
        limit = asyncio.Semaphore(3)
        results = await asyncio.gather(
            *(fetch(browser, url, limit) for url in URLS)
        )
        for url, text in results:
            print(url, text[:200])
    finally:
        await browser.close()

asyncio.run(main())

The correct limit depends on page size, JavaScript workload, available memory, and the site’s rules. Increase it gradually while watching failures and resource use. Async concurrency does not override robots policies, terms, authentication controls, or applicable law; only collect content you are allowed to access.

9. Configure pages for difficult sites

Viewport, locale, timezone, and JavaScript

await page.setViewport({'width': 1440, 'height': 900, 'deviceScaleFactor': 1})
await page.setJavaScriptEnabled(True)
await page.emulate({'viewport': {'width': 1280, 'height': 800}, 'userAgent': 'Mozilla/5.0'})

Use the browser’s supported emulation settings for the version installed. A different viewport can change responsive markup and the data exposed in the DOM.

Headers, cookies, and authentication

await page.setExtraHTTPHeaders({'Accept-Language': 'en-US,en;q=0.9'})
await page.setCookie({
    'name': 'session',
    'value': 'REDACTED',
    'domain': 'example.com',
    'path': '/',
})
await page.goto('https://example.com/account', {'waitUntil': 'networkidle2'})

Do not print session cookies or authorization headers. For a login flow, navigate to the login page, fill fields, submit, wait for a post-login selector, then scrape the protected page.

Block unwanted resources

await page.setRequestInterception(True)
async def handle_request(request):
    if request.resourceType in {'image', 'font', 'media'}:
        await request.abort()
    else:
        await request.continue_()
page.on('request', lambda request: asyncio.ensure_future(handle_request(request)))

Blocking resources can reduce local work, but it can also break pages whose content depends on an image, font, media request, or third-party script. Validate extraction after enabling it.

10. Static HTTP versus browser scraping

Approach Choose it when Trade-off
HTTP client and HTML parser The server response already contains the data and no interaction is needed. Simpler and lighter, but it does not execute page JavaScript.
Pyppeteer Content appears after JavaScript, scrolling, clicking, login, or other browser behavior. More setup and resource use because a browser runs.
page.content() You need the complete rendered document. Returns more data than a focused extractor needs.
Targeted evaluation You need selected text, attributes, or JSON-like values. Selectors and page-specific JavaScript must be maintained.

11. Troubleshooting common errors

Symptom Likely cause Fix
ERR何 or Chromium executable missing The first-run browser download did not complete, or the cache is unavailable. Run a small launch program again with network access, check the Pyppeteer cache path, or provide an explicit supported executable path.
TimeoutError during goto() The page keeps connections open, is slow, or is blocked. Use domcontentloaded, raise the timeout for this page, then wait for a specific selector. Record the final URL and response status.
Empty text or missing selector Extraction ran before the component rendered, the selector changed, or content is inside an iframe. Wait for the selector, inspect await page.content(), verify the selector in DevTools, and switch to the frame that owns the element.
Click hangs while navigation occurs A navigation wait was registered after the click. Use asyncio.gather(page.waitForNavigation(...), page.click(...)).
Works locally but fails in a container Sandbox, shared-memory, missing libraries, or a different Chromium revision. Use the container’s documented browser dependencies, provide only the required launch flags for that environment, and keep the bundled revision aligned with Pyppeteer.
CAPTCHA or bot-check page The target has challenged the browser. Do not attempt to bypass access controls. Respect the site’s terms, reduce load, or request an authorized data interface.
Browser or tab memory grows Pages are not closed, tasks are unbounded, or long-lived pages retain state. Close each page in finally, cap concurrency, and restart the browser between controlled batches if necessary.

The first row intentionally has no invented error code: copy the exact exception from your environment because Chromium download and executable errors differ by package revision and operating system.

12. Reliability, performance, and cost considerations

  • Reliability: close pages and browsers in finally, set explicit navigation and selector timeouts, capture status and final URLs, and retry only transient failures with backoff.
  • Performance: reuse one browser, limit simultaneous pages, block resources only after verifying correctness, and extract the smallest value needed.
  • Reproducibility: pin your Pyppeteer version, record the Chromium revision, and keep viewport, locale, user agent, and wait conditions in configuration.
  • Operational cost: browser CPU and memory are local or server expenses. There are no universal throughput figures in the Pyppeteer documentation; measure your own pages and concurrency limit.
  • Compliance: authentication, rate limits, robots directives, terms, and privacy obligations still apply to browser automation.

13. Or skip the browser setup

If you need a clean screenshot or PDF rather than a custom DOM data pipeline, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture with lazy images loaded, element capture by CSS selector, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture for up to 100 URLs per call, usage reporting, and an OpenAPI specification. Every feature is on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account with 1,000 screenshots each month and no card required.

14. FAQ

Is Pyppeteer official?

No. Its documentation describes it as an unofficial Python port intended to be similar to Puppeteer, with differences.

How do I get only visible text?

Evaluate document.body.innerText. Use textContent when you need text from hidden nodes as well.

Should I launch one browser per URL?

Usually no. Reuse one browser and create a bounded number of pages, closing each page after extraction.

Why does network idle never happen?

Analytics, polling, streaming, or long-lived connections can keep the network active. Wait for a page-specific selector or use domcontentloaded followed by a targeted wait.

Can Pyppeteer bypass a CAPTCHA?

You should not bypass bot checks or other access controls. Obtain permission or use an authorized API.