Scraping with Nodriver: Step-by-Step Tutorial with Examples
Learn nodriver web scraping with async Python, selectors, waits, sessions, screenshots, anti-bot limits, debugging, and production patterns.
Direct answer: Nodriver is an asynchronous Python browser automation and scraping library that talks directly to Chrome DevTools Protocol (CDP). Install the package and a Chromium-based browser, start it with uc.start(), navigate with browser.get(), wait for real page elements, then extract text or attributes with text lookup, CSS selectors, or XPath. This approach works well for JavaScript-rendered pages because extraction happens after the browser executes the page’s scripts.
The project describes itself as the official successor to Undetected-Chromedriver and uses the tagline “No more webdriver, no more selenium.” Those are maintainer descriptions, not a promise that every site will allow automated access. Anti-bot behavior remains site-specific.
What you will build
By the end, you will have a runnable scraper that:
- Creates an isolated browser session.
- Loads a JavaScript page and waits for a meaningful element.
- Extracts data with CSS selectors and text-aware lookup.
- Handles missing content, timeouts, and cleanup.
- Saves HTML and screenshots for debugging.
Nodriver supports Chromium, Chrome, Edge, and Brave. A compatible browser must already be installed; pip install nodriver does not install Chrome. See the official README and the PyPI package page for current compatibility and release information.
Install nodriver and a supported browser
1. Create a virtual environment
python -m venv .venv
source .venv/bin/activate
# Windows PowerShell:
# .venv\\Scripts\\Activate.ps1
python -m pip install -U pip nodriver
Current PyPI metadata requires Python 3.9 or newer and classifies nodriver as alpha. Verify the installed version before deploying version-sensitive code:
python --version
python -m pip show nodriver
2. Install Chromium, Chrome, Edge, or Brave
Install one Chromium-based browser through your operating system or container image. In a headless Linux environment, use nodriver’s headless mode or provide a display through Xvfb. Keep the browser version and your nodriver version under review when upgrading.
Minimal asynchronous scraper
This is the smallest useful pattern from the project documentation. It starts a browser, opens a page, retrieves the rendered HTML, prints it, and stops the browser.
import nodriver as uc
async def main():
browser = await uc.start()
try:
page = await browser.get('https://example.com')
html = await page.get_content()
print(html)
finally:
await browser.stop()
if __name__ == '__main__':
uc.loop().run_until_complete(main())
Save this as scrape.py and run python scrape.py. The finally block matters: it closes the browser when navigation or extraction raises an exception.
A complete JavaScript-site scraping example
The following example demonstrates a resilient workflow. It waits for the page’s content instead of sleeping for an arbitrary number of seconds, extracts cards, and writes both JSON and a screenshot.
import asyncio
import json
from pathlib import Path
import nodriver as uc
URL = 'https://example.com/products'
async def scrape_products():
browser = await uc.start()
try:
page = await browser.get(URL)
# Selector lookup retries while the page is loading.
try:
await page.select('main')
except Exception as exc:
await page.save_screenshot('debug-no-main.png')
html = await page.get_content()
Path('debug-no-main.html').write_text(html, encoding='utf-8')
raise RuntimeError('The main content did not appear') from exc
cards = await page.select_all('article.card')
products = []
for card in cards:
products.append({
'text': card.text,
'href': card.attrs.get('href'),
})
Path('products.json').write_text(
json.dumps(products, ensure_ascii=False, indent=2),
encoding='utf-8',
)
await page.save_screenshot('products.png')
return products
finally:
await browser.stop()
if __name__ == '__main__':
results = uc.loop().run_until_complete(scrape_products())
print(json.dumps(results, ensure_ascii=False, indent=2))
Replace the URL and selectors with the target site’s structure. A selector such as article.card is an example, not a universal selector.
Selectors and extraction methods
Text-aware lookup
Use visible text when the label is stable and meaningful:
button = await page.find('accept all', best_match=True)
if button:
await button.click()
items = await page.find_all('Product')
for item in items:
print(item.text)
Text matching is useful for buttons and headings, but it can break when a site changes wording or localizes content.
CSS selectors
cards = await page.select_all('article.card')
for card in cards:
print(card.text)
print(card.attrs.get('href'))
CSS is usually the clearest choice for repeated structures. Inspect the rendered DOM and choose stable classes, data attributes, or semantic elements instead of generated class names.
XPath
nodes = await page.xpath('//h2[contains(., "Price")]')
for node in nodes:
print(node.text)
XPath helps when you need a relationship that CSS cannot express, such as selecting an element based on nearby text.
Frames and iframes
The project documents iframe-aware lookup and a flat-mode connection. You can inspect frames explicitly:
frames = await page.get_frames()
for frame in frames:
print(frame)
Recent releases changed connection behavior. The 0.50.1 notes say that flat mode includes iframes in more operations and that find() includes iframe content. Test iframe-heavy projects after upgrading.
Waiting for dynamic content
JavaScript applications often render an empty shell first and populate it later. Wait for the state you actually need:
# Wait until the main content exists.
await page.select('main')
# Wait until a result label appears.
await page.find('Results', best_match=True)
# Then extract.
rows = await page.select_all('table tbody tr')
Selector calls retry for the duration of their timeout according to the project documentation. A fixed asyncio.sleep() can be useful for a known animation, but it is less reliable than waiting for a meaningful element. Always handle the case where the element never appears.
When a selector never appears
- Capture a screenshot and HTML immediately.
- Check whether the page redirected or returned an access-denied document.
- Confirm the selector belongs to the rendered DOM, not only the original response.
- Check whether the content is inside an iframe.
- Increase the wait only after confirming that the page is progressing.
Clicks, scrolling, and pagination
Interact with the page before extraction when content is behind a button or lazy-loaded below the fold:
more = await page.find('load more', best_match=True)
if more:
await more.click()
await page.select('article.card')
# Scroll to trigger lazy loading.
await page.evaluate('window.scrollTo(0, document.body.scrollHeight)')
await page.select('footer')
For pagination, extract the current page, click the next control, wait for a page-specific change, and stop when the control is missing. Use a maximum page count and deduplicate records so a broken next link cannot create an endless loop.
Cookies, profiles, and authentication
A fresh profile improves reproducibility and is cleaned up at exit by default. A persistent user_data_dir can preserve login state and local storage between runs, but it also changes privacy, isolation, and repeatability.
import nodriver as uc
async def main():
browser = await uc.start(user_data_dir='./browser-profile')
try:
page = await browser.get('https://example.com/account')
print(await page.get_content())
finally:
await browser.stop()
if __name__ == '__main__':
uc.loop().run_until_complete(main())
Do not commit the profile directory or credentials. Use a dedicated profile per account or workload, restrict filesystem permissions, and decide whether saved cookies are acceptable for your data-handling requirements. Nodriver also documents saving and loading cookies, local-storage access, and connecting to an existing Chrome debugging session.
HTML, screenshots, and debugging
Use HTML and screenshots as evidence when an extraction fails:
html = await page.get_content()
with open('page.html', 'w', encoding='utf-8') as f:
f.write(html)
await page.save_screenshot('page.png')
print(page)
The project documents tab.open_external_debugger() for inspection and element representations intended to make HTML debugging easier. Redact cookies, tokens, personal data, and other secrets before storing artifacts or sharing them.
Multiple tabs and windows
Open detail pages in new tabs when the list page must remain available. Keep references to the tabs you own, bring the required tab to the front before interacting, and close tabs when they are no longer needed. The README documents opening tabs or windows, bringing pages forward, reloading, and closing tabs.
Headless deployment
Local development is easiest with a visible browser. In CI, containers, or servers:
- Install a compatible Chromium binary in the image.
- Run headless mode when no display exists, or use Xvfb.
- Give the browser a writable temporary directory.
- Set CPU, memory, and process limits deliberately.
- Store screenshots and HTML only for the retention period you need.
Test the same browser flags and version in development and production. A page that works in an interactive desktop session can fail in a restricted container because of missing fonts, sandbox permissions, certificates, or shared memory.
Anti-bot systems, Cloudflare, and responsible scraping
Can nodriver bypass Cloudflare? There is no universal yes. The maintainers describe nodriver as designed for anti-bot resistance, but detection and access are probabilistic and site-specific. A bot check, CAPTCHA, login wall, or WAF can still block a run.
The documented tab.cf_verify() helper is a checkbox helper, not a general CAPTCHA-solving service. It works only outside expert mode, is currently English-only, and requires opencv-python. The README warns that expert mode disables web security and origin trials and makes you more detectable.
Respect robots directives, terms of service, authentication boundaries, rate limits, privacy obligations, and applicable law. Use an authorized account for protected data, identify your workload where required, and stop when a site explicitly disallows the activity.
Nodriver versus Selenium and Playwright
| Axis | Nodriver | What to evaluate in alternatives |
|---|---|---|
| Protocol | Direct Chrome DevTools Protocol communication; the project positions it as having no WebDriver dependency. | Selenium commonly uses WebDriver; Playwright uses its own browser automation protocol. |
| Programming model | Asynchronous Python APIs. | Compare async support, language choices, and how your job queue handles concurrency. |
| Browser lifecycle | Start a browser or connect to an existing debugging session; profiles can persist state. | Compare context isolation, profile reuse, and cleanup behavior. |
| Selectors and frames | Text, CSS, XPath, and documented iframe-aware operations. | Compare selector ergonomics and frame APIs for your target sites. |
| Debugging | Rendered HTML, screenshots, element representations, and external debugger support. | Compare trace, video, inspector, and artifact support. |
| Anti-bot constraints | Maintainers describe anti-bot resistance, with no universal access guarantee. | Every tool remains subject to the site’s controls and policies. |
Choose based on the browser, language, isolation model, debugging workflow, and site policies you must support. The official Nodriver README is the source for Nodriver’s side of these comparisons; it publishes no controlled benchmark for speed, detection rate, or CAPTCHA success.
Or skip the browser setup
If you need a clean image or PDF rather than extracted records, ScreenshotNeo provides a single GET request to capture a URL. Its consent step accepts cookie banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Only clean shots are billed: bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G 'https://api.screenshotneo.com/v1/shot' \\
-d access_key=YOUR_API_KEY \\
--data-urlencode url=https://stripe.com \\
-o shot.webp
Python
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await Bun.write('shot.webp', bytes); // In Node, write bytes with fs.promises.writeFile
ScreenshotNeo also supports PNG, JPEG, WebP, and PDF; full-page capture with lazy images loaded; CSS element capture; dark mode; device presets and custom viewports; retina scale; PDF paper, margins, landscape, and page ranges; HTML/CSS rendering; custom JavaScript and CSS; clicks; selector or network-idle waits; blocking ads, trackers, requests, or resource types; custom headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, usage reporting, and an OpenAPI specification. Common screenshot-API parameter names also work, which can simplify migration.
An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Pricing includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Yearly billing gives two months free, and every feature is on every plan.
Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card.
Performance, reliability, and cost notes
Performance
- Reuse one browser for a controlled batch instead of starting a process for every URL.
- Limit concurrent tabs to the memory and CPU available in your environment.
- Wait on specific selectors so fast pages do not pay an unnecessary fixed delay.
- Block optional media or third-party requests only when doing so does not change the data you need.
- Cache inputs and results where the site’s freshness requirements permit it.
No controlled benchmark for nodriver speed is published in the cited sources, so measure your own pages with realistic browser versions, network conditions, and concurrency.
Reliability
- Set a maximum run time and capture diagnostics on failure.
- Retry transient browser or network failures with bounded exponential backoff.
- Do not retry permanent responses such as a clear authorization failure without changing credentials or permissions.
- Deduplicate records when pagination or retries can repeat a page.
- Pin and review nodriver versions; the 0.50.1 flat-mode rewrite specifically asks large projects to test thoroughly.
Cost
Nodriver itself is an open-source Python package under the AGPL-3.0 license according to PyPI metadata, but you still pay for compute, browser processes, bandwidth, proxies, storage, and engineering time. Check that the license fits your distribution model. ScreenshotNeo charges only for clean shots, with the free and paid tiers described above; cache hits and failed or unusable captures are not billed.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
ModuleNotFoundError: nodriver |
The package is not installed in the active environment. | Activate the virtual environment and run python -m pip install -U nodriver. |
| Browser fails to start | No supported Chromium browser, missing dependencies, or an unusable display. | Install Chrome, Chromium, Edge, or Brave; use headless mode or Xvfb in a headless host; inspect container libraries and permissions. |
| Selector times out | The page is still loading, redirected, blocked, localized, or the selector changed. | Save HTML and a screenshot, inspect the rendered DOM, confirm the URL, then wait for a stable selector or handle the missing state. |
| HTML has no records | Data is rendered later, placed in an iframe, or loaded after scrolling. | Wait for a result element, inspect get_frames(), scroll to trigger lazy loading, and verify the network or application state. |
| Click has no effect | The visible text changed, an overlay intercepts the click, or the control is outside the active frame. | Use a stable selector, wait for the overlay to disappear, inspect frames, and capture a screenshot before and after the click. |
| Login disappears between runs | You are using a fresh temporary profile. | Use a dedicated persistent user_data_dir or explicitly save and load cookies, and protect the profile. |
| Cloudflare or CAPTCHA blocks the run | The site’s protection challenged the browser. | Do not assume a bypass. Use authorized access, respect the site’s policy, consider cf_verify() only within its documented limits, or stop the job. |
| Works locally but fails in CI | Missing fonts, display, shared memory, certificates, browser binary, or writable directories. | Reproduce the production image locally, install dependencies, enable headless/Xvfb, and retain redacted diagnostics. |
| ScreenshotNeo response is not a clean image | The page verdict indicates a bot check, blank page, timeout, failed load, or another unusable result. | Inspect X-Page-Verdict and X-Billed, then adjust waits, headers, cookies, viewport, or target-page assumptions. |
Production checklist
- Pin Python, nodriver, and browser versions.
- Use a dedicated profile and secret store for authenticated jobs.
- Wait for a semantic page state, not a guessed delay.
- Set timeouts, retry limits, and concurrency limits.
- Persist redacted HTML and screenshots for failed pages.
- Detect redirects, login pages, empty results, and bot checks.
- Respect site terms, robots directives, rate limits, and privacy rules.
- Re-test iframe and selector behavior after nodriver upgrades.
FAQ
Is nodriver a replacement for Selenium?
It is an alternative with a different protocol and lifecycle model: direct CDP communication and asynchronous Python APIs. Compare it against your browser, language, profile, debugging, and maintenance requirements.
Does nodriver install Chrome?
No. Install Chromium, Chrome, Edge, or Brave separately.
Which Python version does nodriver require?
Current PyPI metadata requires Python 3.9 or newer. Recheck the package page when publishing or upgrading.
Can I scrape a page that needs JavaScript?
Yes. Nodriver controls a real Chromium-based browser, so scripts can render content before you extract it. You still need to wait for the relevant page state.
Is nodriver guaranteed to avoid detection?
No. Anti-bot systems are site-specific, and the project publishes no universal detection or CAPTCHA-success guarantee.
When should I use a screenshot API instead?
Use a screenshot API when your output is an image or PDF and you want to avoid maintaining browser installation, profiles, waits, and capture infrastructure. ScreenshotNeo adds consent cleanup, verdict-aware billing, and an MCP server for AI agents.


