ScreenshotNeo

BlogHow-to

How to Extract Data From Websites Using Selenium and Python

Learn a reliable Selenium and Python workflow for dynamic websites, including waits, selectors, pagination, CSV export, errors, and production practices.

By the ScreenshotNeo team1 October 202611 min read

Short answer: install Selenium, start a browser with webdriver.Chrome(), navigate with get(), wait for the data you need, locate elements with stable selectors, extract text or attributes, validate the records, and always call driver.quit(). A page-load event does not mean JavaScript-rendered data is ready, so explicit waits are the key part of reliable extraction.

This guide shows a complete Python workflow for dynamic sites, including Selenium Manager, selectors, waits, clicks, scrolling, pagination, lazy loading, CSV output, retries, diagnostics, and production considerations.

1. Install Selenium and choose a browser

Selenium’s current Python API supports Python 3.10 and newer. Install or upgrade it in a virtual environment:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
# .venv\\Scripts\\Activate.ps1
python -m pip install -U selenium

The official installation documentation currently shows Selenium 4.49.0 as an example requirement; check PyPI when pinning a version. Selenium supports Chrome, Edge, Firefox, Safari, WebKitGTK, and WPEWebKit. Chrome is used below because it is widely available.

Do you still need ChromeDriver?

Usually, no. Selenium Manager ships with Selenium and can discover, download, and cache compatible drivers and, in supported cases, browsers. webdriver.Chrome() is therefore the normal starting point. If your environment requires a controlled binary, pass an explicit Service path or configure the driver through your deployment environment.

from selenium import webdriver

# Selenium Manager handles the driver in the normal case.
driver = webdriver.Chrome()
driver.quit()

See the Selenium Manager documentation for supported setup and controlled-driver options.

2. A complete extraction script that writes CSV

The following example extracts product cards from a dynamic page. Replace the URL and selectors with those from the site you are allowed to collect from.

from __future__ import annotations

import csv
import logging
import time
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
from urllib.parse import urljoin

from selenium import webdriver
from selenium.common.exceptions import TimeoutException, WebDriverException
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

URL = "https://example.com/products"
CARD_SELECTOR = "article.product"
NAME_SELECTOR = ".product-name"
PRICE_SELECTOR = ".price"
LINK_SELECTOR = "a"
WAIT_SECONDS = 15

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

@dataclass
class Product:
    name: str
    price: str
    url: str
    source_url: str
    retrieved_at: str


def clean(value: str | None) -> str:
    return " ".join((value or "").split())


def extract_products(driver: webdriver.Chrome, page_url: str) -> list[Product]:
    wait = WebDriverWait(driver, WAIT_SECONDS)
    driver.get(page_url)

    cards = wait.until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, CARD_SELECTOR))
    )
    retrieved_at = datetime.now(timezone.utc).isoformat()
    rows: list[Product] = []

    for card in cards:
        name = clean(card.find_element(By.CSS_SELECTOR, NAME_SELECTOR).text)
        price = clean(card.find_element(By.CSS_SELECTOR, PRICE_SELECTOR).text)
        link = card.find_element(By.CSS_SELECTOR, LINK_SELECTOR).get_attribute("href")
        rows.append(Product(
            name=name,
            price=price,
            url=urljoin(page_url, link or ""),
            source_url=page_url,
            retrieved_at=retrieved_at,
        ))

    if not rows:
        raise ValueError(f"No products found at {page_url}; check selectors or page state")
    return rows


def main() -> None:
    options = webdriver.ChromeOptions()
    # Uncomment for a server without a display:
    # options.add_argument("--headless=new")
    options.add_argument("--window-size=1440,1000")

    driver = webdriver.Chrome(options=options)
    try:
        rows = extract_products(driver, URL)
        with open("products.csv", "w", newline="", encoding="utf-8") as output:
            writer = csv.DictWriter(output, fieldnames=asdict(rows[0]).keys())
            writer.writeheader()
            writer.writerows(asdict(row) for row in rows)
        logging.info("Wrote %d records", len(rows))
    except (TimeoutException, WebDriverException, ValueError) as exc:
        logging.exception("Extraction failed: %s", exc)
        driver.save_screenshot("diagnostic.png")
        raise
    finally:
        driver.quit()


if __name__ == "__main__":
    main()

get() waits for the page-load event, while the explicit wait checks for the actual product cards. Selenium documents this navigation and cleanup lifecycle in its WebDriver documentation.

3. Wait for application data, not just page load

The browser’s readyState covers the document and declared assets. JavaScript can still fetch and render records afterward. Selenium’s waiting strategies documentation describes this race: the next command can run before the element exists.

Condition Use it when
presence_of_element_located The element must exist in the DOM, even if hidden.
visibility_of_element_located The element must be displayed and have usable dimensions.
element_to_be_clickable You are about to click an enabled, visible control.
presence_of_all_elements_located A list of cards, rows, or links must exist.
text_to_be_present_in_element A status or result container must contain expected text.
staleness_of After pagination, old content must be replaced.
frame_to_be_available_and_switch_to_it The target element is inside an iframe.
wait = WebDriverWait(driver, 20, poll_frequency=0.5)
results = wait.until(
    EC.visibility_of_all_elements_located((By.CSS_SELECTOR, "div.result"))
)

# A custom condition can wait for a non-empty result count.
def at_least_five_results(d):
    return len(d.find_elements(By.CSS_SELECTOR, "div.result")) >= 5

wait.until(at_least_five_results)

An implicit wait applies to element lookups for the driver’s lifetime. Prefer explicit waits for dynamic extraction, and avoid combining a long implicit wait with explicit waits because their timing interactions are difficult to predict. A fixed time.sleep() can be useful for a known animation, but it should not be your primary synchronization method.

4. Find elements with selectors that survive redesigns

find_element returns the first match; find_elements returns a list, including an empty list when nothing matches. Supported strategies include ID, name, CSS selector, XPath, link text, partial link text, tag name, and class name.

from selenium.webdriver.common.by import By

one = driver.find_element(By.ID, "main-content")
name = driver.find_element(By.NAME, "q")
rows = driver.find_elements(By.CSS_SELECTOR, "table tbody tr")
images = driver.find_elements(By.TAG_NAME, "img")
price = driver.find_element(By.XPATH, "//span[contains(@class, 'price')]")

visible_text = one.text
href = driver.find_element(By.CSS_SELECTOR, "a.product").get_attribute("href")
data_id = one.get_attribute("data-id")

Prefer stable IDs, data-* attributes, semantic classes, and meaningful container relationships. Keep selectors together in configuration so a redesign changes fewer lines. Use XPath when a text relationship or nontrivial structure requires it, but avoid selectors built from generated class names or fragile positions such as div:nth-child(7).

5. Click, scroll, and extract lazy-loaded content

Click a control

button = WebDriverWait(driver, 15).until(
    EC.element_to_be_clickable((By.CSS_SELECTOR, "button.load-more"))
)
driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", button)
button.click()
WebDriverWait(driver, 15).until(
    EC.text_to_be_present_in_element((By.CSS_SELECTOR, ".status"), "loaded")
)

Scroll until no more content appears

from selenium.common.exceptions import TimeoutException

wait = WebDriverWait(driver, 15)
for _ in range(20):
    before = len(driver.find_elements(By.CSS_SELECTOR, "article.product"))
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    try:
        wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, "article.product")) > before)
    except TimeoutException:
        break

Bound the number of scrolls and stop when the count, sentinel element, or end-of-results message indicates completion. Do not assume a single scroll loads everything.

6. Paginate without duplicating or losing records

from selenium.common.exceptions import TimeoutException, StaleElementReferenceException

all_rows = []
seen_urls = set()

while True:
    wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "article.product")))
    cards = driver.find_elements(By.CSS_SELECTOR, "article.product")
    for card in cards:
        link = card.find_element(By.CSS_SELECTOR, "a").get_attribute("href")
        if link and link not in seen_urls:
            seen_urls.add(link)
            all_rows.append({"name": clean(card.text), "url": link})

    old_first = cards[0] if cards else None
    next_buttons = driver.find_elements(By.CSS_SELECTOR, "a[rel='next'], button.next")
    if not next_buttons or not next_buttons[0].is_enabled():
        break
    next_buttons[0].click()
    if old_first is not None:
        try:
            wait.until(EC.staleness_of(old_first))
        except TimeoutException:
            wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "article.product")))

Some applications update rows in place instead of replacing the DOM. In that case, wait for a page number, URL change, result count, or a changed first-record key. Deduplicate with a canonical URL or site ID and record the source page for every row.

7. Frames, shadow DOM, authentication, and attributes

iframes

frame = WebDriverWait(driver, 15).until(
    EC.frame_to_be_available_and_switch_to_it((By.CSS_SELECTOR, "iframe.results"))
)
value = driver.find_element(By.CSS_SELECTOR, ".inside-frame").text
driver.switch_to.default_content()

Shadow DOM

For open shadow roots, retrieve the shadow root and query inside it. Closed shadow roots cannot be queried through normal WebDriver DOM APIs.

host = driver.find_element(By.CSS_SELECTOR, "product-widget")
root = host.shadow_root
name = root.find_element(By.CSS_SELECTOR, ".name").text

Login and session state

Use an account only when you have permission. Wait for the post-login condition, keep credentials out of source control, and do not print cookies or tokens in logs. Selenium can add cookies with driver.add_cookie() after first visiting the cookie’s domain.

Text versus attributes

Use .text for rendered text. Use get_attribute() for links, image URLs, IDs, ARIA values, prices stored in attributes, and machine-readable metadata. Normalize whitespace, numbers, dates, and currencies before writing output.

8. Validation, retries, and diagnostics

Never silently write an empty CSV. Validate required fields and expected counts, log the URL and selector that failed, and save a screenshot or HTML sample when policy permits. Retry only transient navigation or network failures, with a bounded count and backoff; selector failures usually require a code change, not another retry.

import time
from selenium.common.exceptions import WebDriverException

for attempt in range(3):
    try:
        driver.get(URL)
        break
    except WebDriverException:
        if attempt == 2:
            raise
        time.sleep(2 ** attempt)

Keep raw HTML or diagnostic screenshots only when storage is justified and site policy allows it. Store retrieval timestamps and source URLs so downstream users can trace each record.

9. Common errors and fixes

Error Likely cause Fix
TimeoutException Wrong selector, slow API, consent dialog, or blocked page. Inspect the rendered DOM, wait for a meaningful condition, handle the dialog, and check the URL and response state.
NoSuchElementException Element is inside an iframe, not rendered yet, or selector changed. Switch to the frame and use an explicit wait with a stable selector.
StaleElementReferenceException The framework replaced the node after you located it. Wait for staleness, then locate the element again; do not reuse the old reference.
ElementClickInterceptedException Overlay, cookie banner, or another element covers the control. Handle the overlay, wait for clickability, scroll it into view, or use the site’s accessible control.
Empty list Data loads later, selector matches a template, or bot protection returned another page. Log the title and URL, save a diagnostic screenshot, wait for data, and detect challenge pages.
Driver or browser mismatch Old manual driver or unsupported browser installation. Upgrade Selenium and use Selenium Manager, or explicitly provide a compatible controlled driver.
Headless differs from headed mode Viewport, timing, fonts, or anti-automation behavior differs. Set a window size, use current headless mode, and compare a diagnostic screenshot.
CSV has broken characters Wrong encoding or newline handling. Open with encoding="utf-8" and newline="".

10. Performance, reliability, and cost

  • Browser cost: Selenium starts a full browser, so it generally uses more CPU, memory, and startup time than an HTTP client and HTML parser. Use direct requests when the needed data is already in the HTTP response.
  • Reduce work: reuse one driver for a bounded batch, avoid unnecessary screenshots, wait for specific conditions, and collect only required fields.
  • Bound everything: set page and wait timeouts, maximum pages, maximum scrolls, retry counts, and an overall job deadline.
  • Reliability: validate schemas, detect challenge or blank pages, deduplicate stable keys, and close every driver in finally.
  • Concurrency: each independent browser consumes resources. Start with low concurrency, measure your own workload, and respect the site’s rate limits.
  • Policy: review terms of service, robots directives, authentication rules, copyright, privacy obligations, and applicable law before collecting data.

There is no universal Selenium speed or success-rate number for every site. The right approach depends on JavaScript complexity, browser interactions, access controls, and deployment resources.

11. When Selenium is the right tool

Selenium is strongest when the data appears only after JavaScript, requires clicks or scrolling, depends on login state, or is rendered by browser APIs. A direct HTTP client and parser are usually simpler when the required data is already present in the server response. This is a technical comparison, not a benchmark claim.

12. Or skip the browser setup

If your goal is a clean screenshot or PDF rather than structured field extraction, ScreenshotNeo provides a single GET request to render a URL. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Read the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

13. FAQ

Can Selenium scrape a site that renders data after load?

Yes. Navigate first, then wait for the specific result container, row count, text, or other condition that proves the data exists.

Should I use XPath or CSS selectors?

Use the most stable selector available. CSS is often concise; XPath is useful for text relationships and complex structure.

Why is my script reliable locally but not on a server?

Check browser availability, headless viewport size, fonts, network access, timing, authentication, and bot or challenge responses. Save a diagnostic screenshot and page title in the server environment.

How do I prevent duplicate rows?

Choose a stable key such as a canonical URL or site ID, keep a set of seen keys, and record the source page for each row.

Do I need Selenium for every scraping job?

No. Use Selenium when browser execution or interaction is required. Use an HTTP client and parser when the response already contains the data you need.