ScreenshotNeo

BlogHow-to

How to Capture Relevant Webpage Content With Selenium and Python

Extract the exact article, section, or iframe content you need with Selenium waits, stable selectors, scrolling, and robust error handling.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: navigate with driver.get(), wait for the smallest DOM container that represents the content you need, locate it with a stable selector, and read element.text or selected attributes. Always handle JavaScript-rendered content, frames, scrolling, and cleanup explicitly.

Selenium is useful when the content appears only after JavaScript runs or after an interaction. If the required HTML is already in the initial HTTP response, a direct HTTP client and parser will usually be simpler and faster. The examples below focus on rendered pages.

1. Install Selenium and a browser driver

Install the Python package:

python -m pip install selenium

Recent Selenium releases can manage compatible browser drivers automatically when a supported browser is installed. Otherwise install Chrome, Firefox, or another supported browser and configure its driver according to the official WebDriver documentation.

2. Extract one relevant content container

Start with a semantic boundary such as article, main, a stable ID, or a data attribute. Waiting for the container prevents you from extracting an empty shell before the page’s application has rendered its content.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "https://example.com/article"
driver = webdriver.Chrome()

try:
    driver.get(url)
    wait = WebDriverWait(driver, 15)

    article = wait.until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "article"))
    )

    print(article.text)
    canonical = article.get_attribute("data-canonical-url")
    print("Canonical data attribute:", canonical)
finally:
    driver.quit()

driver.get() waits for the page’s onload event, but AJAX content can still be incomplete. The explicit wait above is tied to the content boundary you actually intend to capture. See Selenium’s getting-started guide and waits guide.

3. Choose selectors that survive redesigns

Prefer selectors in this order:

  1. A stable ID such as #article-body.
  2. A semantic element such as article or main.
  3. A purposeful data attribute such as [data-testid="article-body"].
  4. A short class selector whose meaning is clear, such as .article-body.
  5. XPath only when CSS cannot express the relationship you need.

Avoid deeply nested positional selectors such as div:nth-child(3) > div:nth-child(2). They tend to break when an advertisement, recommendation, or new wrapper is inserted. Selenium’s locating-elements documentation covers CSS, XPath, IDs, class names, tags, and link text.

# First matching element
article = driver.find_element(By.CSS_SELECTOR, "main article")

# All matching cards
cards = driver.find_elements(By.CSS_SELECTOR, "main article, [role='main'] article")
for card in cards:
    print(card.text)

# A specific result by data attribute
result = driver.find_element(By.CSS_SELECTOR, "[data-result-id='42']")

find_element raises NoSuchElementException when nothing matches. find_elements returns an empty list, so check the result before publishing or storing extracted data.

4. Wait for meaningful content, not just an element

An element can exist before its text is populated. Use an expected condition that describes readiness:

from selenium.webdriver.support import expected_conditions as EC

wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "main article")))
wait.until(EC.text_to_be_present_in_element(
    (By.ID, "results"),
    "Published"
))

Presence means the node exists in the DOM; visibility means it can be seen; text conditions verify that useful content has arrived. Explicit waits poll for a bounded period (the documented default polling interval is 500 milliseconds) and raise TimeoutException when the condition never succeeds. Avoid using time.sleep() as your only synchronization method: a fixed delay is either wasteful or unreliable.

Wait for a count to increase

from selenium.common.exceptions import TimeoutException

cards_selector = "article.result"
initial_count = len(driver.find_elements(By.CSS_SELECTOR, cards_selector))

try:
    wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, cards_selector)) > initial_count)
except TimeoutException:
    print("No additional results appeared")

Wait for a loading indicator to disappear

wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, ".loading")))

5. Extract visible text, attributes, and live HTML

Use element.text for rendered, visible text. Use get_attribute() for links, labels, dates, data attributes, and other DOM values.

article = driver.find_element(By.CSS_SELECTOR, "article")

record = {
    "text": article.text,
    "url": article.get_attribute("data-canonical-url"),
    "aria_label": article.get_attribute("aria-label"),
}

links = []
for link in article.find_elements(By.CSS_SELECTOR, "a[href]"):
    links.append({
        "text": link.text,
        "href": link.get_attribute("href"),
    })

print(record)
print(links)

When another parser needs the current DOM, retrieve the selected node rather than dumping the whole page:

html = driver.execute_script("return arguments[0].outerHTML;", article)

canonical = driver.execute_script(
    "return document.querySelector('link[rel=canonical]')?.href;"
)
print(canonical)

driver.page_source is useful for diagnostics, but selected-element extraction avoids navigation, cookie banners, sidebars, and footer text that do not belong to the target content.

6. Extract content inside an iframe

An iframe has a separate document. Locate the frame, switch into it, extract the content, then return to the top-level document in a finally block.

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

wait = WebDriverWait(driver, 15)
frame = wait.until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "iframe"))
)

driver.switch_to.frame(frame)
try:
    body = wait.until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "article, main"))
    )
    text = body.text
finally:
    driver.switch_to.default_content()

print(text)

If the frame is replaced during loading, reacquire it before switching. Cross-origin policy does not prevent Selenium from interacting with a frame as a browsing context, but selectors must be evaluated after the switch.

7. Handle infinite scroll and lazy-loaded content

Infinite-scroll pages require a bounded loop. Scroll in steps, wait for a measurable change, and stop after a maximum number of rounds or when the page reports no more results.

from selenium.common.exceptions import TimeoutException

item_selector = "article.result"
max_rounds = 20
previous_count = 0

for _ in range(max_rounds):
    current_count = len(driver.find_elements(By.CSS_SELECTOR, item_selector))
    if current_count == previous_count:
        try:
            wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, item_selector)) > current_count)
        except TimeoutException:
            break

    previous_count = len(driver.find_elements(By.CSS_SELECTOR, item_selector))
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")

    try:
        wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, item_selector)) > previous_count)
    except TimeoutException:
        break

items = driver.find_elements(By.CSS_SELECTOR, item_selector)
text = [item.text for item in items]

Some sites expose a “load more” button instead. Wait for it to be clickable, click it, and wait for the item count to increase. Set a maximum page count or item count so a broken endpoint cannot create an unbounded job.

8. Use JavaScript-rendered state deliberately

For tabs, accordions, or menus, perform the interaction before extraction. Reacquire elements after an interaction that replaces the DOM.

tab = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button[data-tab='details']")))
tab.click()

panel = wait.until(
    EC.visibility_of_element_located((By.CSS_SELECTOR, "section[data-panel='details']"))
)
print(panel.text)

A StaleElementReferenceException means the saved node no longer belongs to the current DOM. Locate it again after navigation, filtering, or client-side rerendering.

9. A reusable extraction function

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException


def capture_content(url, selector, timeout=15):
    driver = webdriver.Chrome()
    driver.set_page_load_timeout(timeout)
    driver.set_script_timeout(timeout)
    try:
        driver.get(url)
        wait = WebDriverWait(driver, timeout)
        element = wait.until(
            EC.visibility_of_element_located((By.CSS_SELECTOR, selector))
        )
        text = element.text.strip()
        if not text:
            raise ValueError(f"Selector matched but produced no visible text: {selector}")
        return {
            "url": driver.current_url,
            "title": driver.title,
            "text": text,
            "html": driver.execute_script(
                "return arguments[0].outerHTML;", element
            ),
        }
    except (TimeoutException, NoSuchElementException) as exc:
        raise RuntimeError(f"Could not capture {selector} from {url}") from exc
    finally:
        driver.quit()


if __name__ == "__main__":
    result = capture_content(
        "https://example.com/article",
        "article",
    )
    print(result["title"])
    print(result["text"])

10. Troubleshooting

Symptom Cause Fix
TimeoutException The selector is wrong, the content is slow, or a page variant was served. Inspect the current DOM, verify the selector, wait for a meaningful condition, and log the URL and selector.
NoSuchElementException The node does not exist in the current document or frame. Check the selector and switch into the correct iframe before locating it.
Empty element.text The node is a shell, hidden, or populated later by JavaScript. Wait for visible text, a result count, or a loading indicator to disappear.
StaleElementReferenceException A rerender replaced the saved element. Reacquire the element after the DOM-changing action.
Only part of an infinite list appears Scrolling happened before the next batch loaded. Wait for item count growth after each scroll and cap the loop.
Unexpected navigation or consent text A banner, redirect, login wall, or bot check changed the page. Record driver.current_url, inspect page state, and handle the site’s required interaction explicitly.
Browser process remains after failure Cleanup was skipped. Put driver.quit() in finally.

11. Performance, reliability, and cost considerations

  • Scope the DOM: extracting one container is cheaper and cleaner than parsing page_source for the entire document.
  • Use explicit waits: short, condition-based waits finish quickly on fast pages while remaining bounded on slow pages.
  • Reuse a session carefully: a single driver can process several pages, but clear state between jobs and always quit it at the end.
  • Limit work: cap scroll rounds, item counts, navigation timeouts, and retries.
  • Make failures observable: log the URL, selector, current URL, title, exception type, and a diagnostic screenshot or HTML snapshot when permitted.
  • Plan for markup changes: keep selectors in configuration or small functions so a redesign does not require rewriting the whole pipeline.
  • Scale deliberately: local WebDriver suits occasional jobs. Parallel or long-running workloads may need a remote browser service, queue, and per-job isolation.
  • Control access: respect authentication, robots policies, terms, rate limits, and privacy requirements for the sites you process.

12. Or skip the browser setup

If you need a rendered screenshot or PDF rather than extracted text, ScreenshotNeo provides a single request endpoint. It handles the browser session for you and supports full-page shots, an element selected by CSS, custom JavaScript and CSS, waits, headers, cookies, user agents, geolocation, and PDF options.

cURL (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/article -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/article"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/article' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in X-Page-Verdict and X-Billed headers. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. FAQ

Does Selenium wait for all AJAX requests?

No. driver.get() waits for onload; use an explicit condition tied to the content or state you need.

Should I use implicit or explicit waits?

Prefer explicit waits for critical content. A global implicit wait can make unrelated lookups slow and can obscure which condition was not met.

How do I extract only links from an article?

Locate the article first, then call article.find_elements(By.CSS_SELECTOR, "a[href]") and read each link’s text and href.

Why does page source differ from what I see?

The live DOM may have been changed by JavaScript after the initial response. Use execute_script or the selected element’s properties after the page reaches the required state.

When should I avoid Selenium?

Use a direct HTTP client and HTML parser when the needed content is present in the server response and no browser interaction is required.