ScreenshotNeo

BlogComparisons

The Best Way to Scrape Website Data: Approaches Compared with Python

Choose Requests, Scrapy, or Selenium by page type and scale, with runnable Python examples, troubleshooting, and a hybrid workflow.

By the ScreenshotNeo team1 October 20267 min read

Short answer: use Requests with Beautiful Soup when the data is already in the initial HTML, Scrapy when you need a repeatable crawl across many URLs, and Selenium when a real browser must run JavaScript or perform interactions. Inspect one page first; the page type and crawl size should decide the tool.

Beautiful Soup parses HTML or XML; it does not fetch pages. Requests performs the fetch. Scrapy combines requests, scheduling, selectors, throttling, caching, and item pipelines for crawls. Selenium drives a browser through WebDriver, so it can execute JavaScript, click controls, submit forms, and keep browser state. See the Beautiful Soup docs, Scrapy overview, Scrapy requests and responses, and Selenium WebDriver.

Choose by page and job

Situation Start with Why Trade-off
One or a few static pages Requests + Beautiful Soup Small, inspectable request/parse pipeline You add retries, throttling, pagination, and storage
Many pages or domains Scrapy Scheduler, spiders, selectors, exports, caching, and pipelines are built in More project structure to learn and maintain
JavaScript-rendered or interactive pages Selenium Runs a real browser and can click, scroll, submit, and retain state Higher CPU/RAM use and more timing failures
Mostly HTTP plus a few difficult views Requests/API discovery + targeted Selenium Use direct HTTP where possible and a browser only for rendered steps Session and hand-off logic adds complexity

1. Inspect the target before writing a scraper

  1. Open View Source or fetch the URL with curl -I and a plain GET.
  2. Search the response for the field you need. If the text is present, it is probably static HTML.
  3. Use browser developer tools’ Network tab to see whether an XHR or fetch API returns structured data.
  4. Check pagination, authentication, rate limits, robots.txt, and terms before scheduling requests.
  5. Define an output schema and rules for missing values.

document.readyState == 'complete' only means the initial document finished loading. A single-page app can still be fetching and rendering data.

2. Static HTML: Requests and Beautiful Soup

python -m pip install requests beautifulsoup4 lxml
from urllib.parse import urljoin
import json, time, requests
from bs4 import BeautifulSoup

URL = 'https://example.com/news'
HEADERS = {'User-Agent': 'research-bot/1.0 (contact: you@example.com)'}
session = requests.Session()
for attempt in range(3):
    try:
        response = session.get(URL, headers=HEADERS, timeout=(10, 30))
        response.raise_for_status()
        break
    except requests.RequestException:
        if attempt == 2: raise
        time.sleep(2 ** attempt)

soup = BeautifulSoup(response.text, 'lxml')
items = []
for card in soup.select('article'):
    link = card.select_one('a[href]')
    heading = card.select_one('h2, h3')
    if not link or not heading: continue
    items.append({'title': heading.get_text(' ', strip=True), 'url': urljoin(response.url, link['href'])})
with open('items.json', 'w', encoding='utf-8') as f:
    json.dump(items, f, ensure_ascii=False, indent=2)
print(f'saved {len(items)} items')

Make a static scraper dependable

  • Use a Session for connection reuse and cookies.
  • Set connect and read timeouts.
  • Retry transient failures such as timeouts and selected 5xx responses, not 401, 403, or 404.
  • Normalize whitespace, resolve relative links, and validate required fields.
  • Throttle requests per host and persist progress.

3. Many URLs: Scrapy

python -m pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.com
import scrapy

class ProductsSpider(scrapy.Spider):
    name = 'products'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/products']
    custom_settings = {
        'ROBOTSTXT_OBEY': True,
        'DOWNLOAD_DELAY': 0.5,
        'AUTOTHROTTLE_ENABLED': True,
        'AUTOTHROTTLE_START_DELAY': 0.5,
        'AUTOTHROTTLE_MAX_DELAY': 10,
        'CONCURRENT_REQUESTS_PER_DOMAIN': 4,
        'FEEDS': {'products.jsonl': {'format': 'jsonlines', 'overwrite': True}},
    }
    def parse(self, response):
        for card in response.css('article.product'):
            yield {'name': card.css('h2::text').get(default='').strip(), 'price': card.css('.price::text').get(default='').strip(), 'url': response.urljoin(card.css('a::attr(href)').get())}
        next_url = response.css('a.next::attr(href)').get()
        if next_url: yield response.follow(next_url, callback=self.parse)
scrapy crawl products

Useful settings include ROBOTSTXT_OBEY, download delay, per-domain concurrency, AutoThrottle, retry settings, HTTP caching, cookies, and feed exports. Prefer stable attributes over generated class names, log rejected items, and checkpoint long crawls.

4. JavaScript and interaction: Selenium

python -m pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = Options()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,1200')
driver = webdriver.Chrome(options=options)
try:
    driver.get('https://example.com/dashboard')
    wait = WebDriverWait(driver, 30)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'main')))
    driver.execute_script('window.scrollTo(0, document.body.scrollHeight);')
    wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, '.results')))
    rows = [el.text for el in driver.find_elements(By.CSS_SELECTOR, '.result-row')]
    print(rows)
finally:
    driver.quit()

Use explicit conditions such as element_to_be_clickable before clicks. Switch into iframes before querying them and switch to new windows after links open tabs. Browser automation is not the first choice when an underlying JSON endpoint provides the same fields.

5. A hybrid workflow

  1. Identify the JSON request supplying the page.
  2. Replay it with Requests, including required headers, cookies, pagination, and bounded timeouts.
  3. Validate authorization and endpoint stability.
  4. Reserve Selenium for login, token acquisition, or views that cannot be obtained through HTTP.
  5. Store raw responses and a parser version for auditing.

6. Pagination, sessions, and data quality

  • Pagination: follow next links or cursors until absent, cap pages, and deduplicate canonical URLs.
  • Sessions: use one Requests session or controlled Selenium browser state; never log session cookies.
  • Encoding: preserve Unicode and use response headers where possible.
  • Missing data: emit null plus a reason instead of shifting columns.
  • Duplicates: key records by a stable ID or normalized URL.
  • Change detection: keep fixture pages and parser checks.

Read terms and robots.txt, identify your user agent, collect only what you need, and keep rates low. Authentication, paywalls, bot checks, and personal data require specific authorization. Robots.txt is an implementation signal, not a complete legal answer.

8. Troubleshooting

Symptom Cause Fix
No expected text JavaScript rendering Find the data API or use Selenium with an explicit wait
403 or 429 Rate, identity, or access policy Stop, check permission, reduce concurrency, and honor Retry-After
Empty selector result Markup changed, wrong frame, or shadow DOM Inspect the live DOM, switch frames, or use an approved API
Intermittent Selenium failures Race with asynchronous rendering Wait for a condition and capture browser logs
StaleElementReferenceException Framework re-rendered the node Locate the element again
Timeouts or memory growth Drivers, tabs, or queues remain open Use finally, quit drivers, and process batches
Wrong prices or dates Locale, hidden text, or formatting Parse machine-readable attributes and validate units

9. Performance, reliability, and cost

HTTP parsing is usually the lightest path: reuse connections, cache immutable responses, and write incrementally. Scrapy adds scheduling and back-pressure for large crawls. Selenium starts a browser per worker, so CPU, RAM, startup time, and synchronization dominate; use a small worker pool and target only pages that need it. Measure your own workload rather than assuming one library is universally faster.

Reliability comes from explicit timeouts, bounded retries with backoff, idempotent output, checkpointing, selector tests, and structured logs containing URL, status, duration, and parser version. Cost includes bandwidth, proxies or hosted browsers, storage, and engineering time.

Or skip the browser setup

If your deliverable is a clean visual capture rather than extracted text, ScreenshotNeo provides a single website screenshot API request. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. An MCP server lets Claude, Cursor, and other MCP clients take screenshots.

See the ScreenshotNeo API docs for full-page and element capture, dark mode, device presets, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, async webhooks, bulk capture, usage, and OpenAPI compatibility.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

The free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account.

FAQ

Is Beautiful Soup a scraper?

It is the parser in a fetch-and-parse pipeline. Requests retrieves the response.

Is Scrapy faster than Selenium?

They solve different workloads. Scrapy avoids browser overhead; Selenium is required for browser execution or interaction.

How do I scrape a single-page application?

Look for its JSON requests first. If no usable endpoint exists, use Selenium with explicit waits.

Should I use fixed sleeps?

Use condition-based waits where possible. Fixed delays are slower and less reliable.

How do I avoid scraping too aggressively?

Identify the client, obey applicable rules, throttle per host, cap concurrency, honor Retry-After, and stop when access is denied.