The Best Way to Scrape Website Data: Approaches Compared with Python
Choose Requests, Scrapy, or Selenium by page type and scale, with runnable Python examples, troubleshooting, and a hybrid workflow.
Short answer: use Requests with Beautiful Soup when the data is already in the initial HTML, Scrapy when you need a repeatable crawl across many URLs, and Selenium when a real browser must run JavaScript or perform interactions. Inspect one page first; the page type and crawl size should decide the tool.
Beautiful Soup parses HTML or XML; it does not fetch pages. Requests performs the fetch. Scrapy combines requests, scheduling, selectors, throttling, caching, and item pipelines for crawls. Selenium drives a browser through WebDriver, so it can execute JavaScript, click controls, submit forms, and keep browser state. See the Beautiful Soup docs, Scrapy overview, Scrapy requests and responses, and Selenium WebDriver.
Choose by page and job
| Situation | Start with | Why | Trade-off |
|---|---|---|---|
| One or a few static pages | Requests + Beautiful Soup | Small, inspectable request/parse pipeline | You add retries, throttling, pagination, and storage |
| Many pages or domains | Scrapy | Scheduler, spiders, selectors, exports, caching, and pipelines are built in | More project structure to learn and maintain |
| JavaScript-rendered or interactive pages | Selenium | Runs a real browser and can click, scroll, submit, and retain state | Higher CPU/RAM use and more timing failures |
| Mostly HTTP plus a few difficult views | Requests/API discovery + targeted Selenium | Use direct HTTP where possible and a browser only for rendered steps | Session and hand-off logic adds complexity |
1. Inspect the target before writing a scraper
- Open View Source or fetch the URL with
curl -Iand a plain GET. - Search the response for the field you need. If the text is present, it is probably static HTML.
- Use browser developer tools’ Network tab to see whether an XHR or fetch API returns structured data.
- Check pagination, authentication, rate limits, robots.txt, and terms before scheduling requests.
- Define an output schema and rules for missing values.
document.readyState == 'complete' only means the initial document finished loading. A single-page app can still be fetching and rendering data.
2. Static HTML: Requests and Beautiful Soup
python -m pip install requests beautifulsoup4 lxml
from urllib.parse import urljoin
import json, time, requests
from bs4 import BeautifulSoup
URL = 'https://example.com/news'
HEADERS = {'User-Agent': 'research-bot/1.0 (contact: you@example.com)'}
session = requests.Session()
for attempt in range(3):
try:
response = session.get(URL, headers=HEADERS, timeout=(10, 30))
response.raise_for_status()
break
except requests.RequestException:
if attempt == 2: raise
time.sleep(2 ** attempt)
soup = BeautifulSoup(response.text, 'lxml')
items = []
for card in soup.select('article'):
link = card.select_one('a[href]')
heading = card.select_one('h2, h3')
if not link or not heading: continue
items.append({'title': heading.get_text(' ', strip=True), 'url': urljoin(response.url, link['href'])})
with open('items.json', 'w', encoding='utf-8') as f:
json.dump(items, f, ensure_ascii=False, indent=2)
print(f'saved {len(items)} items')
Make a static scraper dependable
- Use a
Sessionfor connection reuse and cookies. - Set connect and read timeouts.
- Retry transient failures such as timeouts and selected 5xx responses, not 401, 403, or 404.
- Normalize whitespace, resolve relative links, and validate required fields.
- Throttle requests per host and persist progress.
3. Many URLs: Scrapy
python -m pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.com
import scrapy
class ProductsSpider(scrapy.Spider):
name = 'products'
allowed_domains = ['example.com']
start_urls = ['https://example.com/products']
custom_settings = {
'ROBOTSTXT_OBEY': True,
'DOWNLOAD_DELAY': 0.5,
'AUTOTHROTTLE_ENABLED': True,
'AUTOTHROTTLE_START_DELAY': 0.5,
'AUTOTHROTTLE_MAX_DELAY': 10,
'CONCURRENT_REQUESTS_PER_DOMAIN': 4,
'FEEDS': {'products.jsonl': {'format': 'jsonlines', 'overwrite': True}},
}
def parse(self, response):
for card in response.css('article.product'):
yield {'name': card.css('h2::text').get(default='').strip(), 'price': card.css('.price::text').get(default='').strip(), 'url': response.urljoin(card.css('a::attr(href)').get())}
next_url = response.css('a.next::attr(href)').get()
if next_url: yield response.follow(next_url, callback=self.parse)
scrapy crawl products
Useful settings include ROBOTSTXT_OBEY, download delay, per-domain concurrency, AutoThrottle, retry settings, HTTP caching, cookies, and feed exports. Prefer stable attributes over generated class names, log rejected items, and checkpoint long crawls.
4. JavaScript and interaction: Selenium
python -m pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
options = Options()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,1200')
driver = webdriver.Chrome(options=options)
try:
driver.get('https://example.com/dashboard')
wait = WebDriverWait(driver, 30)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'main')))
driver.execute_script('window.scrollTo(0, document.body.scrollHeight);')
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, '.results')))
rows = [el.text for el in driver.find_elements(By.CSS_SELECTOR, '.result-row')]
print(rows)
finally:
driver.quit()
Use explicit conditions such as element_to_be_clickable before clicks. Switch into iframes before querying them and switch to new windows after links open tabs. Browser automation is not the first choice when an underlying JSON endpoint provides the same fields.
5. A hybrid workflow
- Identify the JSON request supplying the page.
- Replay it with Requests, including required headers, cookies, pagination, and bounded timeouts.
- Validate authorization and endpoint stability.
- Reserve Selenium for login, token acquisition, or views that cannot be obtained through HTTP.
- Store raw responses and a parser version for auditing.
6. Pagination, sessions, and data quality
- Pagination: follow next links or cursors until absent, cap pages, and deduplicate canonical URLs.
- Sessions: use one Requests session or controlled Selenium browser state; never log session cookies.
- Encoding: preserve Unicode and use response headers where possible.
- Missing data: emit null plus a reason instead of shifting columns.
- Duplicates: key records by a stable ID or normalized URL.
- Change detection: keep fixture pages and parser checks.
7. Legal, ethical, and operational limits
Read terms and robots.txt, identify your user agent, collect only what you need, and keep rates low. Authentication, paywalls, bot checks, and personal data require specific authorization. Robots.txt is an implementation signal, not a complete legal answer.
8. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| No expected text | JavaScript rendering | Find the data API or use Selenium with an explicit wait |
| 403 or 429 | Rate, identity, or access policy | Stop, check permission, reduce concurrency, and honor Retry-After |
| Empty selector result | Markup changed, wrong frame, or shadow DOM | Inspect the live DOM, switch frames, or use an approved API |
| Intermittent Selenium failures | Race with asynchronous rendering | Wait for a condition and capture browser logs |
| StaleElementReferenceException | Framework re-rendered the node | Locate the element again |
| Timeouts or memory growth | Drivers, tabs, or queues remain open | Use finally, quit drivers, and process batches |
| Wrong prices or dates | Locale, hidden text, or formatting | Parse machine-readable attributes and validate units |
9. Performance, reliability, and cost
HTTP parsing is usually the lightest path: reuse connections, cache immutable responses, and write incrementally. Scrapy adds scheduling and back-pressure for large crawls. Selenium starts a browser per worker, so CPU, RAM, startup time, and synchronization dominate; use a small worker pool and target only pages that need it. Measure your own workload rather than assuming one library is universally faster.
Reliability comes from explicit timeouts, bounded retries with backoff, idempotent output, checkpointing, selector tests, and structured logs containing URL, status, duration, and parser version. Cost includes bandwidth, proxies or hosted browsers, storage, and engineering time.
Or skip the browser setup
If your deliverable is a clean visual capture rather than extracted text, ScreenshotNeo provides a single website screenshot API request. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. An MCP server lets Claude, Cursor, and other MCP clients take screenshots.
See the ScreenshotNeo API docs for full-page and element capture, dark mode, device presets, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, async webhooks, bulk capture, usage, and OpenAPI compatibility.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
The free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account.
FAQ
Is Beautiful Soup a scraper?
It is the parser in a fetch-and-parse pipeline. Requests retrieves the response.
Is Scrapy faster than Selenium?
They solve different workloads. Scrapy avoids browser overhead; Selenium is required for browser execution or interaction.
How do I scrape a single-page application?
Look for its JSON requests first. If no usable endpoint exists, use Selenium with explicit waits.
Should I use fixed sleeps?
Use condition-based waits where possible. Fixed delays are slower and less reliable.
How do I avoid scraping too aggressively?
Identify the client, obey applicable rules, throttle per host, cap concurrency, honor Retry-After, and stop when access is denied.
