How to Build a Web Scraper in Python
Build a small Python scraper with Requests and Beautiful Soup. Learn how to fetch, parse, validate, paginate and save data responsibly.
A basic Python web scraper fetches an HTML page, parses its markup, extracts fields you choose, validates them and saves structured records. For a small static site, requests plus beautifulsoup4 is a straightforward starting point. Use an official API or dataset when available, and check the target site’s terms and access rules before collecting data.
This guide builds a bounded scraper that checks robots.txt, fetches pages with finite timeouts, follows same-site pagination, validates records and writes JSON Lines. Its example selectors are illustrative: inspect the permitted target’s actual HTML and change the selectors to match it.
1. Define the scope and choose the data source
Before writing code, decide which fields you need, where they come from, and how far the scraper may go. A small, explicit boundary prevents an experiment from turning into an uncontrolled crawl.
- Prefer a documented API or downloadable dataset if it provides the data you need.
- Choose a permitted starting URL and a small set of fields, such as page title and article links.
- Set a domain boundary, a maximum page count and an output format.
- Inspect the site’s published terms and
robots.txt. A robots file communicates crawler access preferences; it does not establish legal permission or keep pages out of search results. - Use conservative request volume, honor applicable restrictions and stop if access is denied. No single rule determines whether scraping is permitted in every jurisdiction or situation.
Python’s urllib package includes URL handling, URL opening and a robots.txt parser. This example uses Requests for HTTP and Beautiful Soup for parsing. Requests documents timeouts, response status and exceptions; Beautiful Soup builds a searchable parse tree and supports multiple parser backends. Specifying the parser makes the choice explicit, especially because parser backends can construct different trees from malformed HTML. Python urllib documentation, Requests Quickstart, Beautiful Soup documentation, and Google’s robots.txt guide.
2. Set up a Python project
Use a virtual environment so dependencies stay isolated from other Python projects. The following commands work in a Unix-like shell; Windows users can activate the environment with .venv\\Scripts\\activate.
python -m venv .venv
source .venv/bin/activate
python -m pip install requests beautifulsoup4
Save the program below as scrape.py. Run it with python scrape.py. It uses only these two installed libraries; urllib.robotparser, JSON and URL utilities come with Python.
3. Build a bounded scraper
This sample parses article cards identified by article h2 a, extracts title and destination URL, and follows a same-site pagination link identified by a.next. Replace these CSS selectors to fit the target page. The script deliberately limits the number of pages and checks the domain before requesting a next link.
from __future__ import annotations
import json
import time
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
START_URL = "https://example.com/articles/"
USER_AGENT = "ExampleResearchBot/1.0 (contact: you@example.com)"
TIMEOUT_SECONDS = 15
MAX_PAGES = 5
DELAY_SECONDS = 1.0
OUTPUT_PATH = "articles.jsonl"
def same_site(url: str, allowed_host: str) -> bool:
"""Allow only HTTP(S) URLs on the starting hostname."""
parsed = urlparse(url)
return parsed.scheme in {"http", "https"} and parsed.hostname == allowed_host
def robots_allows(url: str, user_agent: str) -> bool:
"""Check the site's robots.txt rule for this URL and user agent."""
parts = urlparse(url)
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
parser = RobotFileParser()
parser.set_url(robots_url)
try:
parser.read()
except (OSError, ValueError) as exc:
raise RuntimeError(f"Could not read {robots_url}: {exc}") from exc
return parser.can_fetch(user_agent, url)
def fetch_html(session: requests.Session, url: str) -> str:
response = session.get(url, timeout=TIMEOUT_SECONDS)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "").lower()
if "html" not in content_type:
raise ValueError(f"Expected HTML at {url}; got {content_type or 'unknown content type'}")
return response.text
def clean_text(element) -> str | None:
if element is None:
return None
value = element.get_text(" ", strip=True)
return value or None
def parse_page(html: str, page_url: str) -> tuple[list[dict[str, str]], str | None]:
soup = BeautifulSoup(html, "html.parser")
records = []
for anchor in soup.select("article h2 a[href]"):
title = clean_text(anchor)
article_url = urljoin(page_url, anchor["href"])
if title and article_url:
records.append({"title": title, "url": article_url})
next_anchor = soup.select_one("a.next[href]")
next_url = urljoin(page_url, next_anchor["href"]) if next_anchor else None
return records, next_url
def main() -> None:
allowed_host = urlparse(START_URL).hostname
if not allowed_host:
raise ValueError("START_URL must include a hostname")
headers = {"User-Agent": USER_AGENT}
session = requests.Session()
session.headers.update(headers)
seen_pages: set[str] = set()
seen_records: set[str] = set()
page_url: str | None = START_URL
page_count = 0
with open(OUTPUT_PATH, "w", encoding="utf-8") as output:
while page_url and page_count < MAX_PAGES:
if not same_site(page_url, allowed_host):
print(f"Stopping: next URL is outside the allowed host: {page_url}")
break
if page_url in seen_pages:
print(f"Stopping: pagination loop detected at {page_url}")
break
if not robots_allows(page_url, USER_AGENT):
print(f"Stopping: robots.txt does not allow this URL: {page_url}")
break
seen_pages.add(page_url)
if page_count:
time.sleep(DELAY_SECONDS)
try:
html = fetch_html(session, page_url)
except requests.exceptions.Timeout as exc:
print(f"Timeout fetching {page_url}: {exc}")
break
except requests.exceptions.HTTPError as exc:
print(f"HTTP error fetching {page_url}: {exc}")
break
except requests.exceptions.RequestException as exc:
print(f"Request failed for {page_url}: {exc}")
break
records, next_url = parse_page(html, page_url)
for record in records:
# A stable URL is used here to avoid writing the same item twice.
if record["url"] in seen_records:
continue
if not record["title"].strip():
print(f"Skipping record with empty title on {page_url}")
continue
seen_records.add(record["url"])
output.write(json.dumps(record, ensure_ascii=False) + "\\n")
page_count += 1
page_url = next_url
print(f"Saved {len(seen_records)} records from {page_count} pages to {OUTPUT_PATH}")
if __name__ == "__main__":
main()
The sample uses an identifiable User-Agent, checks robots.txt for each page, sends requests through one Session, waits between pages, and stops on HTTP or network failures. Choose a User-Agent that accurately identifies your script and provides a contact method you monitor. Do not treat robots.txt as a substitute for checking permissions or terms.
What each part does
- Fetch:
Session.get()makes an HTTP request. A finite timeout prevents a request from waiting indefinitely.raise_for_status()turns unsuccessful HTTP status codes into an error path. - Parse:
BeautifulSoup(html, "html.parser")builds a tree from the returned markup. Beautiful Soup also supports other parser backends; install and select one explicitly if your project needs it. - Extract: CSS selectors locate elements.
get_text(" ", strip=True)normalizes whitespace, whileurljoin()converts relative links into absolute URLs. - Validate and bound: the script rejects non-HTML responses, checks expected title values, tracks visited pages and records, restricts the host and stops at a page limit.
- Store: JSON Lines writes one JSON object per line, which is convenient for streaming and later processing without loading all records into memory.
4. Find the right selectors and normalize fields
Inspect the HTML response the scraper actually received, not only the rendered browser view. In developer tools, identify a stable parent element and the fields inside it. Prefer semantic structure and stable attributes over generated class names that may change frequently.
| Need | Beautiful Soup pattern | Notes |
|---|---|---|
| First matching element | soup.select_one("h1") |
Returns None when no match exists; check before reading text. |
| All matching elements | soup.select("article h2 a[href]") |
Returns a list, including an empty list if nothing matched. |
| Attribute value | anchor.get("href") |
Can be absent; resolve relative URLs with urljoin(). |
| Visible text | element.get_text(" ", strip=True) |
Collapses nested text into a readable string with normalized spacing. |
Normalize data at extraction time. For dates, decide on one output representation and parse only known formats; record parse failures rather than silently substituting a wrong date. For numbers, remove presentation separators carefully and preserve the original value if conversion is uncertain. Keep source URLs with records so you can trace unexpected values back to their page.
Selectors are coupled to a page’s markup. If a publisher changes its templates, a scraper may return zero records or incomplete fields without a network error. Validate expected counts or required keys and alert or stop when the page shape changes.
5. Handle pagination without crawling the whole site
Pagination is site-specific. Some sites expose a next link, numbered pages, cursor parameters or an API. Follow only the mechanism you observed, and keep explicit limits. The sample uses one next link, a hostname allowlist, a visited-page set and MAX_PAGES. These guard against loops and accidental traversal into unrelated sections.
For numbered pages, construct each URL from the documented or observed pagination pattern and stop at a known maximum. For cursor pagination, preserve the returned cursor exactly and stop when it is absent or repeated. Avoid guessing URLs at scale. In every approach, normalize URLs consistently if the site presents equivalent URLs with different fragments or query parameters; remove query values only when you know they do not distinguish content.
6. Save and validate the output
JSON Lines is suitable for incremental output: each line is a standalone JSON record. CSV is useful for flat tables, but nested values need an agreed representation. SQLite is a better fit when you need queries, uniqueness constraints or incremental updates across runs.
Before relying on a scrape, inspect a few output lines and check the invariants that matter: required keys are present, URLs resolve to the expected site, values have the expected types, and duplicate records are handled intentionally. For durable pipelines, write to a temporary file and replace the prior output only after a complete successful run; otherwise a failed run can leave a partial file that looks complete.
7. When to use Requests and Beautiful Soup or Scrapy
| Approach | Fits when | Trade-off |
|---|---|---|
| Requests + Beautiful Soup | A one-off task or a small number of static pages | Simple and direct; you own pagination, limits, validation and run scheduling. |
| Scrapy | A recurring, multi-page crawl that benefits from a project structure and crawl workflow | More framework setup; provides request/response abstractions and project/deployment workflows. |
| Official API or dataset | The publisher offers a permitted structured source that meets your needs | Often avoids HTML selector maintenance; access and fields depend on that source. |
Scrapy’s request and response objects expose URL, status, headers and body, and its callbacks can extract data and yield follow-up requests. That makes it a sensible next step when a one-page script grows into repeatable crawling. Compare the tools based on URL volume, scheduling, crawl management, error handling, output integration and maintenance overhead. See the Scrapy request and response documentation and Scrapy tutorial.
8. Troubleshooting common problems
| Symptom | Likely cause | What to do |
|---|---|---|
| Timeout exception | The server or network did not respond before the configured limit. | Keep finite timeouts. Retry only transient failures with a small bounded backoff, reduce request frequency, and check whether the target is reachable. |
| HTTP 403 or 429 | The server denied access or signaled excessive request volume. | Stop and review the site’s access rules and terms. Reduce volume or use an authorized API. Do not attempt to evade access controls. |
| HTTP 404 | The URL is missing, stale or constructed incorrectly. | Check the starting URL and pagination link. Do not retry an unchanged missing URL. |
| Zero extracted records | The selector does not match returned markup, or the response differs from the browser page. | Inspect response.status_code, response.url, content type and a small excerpt of response.text. Update selectors from the received HTML. |
| Text or fields are missing | Markup structure changed, selector is too strict, or the value is not in the fetched HTML. | Check the element and its ancestors in the response source. If content is absent, look for an official structured source; changing selectors cannot extract data that was not returned. |
| Relative links point to the wrong place | The raw href was used without its base URL. |
Resolve with urljoin(page_url, href) and enforce the intended host boundary. |
| Different results across machines | Different parser backends or parser versions may build different trees, especially for malformed HTML. | Specify the parser backend, pin project dependencies where reproducibility matters, and keep a small fixture page for regression checks. |
| JSON decoding or encoding problems | Output is not valid JSON, text encoding was assumed incorrectly, or a non-serializable value was included. | Use json.dumps() per record, open output as UTF-8, and normalize values to strings, numbers, booleans, lists or dictionaries. |
9. Performance, reliability and cost
For small jobs, network latency usually matters more than parsing. Reuse a Requests Session to benefit from connection pooling, request only pages you need, and avoid downloading the same pages repeatedly. Set timeouts, cap pages, track progress and make output incremental. Add retries only for failures likely to be transient, with a small maximum attempt count and backoff; retrying every error can increase load and hide a persistent problem.
Concurrency can increase throughput but also increases load on the target and the chance of being throttled. Start sequentially, respect site guidance and use a conservative rate. Do not add parallel requests simply because the code can support them. For recurring crawls, record run time, page counts, error counts and validation failures so a broken selector or blocked run is visible.
The tools used here are open-source Python packages, but a crawl still has operational costs: compute, storage, maintenance and any infrastructure used to run it. The research sources do not establish a universal legal rule, safe request rate or benchmark. Keep collection within the permission and limits relevant to your target.
10. Or skip the browser setup
If your task is to capture a visual screenshot of a web page rather than extract structured records, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
ScreenshotNeo removes cookie and consent banners, newsletter popups and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information and PDF capture. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. A screenshot captures appearance; it does not replace a scraper that extracts and validates structured fields.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Can I scrape a page with Python’s standard library only?
Yes. Python includes URL opening, URL parsing and robots.txt parsing. A third-party parser such as Beautiful Soup is often more convenient for searching complex HTML, but it is optional.
Why does my scraper get different content than my browser?
The HTTP response may differ from the browser’s rendered page. First inspect the response body and headers. If the needed content is not there, look for a documented API or another permitted data source.
When should I move a script to Scrapy?
Consider it when crawling becomes a recurring project with many URLs, follow-up requests, crawl management or deployment needs that are cumbersome to maintain in a small script.
Does robots.txt grant permission to scrape?
No. It communicates crawler access preferences. Review applicable terms and permissions separately, and stop when access is denied.


