ScreenshotNeo

BlogHow-to

How to Scrape Websites with Beautiful Soup in Python

Fetch a page with Requests, parse it with Beautiful Soup, and extract links or other fields with checks for missing data, errors, and JavaScript-rendered content.

By the ScreenshotNeo team4 October 20269 min read

Beautiful Soup parses HTML into a tree; it does not download web pages. To scrape a page, fetch its HTML with an HTTP client such as Requests, check that the request succeeded, then parse the response with Beautiful Soup and extract the elements you need. This guide uses Python 3, Requests, and Beautiful Soup 4.

1. Install the packages

Install beautifulsoup4 and requests in the same Python environment that will run your script. The import name for Beautiful Soup is bs4. The examples use Python’s built-in html.parser, so no additional parser package is required.

python -m pip install beautifulsoup4 requests

If your system distinguishes Python 3 with a separate command, use python3 -m pip. In a virtual environment, activate it first, then run the installation command. See the Beautiful Soup documentation for installation and parsing details.

This complete script checks the HTTP response before parsing, sets a timeout, handles links without an href, and reports when no links match. Replace the example URL with a page you are allowed to access.

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

URL = "https://example.com/"

try:
    response = requests.get(
        URL,
        headers={"User-Agent": "ExampleResearchBot/1.0"},
        timeout=(5, 20),  # connect timeout, read timeout, in seconds
    )
    response.raise_for_status()
except requests.exceptions.Timeout as exc:
    raise SystemExit(f"The request timed out: {exc}")
except requests.exceptions.HTTPError as exc:
    raise SystemExit(f"The server returned an unsuccessful HTTP status: {exc}")
except requests.exceptions.RequestException as exc:
    raise SystemExit(f"The request failed: {exc}")

soup = BeautifulSoup(response.content, "html.parser")

# find_all returns every matching anchor. href=True skips anchors without href.
links = []
for anchor in soup.find_all("a", href=True):
    label = anchor.get_text(" ", strip=True)
    href = anchor.get("href")
    links.append({
        "text": label,
        "url": urljoin(response.url, href),
    })

if not links:
    print("No links with href attributes were found in the returned HTML.")
else:
    for link in links:
        print(f"{link['text'] or '[no link text]'}: {link['url']}")

response.raise_for_status() raises an exception for unsuccessful HTTP status codes, so an error response does not quietly get treated as the page you meant to scrape. Requests does not set a timeout unless you supply one; its Quickstart recommends using timeouts in production code. The tuple above sets separate connection and response-read limits.

3. Choose elements with Beautiful Soup

Start by inspecting the HTML actually returned by the server. Then target stable tags, attributes, or CSS classes. These common methods cover most extraction tasks:

  • find(name, attrs, ...) returns the first matching tag, or None when there is no match.
  • find_all(name, attrs, ...) returns all matching tags. An empty result is normal when nothing matches.
  • select_one(css_selector) returns the first match for a CSS selector, or None.
  • select(css_selector) returns all matches for a CSS selector.
  • tag.get("attribute") safely reads an attribute and returns None if it is absent.
  • tag.get_text(" ", strip=True) collects a tag’s text and separates nested text with spaces.

Use find and find_all when tag names and attributes are clear. Use CSS selectors when the page’s class or nesting structure makes the target more direct. Beautiful Soup’s selector support is provided by SoupSieve in current releases; available selectors can depend on the installed version. The official documentation describes filters by tag, attribute, string, regular expression, list, function, and True.

# First page title; guard against a missing title tag.
title_tag = soup.find("title")
title = title_tag.get_text(" ", strip=True) if title_tag else None

# Find elements with a known id or class.
main = soup.find(id="main-content")
headings = soup.find_all("h2")

# Equivalent CSS selector style for many tasks.
first_article = soup.select_one("article.post")
article_links = soup.select("article.post a[href]")

# Read a potentially absent attribute safely.
image = soup.find("img")
image_src = image.get("src") if image else None

Classes can contain multiple values, and site markup changes over time. Prefer a meaningful identifier or a selector anchored to a stable part of the page; verify your results against a small sample before collecting many pages.

4. Extract structured records safely

For repeated items such as article cards, find each container first, then extract fields within that container. This avoids accidentally pairing a title from one item with a link from another. Missing fields should be represented explicitly or skipped according to your use case.

records = []

for card in soup.select("article"):
    heading = card.select_one("h2, h3")
    link = card.select_one("a[href]")

    if heading is None or link is None:
        # The page may have an advert, a different card type, or changed markup.
        continue

    records.append({
        "title": heading.get_text(" ", strip=True),
        "url": urljoin(response.url, link["href"]),
    })

if not records:
    print("No complete article records found; inspect the response and selectors.")
else:
    for record in records:
        print(record)

Do not assume every tag has every attribute. For example, find_all("a", href=True) selects only anchors with an href, while calling find("a")["href"] can fail if there is no anchor or no attribute.

5. Pick a parser deliberately

Pass a parser name to BeautifulSoup so the script’s behavior is clear. The examples use html.parser, included with Python. Beautiful Soup also supports third-party parsers such as lxml and html5lib; install the one you choose in the active environment.

Parser When to consider it Trade-off
html.parser A simple setup with no parser dependency beyond Python. For malformed HTML, its constructed tree can differ from other parsers.
lxml You want a third-party parser; the Beautiful Soup docs note its parsing speed advantage. It must be installed. For raw parsing speed, the docs recommend using lxml directly instead of Beautiful Soup.
html5lib You want parsing behavior modeled on a browser. It must be installed and may have different speed and tree behavior.

Different parsers can produce different trees from malformed markup. If repeatability matters, specify the parser and keep the dependency available in every environment. Parser availability and behavior can vary with installed versions, so check the current project documentation when selecting or upgrading one.

6. Know what the server returned

Beautiful Soup can only inspect the HTML you give it. A page may return a successful status while showing an access-denied page, a challenge, or a shell that expects JavaScript to populate the content later. A static Requests response will not execute that page’s JavaScript. Inspect a short excerpt of the response, its final URL, status, and relevant tags before concluding that the selector is wrong.

print("status:", response.status_code)
print("final URL:", response.url)
print("content type:", response.headers.get("Content-Type"))
print(response.text[:1000])  # inspect a small prefix, not the full page

If the data is absent from the returned HTML because it is generated in the browser, identify whether the site exposes an appropriate documented data endpoint or whether a browser-rendering workflow is needed. Follow the site’s access guidance, keep request volume modest, avoid unnecessary personal data, and stop if access is blocked. These are practical precautions, not a claim that scraping is permitted on every site.

7. cURL and Node.js equivalents for fetching

Beautiful Soup is a Python parser, so these examples fetch the same kind of HTML without using Beautiful Soup. They are useful for checking the server response or fitting the fetch step into another toolchain.

cURL

curl --fail --show-error --location --max-time 25 \
  -H 'User-Agent: ExampleResearchBot/1.0' \
  'https://example.com/' \
  -o page.html

--fail returns an error for unsuccessful HTTP responses, --location follows redirects, and --max-time bounds the total transfer time. Inspect the saved HTML before building selectors against it.

Node.js

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 25000);

try {
  const response = await fetch('https://example.com/', {
    headers: { 'User-Agent': 'ExampleResearchBot/1.0' },
    signal: controller.signal,
  });

  if (!response.ok) {
    throw new Error(`HTTP ${response.status} ${response.statusText}`);
  }

  const html = await response.text();
  console.log(html.slice(0, 1000));
} finally {
  clearTimeout(timer);
}

Run this in a Node.js environment that supports the built-in Fetch API. It fetches and prints HTML; parsing it into a DOM requires a separate HTML parsing library.

8. Troubleshooting

Symptom Likely cause What to check or change
ModuleNotFoundError: No module named 'bs4' Beautiful Soup is not installed in the Python environment running the script. Run python -m pip install beautifulsoup4 using that same interpreter. Install the distribution name beautifulsoup4; import it as bs4.
find() result causes an attribute error find() returned None. Check the result before reading its text or attributes. Inspect the response and confirm the tag and filter.
find_all() or select() returns an empty list The selector does not match, the markup changed, or the desired content is not in the returned HTML. Print a small response excerpt; inspect the exact tag, attributes, classes, status and final URL. Test a broader selector, then narrow it.
Script hangs or takes too long No request timeout was set, or the server is slow. Set a connect/read timeout in Requests. Consider retrying only transient failures with a bounded retry policy and modest delays.
Parser output differs across machines A parser was left implicit, or parser libraries differ between environments. Specify html.parser, lxml, or html5lib and install the chosen dependency consistently.
HTML is an error or challenge page The server returned a block, error, redirect destination, or other unexpected document. Check status, final URL, response headers, and a small HTML excerpt. Respect access guidance and stop if the site blocks the request.
Content appears in a browser but not in Requests It may be inserted after the initial HTML is loaded by JavaScript. Compare the browser’s rendered content with the raw response. Use an appropriate documented data source or a browser-rendering workflow if needed.
KeyError: 'href' The selected tag has no href attribute. Filter with href=True or read it with tag.get("href") and handle None.

9. Performance, reliability, and cost

For small jobs, a single request followed by a few targeted selectors is usually the simplest approach. Parsing speed depends on document size, parser choice, and how much work your code performs; the Beautiful Soup documentation notes lxml’s speed advantage and recommends direct lxml use when raw parsing speed dominates. There are no universal timing figures for a page or machine.

  • Bound waiting time: set request timeouts and handle network and HTTP exceptions.
  • Keep traffic controlled: avoid unnecessary repeat requests, pace multi-page work, and do not retry indefinitely.
  • Make extraction observable: record status, final URL, result counts, and a small sample so markup changes are easier to detect.
  • Make output robust: handle missing tags and attributes, normalize whitespace, and validate important fields before saving records.
  • Account for work: Requests and Beautiful Soup are open-source Python packages; your practical costs are development, compute, storage, and any infrastructure or browser service you choose. No benchmark or fixed cost follows from the libraries alone.

10. Or skip the browser setup

If your actual goal is a screenshot rather than structured HTML data, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Cookie banners, newsletter popups, and chat widgets are removed before the shot, and each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status. An MCP server lets AI agents use screenshot tools, including take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

11. Frequently asked questions

Can Beautiful Soup scrape a whole website by itself?

No. It parses documents you provide. You need separate code to fetch pages, discover which URLs to visit, and decide when to stop.

Can I use it with a local HTML file?

Yes. Read the file and pass its contents to BeautifulSoup with a parser name, just as you would pass a response body.

No. Use a URL resolver such as Python’s urllib.parse.urljoin with the page URL and the link’s href.

Should I use CSS selectors or find_all?

Use the form that makes the target easiest to understand and maintain. Both can express many common extraction tasks; verify selector support in the version installed in your environment.