ScreenshotNeo

BlogHow-to

How to Use Beautiful Soup for Web Scraping

Learn to fetch a web page with Requests, parse it with Beautiful Soup, extract text and links, and troubleshoot missing or unexpected results.

By the ScreenshotNeo team4 October 20268 min read

Beautiful Soup parses HTML or XML that you already have; it does not download pages or run their JavaScript. A basic scraping workflow uses an HTTP client such as Requests to retrieve a page, checks the response, passes its content to Beautiful Soup, and then searches the parsed tree for the data you need. Beautiful Soup describes itself as “a Python library for pulling data out of HTML and XML files.” Beautiful Soup documentation; Requests Quickstart.

1. Install Beautiful Soup and Requests

Use Python 3. The package is named beautifulsoup4, but you import it from bs4. Requests handles HTTP retrieval.

python -m pip install beautifulsoup4 requests

Save the example below as scrape.py and run it with python scrape.py.

2. Fetch a page, parse it, and extract data

This complete script retrieves a page, checks for an HTTP error, parses the returned HTML with Python’s built-in html.parser, and handles missing elements without crashing. The example extracts a page title and links from https://example.com/; adapt its selectors to the actual markup you receive.

import requests
from bs4 import BeautifulSoup

URL = "https://example.com/"


def main():
    try:
        response = requests.get(
            URL,
            headers={"User-Agent": "ExampleResearchBot/1.0"},
            timeout=(5, 20),
        )
        response.raise_for_status()
    except requests.RequestException as exc:
        raise SystemExit(f"Could not retrieve {URL}: {exc}")

    # Use response.content (bytes) and let Beautiful Soup handle document encoding.
    soup = BeautifulSoup(response.content, "html.parser")

    title = soup.title.get_text(" ", strip=True) if soup.title else "(no title)"
    print(f"Title: {title}")

    for link in soup.select("a[href]"):
        text = link.get_text(" ", strip=True)
        href = link.get("href")
        print({"text": text, "href": href})


if __name__ == "__main__":
    main()

Requests returns a Response; raise_for_status() raises an exception for unsuccessful HTTP status codes. The parser receives the body only after retrieval succeeds. Requests’ Quickstart documents response status, headers, and content.

3. Choose a parser deliberately

Beautiful Soup supports Python’s built-in html.parser and optional lxml and html5lib parsers. Install optional parsers separately, then name the parser explicitly in the constructor. Parser choice can change the tree produced from malformed markup, so use the same choice in development and deployment. For XML, the Beautiful Soup documentation directs users to use lxml in XML mode.

Parser When to consider it Example
html.parser Start with Python’s built-in HTML parser and avoid an extra parser dependency. BeautifulSoup(markup, "html.parser")
lxml Use when it is part of your environment or you need the documented XML mode. Install and pin it in your project. BeautifulSoup(markup, "lxml")
html5lib Use when its HTML parsing behavior suits your input. Install and pin it explicitly. BeautifulSoup(markup, "html5lib")
XML with lxml Parse XML as XML instead of letting HTML rules interpret it. BeautifulSoup(markup, "xml")

For example, install an optional parser with python -m pip install lxml or python -m pip install html5lib. Choose based on compatibility with your markup, the parse tree you need, and dependencies you can maintain. Do not assume one parser is always faster or produces the same result: performance depends on the input and versions, and the cited documentation does not establish a current benchmark.

4. Find elements and extract their values

Use find() for one match

find() returns the first matching tag or None if it finds no match. Check for None before reading its contents.

heading = soup.find("h1")
if heading is None:
    print("No h1 found")
else:
    print(heading.get_text(" ", strip=True))

Use find_all() for repeated matches

find_all() returns all matching tags as a result you can iterate over. Filter or validate the extracted values when a page may contain empty or incomplete elements.

for paragraph in soup.find_all("p"):
    text = paragraph.get_text(" ", strip=True)
    if text:
        print(text)

Use CSS selectors when relationships read more clearly

select() accepts CSS selectors. Use it when a class, attribute, or relationship expresses the target more clearly than nested searches. These examples find article cards and links with an href; the selectors must match the returned markup.

for card in soup.select("article.product-card"):
    name_tag = card.select_one("h2")
    link_tag = card.select_one("a[href]")
    name = name_tag.get_text(" ", strip=True) if name_tag else None
    href = link_tag.get("href") if link_tag else None
    print({"name": name, "href": href})

Read text and attributes

Use get_text(" ", strip=True) to join descendant text with spaces and trim surrounding whitespace. Use get() for an attribute so a missing attribute returns None instead of raising a key error.

for image in soup.select("img"):
    print({
        "alt": image.get("alt"),
        "src": image.get("src"),
    })

Do not depend on positions such as “the third paragraph is the price” unless the page format guarantees that structure. Prefer meaningful tags, attributes, and relationships, and check for missing matches.

5. Inspect and refine your selectors

  1. Print or save the response status, content type, and a short portion of the response body.
  2. Confirm the response actually contains the content you expect; an error page or different page can still be valid HTML.
  3. Inspect the relevant returned markup, then choose a tag, class, ID, attribute, or relationship that exists there.
  4. Parse with the same explicit parser used in production, and check whether the parsed tree contains the target.
  5. Test the no-match case and decide whether to skip, log, retry retrieval, or report a changed page structure.

A browser’s rendered appearance is not proof that the same content appears in the initial HTTP response. A page can add content after JavaScript runs; the simple Requests-plus-Beautiful-Soup workflow does not run page scripts. If required data is absent from returned markup, you need an authorized retrieval method that supplies it, such as a documented data endpoint or browser-based rendering.

6. Handle encoding and response details

Requests exposes decoded text as response.text and the response body bytes as response.content. Its encoding choice is influenced by the HTTP response headers. Passing bytes to Beautiful Soup, as in the main example, lets the parser handle document encoding information. If characters look corrupted, inspect response.headers, response.encoding, and the original bytes before changing selectors.

print("status:", response.status_code)
print("content-type:", response.headers.get("Content-Type"))
print("requests encoding:", response.encoding)
print("body preview:", response.text[:500])

Use the response content that was actually returned. Do not change an encoding by guesswork without checking the response headers and document.

7. Common errors and fixes

Symptom Likely cause What to do
ModuleNotFoundError: No module named 'bs4' Beautiful Soup is missing from the Python environment running the script. Run python -m pip install beautifulsoup4 with the same Python interpreter used to run the script.
ModuleNotFoundError: No module named 'requests' Requests is not installed in that environment. Run python -m pip install requests with that interpreter.
FeatureNotFound: Couldn't find a tree builder The requested optional parser is not installed, or the parser name is unavailable. Install the matching package, such as lxml or html5lib, or use html.parser.
Lookup returns None or an empty list The response lacks the expected markup, the selector is wrong, or the page changed. Inspect the response and parsed tree; update the selector and handle absent results.
Text is an error message or access notice The request returned a different page than expected, possibly due to status, headers, or site access behavior. Check status, content type, headers, and body. Follow the site’s access rules; do not try to bypass access controls.
Browser shows data that the script cannot find JavaScript may populate it after the initial HTML response. Check whether the data is available through an authorized documented endpoint or an appropriate rendering method.
Accented or non-Latin characters look wrong Response encoding metadata and the decoding path may not match the document. Inspect response headers, Requests’ encoding, and response bytes before changing parsing logic.
Results differ across machines Parser choice or parser versions may differ. Specify the parser and pin project dependencies so environments use consistent versions.
Request hangs or fails intermittently Network delays, unavailable host, or missing timeout handling. Set a timeout, handle Requests exceptions, and use a measured retry policy only when appropriate.

8. Reliability, performance, and responsible use

  • Bound network waits: use explicit connect and read timeouts. A timeout limits how long a request waits; it does not ensure the server will respond successfully.
  • Check every response: inspect status and content type where relevant, and call raise_for_status() before treating a body as the expected page.
  • Expect page changes: selectors are coupled to markup. Validate required fields, handle absent values, and log enough context to diagnose changes.
  • Keep dependencies reproducible: specify a parser and maintain the dependencies used in deployment.
  • Avoid excessive requests: reuse an HTTP session when making many requests to the same service, pace requests, and avoid fetching pages you do not need. No throughput benchmark is implied here.
  • Check site-specific rules: review the target site’s current terms, robots directives, access controls, privacy and copyright requirements, and applicable rules for your situation. Obtain appropriate authorization and avoid overloading the service. The library documentation does not determine whether scraping a particular site is permitted.

Or skip the browser setup

Beautiful Soup is for extracting data from returned HTML. If what you need is a screenshot of the rendered page, ScreenshotNeo provides a website screenshot API and MCP server. The self-hosted approach above gives you control over retrieval and parsing; this one-call option returns a capture without setting up browser automation. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. Response headers report the page verdict and billing status.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

FAQ

Can Beautiful Soup scrape a site by itself?

No. It parses markup supplied to it. Use an HTTP client to retrieve the page, then parse the response.

Does Beautiful Soup execute JavaScript?

No. It parses the markup it receives. JavaScript-generated content may not exist in the initial response.

Should I use find() or a CSS selector?

Use whichever expresses the target structure clearly and remains easy to maintain. Both depend on the actual returned markup.

Can I use it with XML?

Yes. Beautiful Soup handles XML; its documentation recommends the lxml parser in XML mode.

Why did my selector stop working?

The response may have changed, differ from the browser view, or be parsed differently by another parser. Inspect the response and parsed tree first.