ScreenshotNeo

BlogGuides

How to Extract Text from HTML with Python: Library Guide for Developers

Learn reliable HTML text extraction in Python with BeautifulSoup, lxml, html.parser, cleanup patterns, troubleshooting, and production guidance.

By the ScreenshotNeo team30 September 20269 min read

How to Extract Text from HTML with Python: Library Guide for Developers

To extract readable text from HTML with Python, parse the document and call Beautiful Soup’s get_text() method with a separator and whitespace trimming:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "lxml")
text = soup.get_text(" ", strip=True)

The separator keeps words from running together when inline tags or block elements sit next to each other. The explicit lxml parser makes behavior consistent across machines. This guide covers complete scripts, parser choices, selecting the main content, cleanup, malformed markup, dynamic pages, performance, and production troubleshooting.

1. Install a parser and fetch the HTML

Beautiful Soup provides the tree API; a parser backend turns the byte or string input into that tree. Install Beautiful Soup and lxml:

python -m pip install beautifulsoup4 lxml requests

Here is a runnable script that downloads a page, checks the response, extracts visible-looking text, and writes it to a file:

from pathlib import Path

import requests
from bs4 import BeautifulSoup

url = "https://example.com/"
response = requests.get(
    url,
    timeout=30,
    headers={"User-Agent": "text-extractor/1.0"},
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "lxml")
text = soup.get_text(" ", strip=True)
Path("page.txt").write_text(text + "\n", encoding="utf-8")
print(text)

Keep network retrieval and parsing as separate steps. That makes it easier to retry transient HTTP failures, cache the original response, and test extraction against a saved fixture.

2. Beautiful Soup extraction patterns

Extract the whole document

text = soup.get_text(" ", strip=True)

get_text() returns the text beneath a document or tag. The first argument is the separator inserted between text fragments; strip=True removes surrounding whitespace. See the Beautiful Soup documentation for get_text() and related APIs.

The extraction pipeline: fetch HTML, parse it with an explicit backend, select content, then normalize text.
The extraction pipeline: fetch HTML, parse it with an explicit backend, select content, then normalize text.

Extract only the article or main element

Whole-document extraction often includes navigation, footer links, cookie notices, comments, and duplicated responsive markup. Select the region you need before converting it to text:

main = soup.select_one("main")
if main is None:
    raise ValueError("The page has no main element")

article_text = main.get_text(" ", strip=True)

Selectors can target an article, a known class, or an ID:

node = soup.select_one("article.post, .article-body, #content")
text = node.get_text(" ", strip=True) if node else ""

For one known element, this approach prevents menus and related-content modules from entering the result. A selector is site-specific, so keep it in configuration rather than scattering it through your code.

Process fragments with stripped_strings

Use stripped_strings when you need to inspect, filter, or transform fragments individually:

parts = [part for part in soup.stripped_strings]
text = " ".join(parts)

This is useful when you want to discard a particular heading, stop after a paywall marker, or apply custom normalization to each fragment.

Remove unwanted regions before extraction

for selector in ["script", "style", "template", "nav", "footer", ".cookie-banner"]:
    for node in soup.select(selector):
        node.decompose()

text = soup.get_text(" ", strip=True)

Use decompose() when a node and its contents should be removed. Make sure selectors match the pages you process; a broad selector can delete legitimate article text.

3. Choosing between lxml, html.parser, and html5lib

Approach Strength Trade-off Best fit
Beautiful Soup + lxml Friendly tree API with a robust parser backend Extra dependency General extraction from messy pages
Beautiful Soup + html5lib HTML5-style parsing and browser-like error recovery Usually slower and adds a dependency Inputs where HTML5 recovery matters
Beautiful Soup + html.parser Simple installation and familiar API Different recovery behavior on invalid markup Small scripts and controlled input
html.parser.HTMLParser Python standard library and callback control You implement collection and cleanup Dependency-light, event-driven processing

Beautiful Soup documents all three selectable parser families and warns that malformed markup can produce different trees. Choose one explicitly instead of relying on whichever parser happens to be installed:

soup = BeautifulSoup(html, "lxml")

Pin the parser dependency in your requirements file and keep representative malformed fixtures in your test suite. That recommendation follows from the documented parser differences: a parser upgrade or a different environment can change the resulting tree.

4. Standard-library extraction with HTMLParser

If adding Beautiful Soup is undesirable, Python includes html.parser.HTMLParser, an event-driven parser. Its callbacks receive start tags, text, comments, and other markup events. The following implementation collects text and normalizes whitespace:

from html.parser import HTMLParser

class TextExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []

    def handle_data(self, data):
        self.parts.append(data)

html = "<main><h1>Hello</h1><p>A <strong>short</strong> page.</p></main>"
extractor = TextExtractor()
extractor.feed(html)
extractor.close()
text = " ".join(" ".join(extractor.parts).split())
print(text)

The standard library documentation describes HTMLParser as a simple HTML and XHTML parser. This callback approach gives you control over which tags start or stop collection, but you must add those rules yourself.

Collect only a selected region

from html.parser import HTMLParser

class MainTextExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []
        self.depth = 0

    def handle_starttag(self, tag, attrs):
        if tag == "main":
            self.depth += 1
        elif self.depth:
            self.depth += 1

    def handle_endtag(self, tag):
        if self.depth:
            self.depth -= 1

    def handle_data(self, data):
        if self.depth:
            self.parts.append(data)

extractor = MainTextExtractor()
extractor.feed(html)
extractor.close()
text = " ".join(" ".join(extractor.parts).split())

For complex selectors, malformed documents, or nested exclusion rules, Beautiful Soup’s tree is usually simpler to maintain.

5. Whitespace, entities, and readable output

HTML whitespace is not the same as human-readable paragraph spacing. Start with get_text(" ", strip=True), then apply only the normalization your downstream system needs:

import re

text = soup.get_text(" ", strip=True)
text = re.sub(r"\s+", " ", text)
text = text.strip()

Use a newline separator when preserving block boundaries is more important than compact output:

text = soup.get_text("\n", strip=True)
lines = [line.strip() for line in text.splitlines() if line.strip()]
text = "\n".join(lines)

Do not blindly collapse every space if the result feeds code, poetry, tables, or preformatted content. Decide whether your consumer needs paragraphs, individual fragments, or one search-friendly string.

HTML character references are decoded by the parser. If you receive an already escaped string, identify whether it needs one decoding pass before parsing; decoding repeatedly can turn literal text into markup.

6. Static HTML versus JavaScript-rendered pages

requests downloads the server response. It does not run the page’s JavaScript, so an application that inserts article text after load may produce an empty or incomplete result. Diagnose this by saving and inspecting the response:

Selecting the article container prevents navigation and repeated page elements from entering the extracted text.
Selecting the article container prevents navigation and repeated page elements from entering the extracted text.
Path("response.html").write_text(response.text, encoding="utf-8")
print(len(response.text), response.url, response.headers.get("content-type"))

If the text exists in response.html, parse it directly. If the response contains only an application shell, use the site’s documented data endpoint or a browser automation tool that is appropriate for your access rules. A screenshot service captures pixels or PDFs; it does not replace semantic HTML extraction when you need text nodes.

7. Fetching and parsing with cURL, Python, and Node.js

Use cURL to inspect the server response before writing extraction code:

curl -L --max-time 30 -A 'text-extractor/1.0' https://example.com/ -o page.html

Python remains the most direct option for this guide:

import requests
from bs4 import BeautifulSoup

html = requests.get("https://example.com/", timeout=30).text
soup = BeautifulSoup(html, "lxml")
print(soup.get_text(" ", strip=True))

Node.js can fetch the HTML, but you need an HTML parser package for equivalent tree operations. With a package such as Cheerio installed:

import * as cheerio from 'cheerio';

const res = await fetch('https://example.com/');
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
const $ = cheerio.load(html);
console.log($('main').length ? $('main').text().trim() : $('body').text().trim());

Choose one runtime for a production pipeline unless you have a clear reason to operate several. Consistent parser versions and cleanup rules make output easier to compare.

8. Or skip the browser setup

If your workflow starts with a URL and you need a reliable page artifact before processing it, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, cookie and consent banners, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. The API also supports full-page and element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture, usage reporting, PDFs, and HTML/CSS-to-image. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Free accounts include 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account and try the API with no card.

9. Troubleshooting common extraction failures

Symptom Likely cause Fix
Words run together No separator between fragments Use get_text(" ", strip=True) or a newline separator.
Menus and footers pollute output Whole-document extraction Select main or article; remove known regions first.
Empty text Content is inserted by JavaScript Inspect the saved response; use a data endpoint or browser-capable workflow.
FeatureNotFound Requested parser is not installed Install its package or use an installed parser explicitly.
Different output on two machines Implicit or different parser backend Name the parser and pin dependency versions.
Encoding appears corrupted Incorrectly decoded bytes Check response.encoding, content type, and write files as UTF-8.
Request hangs No network timeout Set a finite timeout and add bounded retries for transient failures.
Important text is missing Overbroad decompose() selector Log matched selectors and narrow them to the unwanted container.

10. Performance, reliability, and cost

Parsing is usually cheaper than downloading the page. Reduce total work by fetching only once, reusing a session for many URLs, selecting a smaller subtree before extraction, and avoiding repeated CSS queries. For large documents, process saved responses in a worker queue and write normalized text incrementally.

Reliability comes from explicit timeouts, checked HTTP status codes, bounded retries, and fixtures that represent malformed markup. Cache responses when the source permits it, record the final URL after redirects, and retain parser and cleanup configuration with each output so a later change can be explained.

Requests to a third-party site can incur bandwidth, proxy, or browser costs outside Python itself. Beautiful Soup and html.parser do not charge per extraction. If you use ScreenshotNeo, only clean shots are billed; failed loads, bot checks, blank pages, timeouts, and cache hits are free according to the response verdict and billing headers. Choose caching TTL, asynchronous jobs, or bulk capture when they match your workload.

11. A production checklist

  • Fetch with a timeout and check the HTTP status.
  • Record the final URL, content type, parser name, and parser version.
  • Choose lxml, html5lib, or html.parser explicitly.
  • Target main, article, or a site-specific container when possible.
  • Remove navigation, scripts, styles, templates, and consent elements only with tested selectors.
  • Use a separator and deliberate whitespace normalization.
  • Keep saved fixtures for malformed and JavaScript-rendered pages.
  • Measure output length and flag unexpectedly empty or tiny results.
  • Respect the site’s terms, robots guidance, authentication, and rate limits.

12. FAQ

Is Beautiful Soup a text extractor?

It is an HTML parsing library with a tree API. Its get_text() method extracts descendant text, while selecting the correct content region remains your responsibility.

Should I use lxml or html.parser?

Use lxml for a friendly general-purpose workflow when an extra dependency is acceptable. Use the standard-library parser for dependency-light scripts and callback control. Name the parser explicitly either way.

How do I preserve paragraphs?

Extract with a newline separator, remove blank lines, and keep the resulting lines instead of collapsing all whitespace into one string.

Can HTML extraction read text hidden behind a login?

Only if your HTTP or browser session is authenticated and you are authorized to access it. Supply the required session cookies or use the site’s supported API.

Why does screenshot text differ from extracted HTML?

A screenshot represents rendered pixels after layout and JavaScript. HTML extraction reads the response or a parsed DOM. They answer different questions and can legitimately produce different content.