How to Extract Text from HTML with Python: Library Guide for Developers
Learn reliable HTML text extraction in Python with BeautifulSoup, lxml, html.parser, cleanup patterns, troubleshooting, and production guidance.

To extract readable text from HTML with Python, parse the document and call Beautiful Soup’s get_text() method with a separator and whitespace trimming:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "lxml")
text = soup.get_text(" ", strip=True)
The separator keeps words from running together when inline tags or block elements sit next to each other. The explicit lxml parser makes behavior consistent across machines. This guide covers complete scripts, parser choices, selecting the main content, cleanup, malformed markup, dynamic pages, performance, and production troubleshooting.
1. Install a parser and fetch the HTML
Beautiful Soup provides the tree API; a parser backend turns the byte or string input into that tree. Install Beautiful Soup and lxml:
python -m pip install beautifulsoup4 lxml requests
Here is a runnable script that downloads a page, checks the response, extracts visible-looking text, and writes it to a file:
from pathlib import Path
import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
response = requests.get(
url,
timeout=30,
headers={"User-Agent": "text-extractor/1.0"},
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "lxml")
text = soup.get_text(" ", strip=True)
Path("page.txt").write_text(text + "\n", encoding="utf-8")
print(text)
Keep network retrieval and parsing as separate steps. That makes it easier to retry transient HTTP failures, cache the original response, and test extraction against a saved fixture.
2. Beautiful Soup extraction patterns
Extract the whole document
text = soup.get_text(" ", strip=True)
get_text() returns the text beneath a document or tag. The first argument is the separator inserted between text fragments; strip=True removes surrounding whitespace. See the Beautiful Soup documentation for get_text() and related APIs.

Extract only the article or main element
Whole-document extraction often includes navigation, footer links, cookie notices, comments, and duplicated responsive markup. Select the region you need before converting it to text:
main = soup.select_one("main")
if main is None:
raise ValueError("The page has no main element")
article_text = main.get_text(" ", strip=True)
Selectors can target an article, a known class, or an ID:
node = soup.select_one("article.post, .article-body, #content")
text = node.get_text(" ", strip=True) if node else ""
For one known element, this approach prevents menus and related-content modules from entering the result. A selector is site-specific, so keep it in configuration rather than scattering it through your code.
Process fragments with stripped_strings
Use stripped_strings when you need to inspect, filter, or transform fragments individually:
parts = [part for part in soup.stripped_strings]
text = " ".join(parts)
This is useful when you want to discard a particular heading, stop after a paywall marker, or apply custom normalization to each fragment.
Remove unwanted regions before extraction
for selector in ["script", "style", "template", "nav", "footer", ".cookie-banner"]:
for node in soup.select(selector):
node.decompose()
text = soup.get_text(" ", strip=True)
Use decompose() when a node and its contents should be removed. Make sure selectors match the pages you process; a broad selector can delete legitimate article text.
3. Choosing between lxml, html.parser, and html5lib
| Approach | Strength | Trade-off | Best fit |
|---|---|---|---|
| Beautiful Soup + lxml | Friendly tree API with a robust parser backend | Extra dependency | General extraction from messy pages |
| Beautiful Soup + html5lib | HTML5-style parsing and browser-like error recovery | Usually slower and adds a dependency | Inputs where HTML5 recovery matters |
| Beautiful Soup + html.parser | Simple installation and familiar API | Different recovery behavior on invalid markup | Small scripts and controlled input |
| html.parser.HTMLParser | Python standard library and callback control | You implement collection and cleanup | Dependency-light, event-driven processing |
Beautiful Soup documents all three selectable parser families and warns that malformed markup can produce different trees. Choose one explicitly instead of relying on whichever parser happens to be installed:
soup = BeautifulSoup(html, "lxml")
Pin the parser dependency in your requirements file and keep representative malformed fixtures in your test suite. That recommendation follows from the documented parser differences: a parser upgrade or a different environment can change the resulting tree.
4. Standard-library extraction with HTMLParser
If adding Beautiful Soup is undesirable, Python includes html.parser.HTMLParser, an event-driven parser. Its callbacks receive start tags, text, comments, and other markup events. The following implementation collects text and normalizes whitespace:
from html.parser import HTMLParser
class TextExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.parts = []
def handle_data(self, data):
self.parts.append(data)
html = "<main><h1>Hello</h1><p>A <strong>short</strong> page.</p></main>"
extractor = TextExtractor()
extractor.feed(html)
extractor.close()
text = " ".join(" ".join(extractor.parts).split())
print(text)
The standard library documentation describes HTMLParser as a simple HTML and XHTML parser. This callback approach gives you control over which tags start or stop collection, but you must add those rules yourself.
Collect only a selected region
from html.parser import HTMLParser
class MainTextExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.parts = []
self.depth = 0
def handle_starttag(self, tag, attrs):
if tag == "main":
self.depth += 1
elif self.depth:
self.depth += 1
def handle_endtag(self, tag):
if self.depth:
self.depth -= 1
def handle_data(self, data):
if self.depth:
self.parts.append(data)
extractor = MainTextExtractor()
extractor.feed(html)
extractor.close()
text = " ".join(" ".join(extractor.parts).split())
For complex selectors, malformed documents, or nested exclusion rules, Beautiful Soup’s tree is usually simpler to maintain.
5. Whitespace, entities, and readable output
HTML whitespace is not the same as human-readable paragraph spacing. Start with get_text(" ", strip=True), then apply only the normalization your downstream system needs:
import re
text = soup.get_text(" ", strip=True)
text = re.sub(r"\s+", " ", text)
text = text.strip()
Use a newline separator when preserving block boundaries is more important than compact output:
text = soup.get_text("\n", strip=True)
lines = [line.strip() for line in text.splitlines() if line.strip()]
text = "\n".join(lines)
Do not blindly collapse every space if the result feeds code, poetry, tables, or preformatted content. Decide whether your consumer needs paragraphs, individual fragments, or one search-friendly string.
HTML character references are decoded by the parser. If you receive an already escaped string, identify whether it needs one decoding pass before parsing; decoding repeatedly can turn literal text into markup.
6. Static HTML versus JavaScript-rendered pages
requests downloads the server response. It does not run the page’s JavaScript, so an application that inserts article text after load may produce an empty or incomplete result. Diagnose this by saving and inspecting the response:

Path("response.html").write_text(response.text, encoding="utf-8")
print(len(response.text), response.url, response.headers.get("content-type"))
If the text exists in response.html, parse it directly. If the response contains only an application shell, use the site’s documented data endpoint or a browser automation tool that is appropriate for your access rules. A screenshot service captures pixels or PDFs; it does not replace semantic HTML extraction when you need text nodes.
7. Fetching and parsing with cURL, Python, and Node.js
Use cURL to inspect the server response before writing extraction code:
curl -L --max-time 30 -A 'text-extractor/1.0' https://example.com/ -o page.html
Python remains the most direct option for this guide:
import requests
from bs4 import BeautifulSoup
html = requests.get("https://example.com/", timeout=30).text
soup = BeautifulSoup(html, "lxml")
print(soup.get_text(" ", strip=True))
Node.js can fetch the HTML, but you need an HTML parser package for equivalent tree operations. With a package such as Cheerio installed:
import * as cheerio from 'cheerio';
const res = await fetch('https://example.com/');
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
const $ = cheerio.load(html);
console.log($('main').length ? $('main').text().trim() : $('body').text().trim());
Choose one runtime for a production pipeline unless you have a clear reason to operate several. Consistent parser versions and cleanup rules make output easier to compare.
8. Or skip the browser setup
If your workflow starts with a URL and you need a reliable page artifact before processing it, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, cookie and consent banners, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. The API also supports full-page and element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture, usage reporting, PDFs, and HTML/CSS-to-image. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
Free accounts include 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account and try the API with no card.
9. Troubleshooting common extraction failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Words run together | No separator between fragments | Use get_text(" ", strip=True) or a newline separator. |
| Menus and footers pollute output | Whole-document extraction | Select main or article; remove known regions first. |
| Empty text | Content is inserted by JavaScript | Inspect the saved response; use a data endpoint or browser-capable workflow. |
FeatureNotFound |
Requested parser is not installed | Install its package or use an installed parser explicitly. |
| Different output on two machines | Implicit or different parser backend | Name the parser and pin dependency versions. |
| Encoding appears corrupted | Incorrectly decoded bytes | Check response.encoding, content type, and write files as UTF-8. |
| Request hangs | No network timeout | Set a finite timeout and add bounded retries for transient failures. |
| Important text is missing | Overbroad decompose() selector |
Log matched selectors and narrow them to the unwanted container. |
10. Performance, reliability, and cost
Parsing is usually cheaper than downloading the page. Reduce total work by fetching only once, reusing a session for many URLs, selecting a smaller subtree before extraction, and avoiding repeated CSS queries. For large documents, process saved responses in a worker queue and write normalized text incrementally.
Reliability comes from explicit timeouts, checked HTTP status codes, bounded retries, and fixtures that represent malformed markup. Cache responses when the source permits it, record the final URL after redirects, and retain parser and cleanup configuration with each output so a later change can be explained.
Requests to a third-party site can incur bandwidth, proxy, or browser costs outside Python itself. Beautiful Soup and html.parser do not charge per extraction. If you use ScreenshotNeo, only clean shots are billed; failed loads, bot checks, blank pages, timeouts, and cache hits are free according to the response verdict and billing headers. Choose caching TTL, asynchronous jobs, or bulk capture when they match your workload.
11. A production checklist
- Fetch with a timeout and check the HTTP status.
- Record the final URL, content type, parser name, and parser version.
- Choose
lxml,html5lib, orhtml.parserexplicitly. - Target
main,article, or a site-specific container when possible. - Remove navigation, scripts, styles, templates, and consent elements only with tested selectors.
- Use a separator and deliberate whitespace normalization.
- Keep saved fixtures for malformed and JavaScript-rendered pages.
- Measure output length and flag unexpectedly empty or tiny results.
- Respect the site’s terms, robots guidance, authentication, and rate limits.
12. FAQ
Is Beautiful Soup a text extractor?
It is an HTML parsing library with a tree API. Its get_text() method extracts descendant text, while selecting the correct content region remains your responsibility.
Should I use lxml or html.parser?
Use lxml for a friendly general-purpose workflow when an extra dependency is acceptable. Use the standard-library parser for dependency-light scripts and callback control. Name the parser explicitly either way.
How do I preserve paragraphs?
Extract with a newline separator, remove blank lines, and keep the resulting lines instead of collapsing all whitespace into one string.
Can HTML extraction read text hidden behind a login?
Only if your HTTP or browser session is authenticated and you are authorized to access it. Supply the required session cookies or use the site’s supported API.
Why does screenshot text differ from extracted HTML?
A screenshot represents rendered pixels after layout and JavaScript. HTML extraction reads the response or a parsed DOM. They answer different questions and can legitimately produce different content.


