How to Parse HTML in Python: A Step-by-Step Guide for Beginners
Learn how to parse HTML in Python with html.parser, Beautiful Soup, and lxml, plus extraction patterns, fixes, and runnable code.
How do I parse HTML in Python? Start with HTML you already have as a string or file, choose a parser, build a structure (or handle events), then inspect tags, text, and attributes. Parsing does not download a page or run JavaScript; obtaining markup is separate.
1. Identify your input
Beautiful Soup accepts strings and file-like objects. Read a local file as text, or pass an existing response body.
from pathlib import Path
html_text = "<article><h1>Hello</h1></article>"
file_text = Path("page.html").read_text(encoding="utf-8")
2. Parse with Python’s html.parser
Python’s HTMLParser is event driven: it calls handlers for start tags, end tags, text, comments, and other markup.
from html.parser import HTMLParser
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.in_title = False
self.title = []
self.links = []
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
if tag == "title": self.in_title = True
if tag == "a" and "href" in attrs: self.links.append(attrs["href"])
def handle_endtag(self, tag):
if tag == "title": self.in_title = False
def handle_data(self, data):
if self.in_title: self.title.append(data)
html = "<title>Example</title><a href='/docs'>Docs</a>"
p = LinkParser(); p.feed(html); p.close()
print("".join(p.title).strip(), p.links)
This standard-library approach suits small callback workflows. It does not check that end tags match start tags. Feed chunks when streaming and call close() at the end.
3. Parse with Beautiful Soup
Install it with python -m pip install beautifulsoup4. Beautiful Soup creates Unicode-backed objects in a navigable tree. Name the parser explicitly: html.parser, lxml, or html5lib.
from bs4 import BeautifulSoup
html = "<article class='post'><h1>Parsing HTML</h1><p class='summary'>A guide.</p><a href='/next' data-kind='internal'>Next</a></article>"
soup = BeautifulSoup(html, "html.parser")
print(soup.find("h1").get_text(" ", strip=True))
print(soup.select_one("a[data-kind='internal']")["href"])
How do I extract text from HTML in Python?
article = soup.select_one("article.post")
print(article.get_text(" ", strip=True) if article else "")
for paragraph in soup.select("p"):
print(paragraph.get_text(" ", strip=True))
Extract attributes and links
for anchor in soup.select("a[href]"):
print(anchor.get_text(" ", strip=True), anchor["href"])
for node in soup.select("article.post a[href^='/']"):
print(node.get("href"))
Parse a file
from pathlib import Path
from bs4 import BeautifulSoup
with Path("page.html").open("rb") as handle:
soup = BeautifulSoup(handle, "html.parser")
print(soup.get_text(" ", strip=True))
4. Parse with lxml
lxml supplies HTML and XML APIs. Install it with python -m pip install lxml. If XHTML must follow XML rules, parse it as XML; parsing it as HTML can change the tree.
from lxml import html
doc = html.fromstring("<article><h1>Hello</h1><p>Body</p></article>")
print(doc.xpath("string(//h1)"))
print(doc.xpath("//p/text()"))
5. Choose deliberately
| Option | Useful when | Trade-off |
|---|---|---|
html.parser |
Standard library and callbacks | You implement state; no end-tag matching validation. |
| Beautiful Soup | Convenient tree search/navigation | It wraps a selected parser; malformed HTML can produce different trees. |
| lxml | Its XPath or HTML/XML APIs fit | Keep HTML and XML semantics distinct. |
No cited source establishes a universal performance winner. Base the choice on workflow, dependencies, malformed input, and HTML versus XHTML/XML.
6. Inspect and normalize
Different parsers can build different trees from malformed markup. Explicit parser selection improves repeatability.
print(soup.prettify()[:2000])
from urllib.parse import urljoin
for anchor in soup.select("a[href]"):
print(urljoin("https://example.com/docs/", anchor["href"]))
7. Complete extraction script
from pathlib import Path
from bs4 import BeautifulSoup
soup = BeautifulSoup(Path("page.html").read_text(encoding="utf-8"), "html.parser")
result = {
"title": soup.title.get_text(" ", strip=True) if soup.title else None,
"headings": [h.get_text(" ", strip=True) for h in soup.select("h1,h2,h3")],
"paragraphs": [p.get_text(" ", strip=True) for p in soup.select("p")],
"links": [{"text": a.get_text(" ", strip=True), "href": a.get("href")} for a in soup.select("a[href]")],
}
print(result)
8. Fetching is separate from parsing
An HTTP client can provide HTML to a parser, subject to site terms, robots rules, authentication, and rate limits. JavaScript-generated content may not be in the initial response.
import requests
from bs4 import BeautifulSoup
response = requests.get("https://example.com", timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
print(soup.get_text(" ", strip=True)[:500])
9. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
ModuleNotFoundError: bs4 |
Package absent | python -m pip install beautifulsoup4 |
| No selector result | Wrong selector, nesting, or JavaScript content | Print soup.prettify() and verify source HTML. |
| Different machines differ | Implicit parser/version | Name the parser and pin dependencies. |
| Text runs together | No separator | Use get_text(" ", strip=True). |
| Broken nesting | Malformed source recovery | Compare parsers and inspect the resulting tree. |
| XHTML XPath misses | Namespaces | Use a namespace map or local-name(); parse as XML when required. |
| Huge input is slow | Unbounded input or expensive queries | Limit size, stream callbacks, and narrow selectors. |
10. Performance, reliability, and cost
- Measure representative documents; the references provide no comparable benchmark.
- Reuse one parsed tree for multiple queries.
- Set size and time limits for untrusted input and never execute scripts as part of parsing.
- Explicit parser choice and malformed fixtures improve reproducibility.
- Local parsing has no API charge; network retrieval and third-party rendering have their own limits and terms.
11. Or skip the browser setup
If you need a clean image or PDF of a live page, ScreenshotNeo returns PNG, JPEG, WebP, or PDF from one GET request. It accepts cookie and consent banners and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; headers report the verdict and billing status.
See the ScreenshotNeo API docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Free includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free account.
12. FAQ
How do I parse HTML without installing a package?
Subclass HTMLParser and override the callbacks you need.
How do I use Beautiful Soup?
Install beautifulsoup4, construct BeautifulSoup(markup, "html.parser"), then use find, find_all, or CSS selectors.
Can a parser see JavaScript content?
Only if that content is in the HTML supplied to it. Parsing does not execute JavaScript.
Is Beautiful Soup itself a parser?
It is a tree interface over a selected underlying parser, so state the parser name.
Further reading
Read the Python documentation, Beautiful Soup documentation, structured markup tools, and lxml guide. O’Reilly lists Ryan Mitchell’s Web Scraping with Python, 3rd Edition (February 2024) for intermediate to advanced readers.


