ScreenshotNeo

BlogComparisons

Top 5 Python HTML Parsers

Compare Beautiful Soup, lxml, html5lib, html.parser and selectolax with runnable examples, trade-offs and guidance for choosing the right parser.

By the ScreenshotNeo team1 October 20268 min read

Short answer: choose Beautiful Soup for readable extraction code, lxml for direct tree work and performance-sensitive workloads, html5lib when browser-like WHATWG HTML parsing matters, Python’s built-in html.parser when you want no extra dependency, and selectolax when CSS selectors and throughput are priorities.

There is no universal winner. A parser’s error-recovery rules change the tree produced from malformed HTML, and Beautiful Soup is an interface that delegates parsing to a backend you choose. Pin that backend when reproducibility matters.

1. Comparison at a glance

Parser Best fit Main trade-off Install
Beautiful Soup Approachable extraction API Backend affects speed and tree shape pip install beautifulsoup4 lxml
lxml Direct HTML/XML trees and speed-sensitive work You must work with lxml’s API and semantics pip install lxml
html5lib WHATWG/browser-style HTML parsing Standards-oriented parsing can be slower pip install html5lib
html.parser Standard-library-only projects Different recovery behavior and a lower-level API Included with Python
selectolax CSS selectors and high-throughput extraction Benchmark and API choices need validation for your workload pip install selectolax

2. Beautiful Soup: the easiest extraction interface

Beautiful Soup gives Python code a consistent, readable API for finding tags, attributes and text. It is not one fixed parsing engine: it delegates to a selected backend such as html.parser, lxml or html5lib.

from bs4 import BeautifulSoup

html = """<article>
  <h1>Parser guide</h1>
  <a href='/docs'>Read docs</a>
</article>"""

# Pin the backend so every environment builds the same kind of tree.
soup = BeautifulSoup(html, "lxml")
print(soup.h1.get_text(strip=True))
for link in soup.select("a[href]"):
    print(link.get_text(" ", strip=True), link["href"])

Use BeautifulSoup(markup, "lxml") or another explicit backend in distributed code. If you omit the backend, Beautiful Soup uses the best installed parser, so dependency differences between machines can change results.

The project documentation says Beautiful Soup will never be as fast as the parsers beneath it. If response time is critical, it recommends working directly with lxml; if you want Beautiful Soup’s API but more speed, use lxml as its backend.

3. lxml: direct HTML and XML tree processing

lxml exposes efficient element trees, XPath and CSS-style querying through its HTML facilities. It is a strong default when you control the input shape, need XML as well as HTML, or want to minimize abstraction overhead.

from lxml import html

source = """<main>
  <h1>Parser guide</h1>
  <ul><li>One</li><li>Two</li></ul>
</main>"""

tree = html.fromstring(source)
title = tree.xpath("string(//h1)").strip()
items = [text.strip() for text in tree.xpath("//li//text()")]
print(title)
print(items)

Use XPath when you need precise structural queries. Compare lxml’s recovery behavior with your required semantics before accepting malformed third-party HTML.

4. html5lib: WHATWG parsing rules

html5lib is designed to conform to the WHATWG HTML specification as implemented by major browsers. Choose it when browser-like error recovery is more important than parsing speed. It can build different tree types, including ElementTree, minidom and lxml.etree.

import html5lib

source = "<!doctype html><p>Hello <b>world"
document = html5lib.parse(source, treebuilder="etree")

for element in document.iter():
    if element.text and element.text.strip():
        print(element.tag, element.text.strip())

html5lib’s standards behavior is useful for hostile or heavily malformed pages, but do not assume it is the fastest choice. Measure with your own documents and extraction operations.

5. Python’s built-in html.parser

html.parser avoids an additional package and is suitable for simple, controlled input. It is a callback-oriented parser, so you generally collect the data you need while start tags, end tags and text are visited.

from html.parser import HTMLParser

class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            attributes = dict(attrs)
            if "href" in attributes:
                self.links.append(attributes["href"])

parser = LinkParser()
parser.feed("<a href='/one'>One</a><a href='/two'>Two</a>")
parser.close()
print(parser.links)

Its tree and recovery behavior differ from lxml and html5lib, especially for invalid markup. Select it when avoiding dependencies matters more than a rich query API.

6. selectolax: CSS selectors with a fast HTML5 parser

selectolax provides CSS selection and HTML5 parsing. Its project currently recommends the Lexbor backend for its documented workflow.

from selectolax.lexbor import LexborHTMLParser

source = """<html><head><title>Example</title></head>
<body><article><h1>Hello</h1><a href='/docs'>Docs</a></article></body></html>"""

parser = LexborHTMLParser(source)
print(parser.css_first("title").text())
for node in parser.css("article a[href]"):
    print(node.text(strip=True), node.attributes.get("href"))

Use the project’s benchmark as a workload-specific data point, not a universal ranking. In its sample extraction task across 754 domains, reported times were 61.02 seconds for Beautiful Soup with html.parser, 9.09 for lxml/Beautiful Soup with lxml, 16.10 for html5_parser, 2.94 for selectolax Modest and 2.39 for selectolax Lexbor. Your documents, selectors, Python version and I/O pattern can change the result.

7. How malformed HTML changes the answer

Consider the fragment <a></p>. Beautiful Soup’s documentation shows that lxml drops the dangling closing paragraph and adds html/body, html5lib creates a paragraph and adds html/head/body, while html.parser leaves a simpler structure. None is universally correct until you define the parsing rules you need.

from bs4 import BeautifulSoup
from bs4.diagnose import diagnose

fragment = "<a></p>"
for backend in ("lxml", "html5lib", "html.parser"):
    soup = BeautifulSoup(fragment, backend)
    print(backend, soup.prettify())

# For difficult input, inspect all available parser interpretations.
diagnose(fragment)

When output surprises you, print or serialize the generated tree, run Beautiful Soup’s diagnose(), and choose the backend whose recovery rules match your application.

8. Installation and a repeatable extraction pattern

python -m venv .venv
. .venv/bin/activate
python -m pip install beautifulsoup4 lxml html5lib selectolax requests
from pathlib import Path
import requests
from bs4 import BeautifulSoup

url = "https://example.com/"
response = requests.get(url, timeout=30)
response.raise_for_status()

# Explicit backend: reproducible across machines.
soup = BeautifulSoup(response.content, "lxml")
records = []
for link in soup.select("a[href]"):
    records.append({
        "text": link.get_text(" ", strip=True),
        "href": link.get("href"),
    })
Path("links.json").write_text(__import__("json").dumps(records, indent=2), encoding="utf-8")

Parsing libraries do not execute JavaScript or render a page like a browser. If the content appears only after scripts run, obtain the rendered HTML with a browser automation tool or a rendering service before passing it to a parser.

9. Choosing by requirement

  • Readable scraping and quick prototypes: Beautiful Soup.
  • Beautiful Soup with better backend speed: Beautiful Soup plus lxml.
  • Maximum control over HTML/XML trees: lxml directly.
  • Browser-compatible handling of broken markup: html5lib.
  • No third-party dependency: html.parser.
  • CSS selectors and throughput: benchmark selectolax with Lexbor on representative pages.

Pin versions and parser backends in your lockfile. Add fixtures containing the malformed constructs your application receives, then assert on the resulting tree or extracted fields.

10. Performance, reliability and cost notes

  • Separate network time from parse time when profiling. A fast parser cannot compensate for slow downloads, retries or browser rendering.
  • Reuse HTTP connections, set finite timeouts and cap response sizes before parsing untrusted pages.
  • Benchmark complete extraction functions, not an empty parse loop. Include nested selectors, text cleanup and serialization.
  • Do not generalize the selectolax project benchmark to every workload; it is one project-produced sample.
  • For high-volume jobs, process documents incrementally where your parser supports it, bound concurrency and record parser version plus backend.

11. Troubleshooting

Symptom Likely cause Fix
Different output on two machines Beautiful Soup selected different installed backends Pass the backend explicitly and pin dependencies.
Missing content Content is inserted by JavaScript Render or fetch the post-render HTML, then parse it.
Unexpected extra html, head or body Backend-specific error recovery Inspect serialized trees and select rules that match your input contract.
FeatureNotFound in Beautiful Soup Requested backend is not installed Install the package, for example python -m pip install lxml.
Selector returns nothing Wrong selector, namespace, or response is not the expected document Log status, content type and a safe prefix of the response; test the selector against a saved fixture.
Memory usage grows Entire responses or trees are retained Release references, limit input sizes and process documents in bounded batches.

12. Or skip the browser setup

If your goal is a clean image or PDF of a page rather than an HTML tree, ScreenshotNeo is the alternative to try first. It captures a URL with one request, accepts cookie and consent banners before capture, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and bills only clean shots. Bot checks, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers identify the verdict and billing status.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also provides an MCP server for Claude, Cursor and other MCP clients, with tools for screenshots, page information and PDFs. It includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. FAQ

Is Beautiful Soup a parser?

It is a Python-facing parsing and extraction interface that delegates to a parser backend.

Which parser handles browser HTML most like a browser?

html5lib is designed around the WHATWG HTML parsing specification.

Should I use lxml directly or through Beautiful Soup?

Use lxml directly for maximum control and performance sensitivity; use it through Beautiful Soup when the simpler API is more valuable.

Can any of these parsers scrape JavaScript-rendered content?

No. They parse supplied HTML; they do not run page JavaScript.

How do I make parser behavior reproducible?

Pin package versions, pass the backend explicitly, and test against saved malformed and representative documents.