BeautifulSoup Alternatives in Python: Parsers, Selectors, and Crawlers
Compare Python’s best BeautifulSoup alternatives by speed, HTML repair, selectors, and crawling scope, with runnable examples to help you choose.

For a faster HTML parser with XPath, choose lxml. For a parser already included with Python, choose html.parser. For browser-like repair of badly malformed HTML, choose html5lib. If you need CSS or XPath selectors without adopting a crawler, choose Parsel; if the job includes spider orchestration, choose Scrapy. For a requests-based workflow that keeps cookies and form state, choose MechanicalSoup.
These options do different jobs. A parser turns markup into a tree; a selector extracts nodes from that tree; a crawler manages visiting many pages. Scrapy is a framework, so comparing it directly with a parser like lxml or BeautifulSoup can mislead. This guide gives you working examples and a decision process for picking the smallest suitable tool.
1. Decide what you need to replace
BeautifulSoup is convenient because it combines parsing with a forgiving API for navigating the resulting tree. You may want to replace it for one of several reasons: throughput, XPath, dependency constraints, HTML5 repair, CSS selector ergonomics, crawling, or stateful form interaction.
| Tool | Best fit | Main tradeoff |
|---|---|---|
| lxml | High-throughput HTML/XML parsing and XPath | Requires an external library with a C dependency; its tree and APIs differ from BeautifulSoup. |
html.parser |
Small scripts and restricted environments | Built into Python, but comparatively less fast and less lenient. |
| html5lib | Malformed markup that needs browser-like HTML5 recovery | Very lenient, but very slow. |
| Parsel | Standalone CSS and XPath extraction | Uses lxml underneath; it is a selector layer, not a crawler. |
| Scrapy | Spiders, crawling, scheduling, and extraction | A complete crawling framework, more than a parser. |
| MechanicalSoup | Requests sessions, forms, and stateful browsing | Still uses BeautifulSoup for parsing; it is not a BeautifulSoup-free parser. |
BeautifulSoup’s documentation recommends lxml when speed matters, and Scrapy’s documentation also calls out BeautifulSoup’s speed as a drawback. Neither source gives a reproducible apples-to-apples benchmark figure, so measure your own representative documents rather than relying on a claimed multiplier. See BeautifulSoup’s parser guidance and Scrapy’s selectors documentation.
2. Install only what your choice needs
Use a virtual environment so the parser and selector dependencies are reproducible. The standard-library option needs no package install. Install just the selected alternative for smaller deployments.
python -m venv .venv
# macOS/Linux:
. .venv/bin/activate
# Windows PowerShell:
.venv\Scripts\Activate.ps1
python -m pip install lxml
# Or, for HTML5 recovery:
python -m pip install html5lib
# Or, for standalone selectors:
python -m pip install parsel
# Or, for crawling:
python -m pip install scrapy
# Or, for stateful requests-based browsing:
python -m pip install MechanicalSoup
On a server or in a deployment image, pin tested dependency versions in your project’s lockfile or requirements file. Avoid selecting a parser implicitly: the same invalid input can produce different trees with different parsers. BeautifulSoup’s documentation demonstrates this behavior and recommends specifying the parser explicitly. That matters even if you later migrate away from BeautifulSoup, because changing tree construction can change what your selectors find.
3. Use lxml for speed and XPath
Use lxml when you need direct access to an HTML tree, XPath expressions, or high-throughput parsing. It supports HTML and XML parsing. Its API is lower-level than BeautifulSoup’s, so write small helper functions for repeated extraction and test your XPath against representative input.

from lxml import html
markup = """
<html><body>
<article class="story">
<h1>A parser comparison</h1>
<a href="/guide">Read guide</a>
</article>
</body></html>
"""
tree = html.fromstring(markup)
title = tree.xpath("string(//article[contains(concat(' ', normalize-space(@class), ' '), ' story ')]/h1)").strip()
links = tree.xpath("//article//a/@href")
print(title)
print(links)
For HTML obtained over HTTP, fetch it separately and pass the response body to html.fromstring. Handle HTTP status, timeouts, and character encoding at the request boundary. For XML, use the XML parser rather than treating the document as HTML; HTML parsing has different recovery and tree rules.
XPath notes: //a/@href returns attributes; string(//h1) gives the combined text value of the first matching node; and contains(concat(' ', normalize-space(@class), ' '), ' story ') checks a class token rather than a partial substring. Prefer an explicit class-token expression when class names may overlap. If XPath returns nothing, inspect the parsed tree and check whether the site’s markup differs from your assumption.
4. Use Python’s built-in html.parser when dependencies matter
html.parser is included in Python and is appropriate for basic extraction in constrained environments. It is a callback parser, rather than a navigable document tree. Subclass it and collect the events or values your task needs. This simple example collects text inside title elements; it deliberately does not attempt to implement a general-purpose nested DOM.
from html.parser import HTMLParser
class TitleCollector(HTMLParser):
def __init__(self):
super().__init__()
self.in_title = False
self.parts = []
def handle_starttag(self, tag, attrs):
if tag.lower() == "title":
self.in_title = True
def handle_endtag(self, tag):
if tag.lower() == "title":
self.in_title = False
def handle_data(self, data):
if self.in_title:
self.parts.append(data)
parser = TitleCollector()
parser.feed("<html><head><title>Example page</title></head></html>")
print("".join(parser.parts).strip())
Use this for small, known extraction tasks, not when you need arbitrary CSS or XPath queries. Python documents it as a simple HTML and XHTML parser; it does not offer BeautifulSoup’s tree-navigation interface. See the Python standard-library documentation.
5. Use html5lib when recovery matters most
Malformed HTML has no single inevitable parse tree. For example, parsers can handle a dangling closing tag differently. html5lib follows browser-like HTML5 parsing and recovery behavior, which can be useful when matching how a browser interprets broken markup. The cost is speed: it is very slow relative to the faster parser options described in the BeautifulSoup documentation.
import html5lib
markup = "<a>link</p>"
document = html5lib.parse(markup)
# html5lib returns an ElementTree. Use its supported tree API to inspect it.
root = document.getroot()
print(root.tag)
for element in root.iter():
if element.tag.endswith("a"):
print("anchor text:", "".join(element.itertext()).strip())
The exact tree API and namespace representation can matter when extracting elements; inspect the root and tags on your input. Choose html5lib because you need its repair behavior, not as a default for every page. If correctness depends on browser-equivalent recovery, add fixtures for representative malformed documents so parser upgrades or changes do not silently alter extraction.
6. Use Parsel for CSS and XPath without a crawler
Parsel provides a compact selector API for HTML and XML. It uses lxml underneath and supports both CSS and XPath. Use it when extraction expressions are the main need and Scrapy’s crawling framework would be unnecessary. Scrapy’s selectors are a thin wrapper around Parsel for integration with Scrapy responses.
from parsel import Selector
markup = """
<main>
<article class="story">
<h1>A selector example</h1>
<a href="/guide">Read guide</a>
</article>
</main>
"""
sel = Selector(text=markup)
title = sel.css("article.story h1::text").get()
links = sel.xpath("//article[contains(@class, 'story')]//a/@href").getall()
print(title)
print(links)
.get() returns the first result, or None when there is no match; .getall() returns a list. CSS is concise for common selectors. XPath is useful for attribute conditions, parent relationships, and more complex document queries. For nested text, use string(.) or normalize-space(.) on the selected element rather than assuming a single direct text node. Consult the Parsel usage documentation for selector behavior and examples.
7. Use Scrapy when the job is crawling
Scrapy is the right step when your task includes visiting many pages, following links, organizing extraction into spiders, and managing crawl behavior. It is not just an alternate parser. Scrapy’s selectors use CSS and XPath through Parsel, while the framework manages the crawl workflow.

import scrapy
class ArticleSpider(scrapy.Spider):
name = "articles"
start_urls = ["https://example.com/"]
def parse(self, response):
for article in response.css("article"):
yield {
"title": article.css("h2::text").get(),
"href": article.css("a::attr(href)").get(),
}
for href in response.css("a.next::attr(href)").getall():
yield response.follow(href, callback=self.parse)
Save this as spider.py in a Scrapy project’s spiders directory, then run it from the project with scrapy crawl articles. Replace the example domain and selectors with the target site’s actual pages and structure. Scrapy’s FAQ explains the framework-versus-parser distinction; its selector docs cover supported CSS and XPath behavior. Crawling also requires you to set suitable request and concurrency behavior for the site and workload.
8. Use MechanicalSoup for state and forms
MechanicalSoup suits workflows that need a requests-backed session, such as following links and submitting forms while retaining browser-like state. It still parses with BeautifulSoup; choose it to simplify stateful browsing, not to eliminate BeautifulSoup. Its StatefulBrowser accepts parser configuration. The documented default is lxml, and the documentation recommends specifying a parser so behavior is consistent across environments.
import mechanicalsoup
with mechanicalsoup.StatefulBrowser(
soup_config={"features": "lxml"},
raise_on_404=True,
) as browser:
page = browser.open("https://example.com/")
print(page.status_code)
print(browser.get_current_page().title.string)
Install both MechanicalSoup and lxml for this configuration. If lxml is unavailable, specify {"features": "html.parser"} instead. For forms, inspect the page’s actual form fields, select the intended form, fill the fields, and submit through the browser object; do not assume a site’s authentication or form flow is static. See the MechanicalSoup API documentation.
9. Migrate selectors without changing the extraction contract
- Capture representative source HTML. Include normal pages, missing fields, unusual nesting, and malformed examples that occur in your input.
- Choose the parser separately from selectors. Decide whether speed, built-in availability, or HTML5 repair is the key requirement.
- Rewrite one extraction at a time. Map each BeautifulSoup operation to a CSS or XPath expression, or to explicit parser callbacks for the standard library.
- Define missing-value behavior. BeautifulSoup code may rely on optional tags or attributes. Make your replacement return a deliberate default, skip the record, or report a structured extraction error.
- Compare output values and types. Check whitespace, entity decoding, relative URLs, duplicate matches, and empty selections.
- Deploy with explicit dependencies and parser choice. Keep local and production parser configuration aligned.
BeautifulSoup’s documentation notes that different parsers can generate different trees for invalid input. This means a migration can alter results even when the CSS-like query appears equivalent. Avoid making parser choice implicit in production code. If selectors are brittle, prefer stable attributes or semantic structure over positional paths such as “the third div.”
10. Performance, reliability, and cost
Performance: lxml is the usual first alternative to evaluate when parsing throughput matters; BeautifulSoup itself recommends it for speed. html5lib trades speed for browser-like recovery, while html.parser avoids an external dependency. Parsel inherits lxml’s parsing foundation. Scrapy’s framework adds capabilities around crawling, so benchmark the whole workload if you are comparing it with a one-document parser.
Reliability: parsing is deterministic only relative to the bytes, parser version, configuration, and extraction logic. Record or pin dependencies, make HTTP timeouts explicit, handle non-success status codes, and distinguish an empty result from a request failure. For evolving websites, monitor extraction completeness rather than treating a syntactically valid empty list as success. Different parsers can repair broken markup into different trees, so retain example inputs and expected outputs for key edge cases.
Cost: the libraries discussed here are software dependencies; operational costs generally come from your runtime, network requests, storage, and maintenance. A crawler can make many requests, so control request volume and concurrency. A browser-rendered screenshot service is a separate tool category: it returns rendered visual output rather than a parser tree for structured extraction. If your goal is a screenshot, consider [ScreenshotNeo](https://screenshotneo.com), which accepts a URL and returns a PNG, JPEG, WebP, or PDF.
11. Troubleshooting common problems
| Symptom | Likely cause | Fix |
|---|---|---|
| lxml import fails | Package missing from the active environment or installation failed. | Activate the intended virtual environment, install lxml there, and confirm the same interpreter runs your script. |
| XPath returns an empty list | Selector does not match the parsed tree, document differs from expectation, or the content is not in the downloaded HTML. | Inspect a saved response and parsed structure; verify the page’s actual markup and use a stable selector. |
| CSS selection works on one page but not another | Markup is inconsistent, multiple roots exist, or the content is generated after the initial response. | Check each response and determine whether server HTML contains the content. Parsel documents that CSS queries on multi-root documents start from the first root; use an XPath root query where appropriate. |
| Different results on two machines | Parser was selected implicitly or dependency versions differ. | Specify parser explicitly and lock dependencies. Compare a malformed fixture with both environments. |
| html5lib seems slow | Browser-like repair costs more than faster parsing. | Use it only for documents requiring that recovery behavior; evaluate lxml for ordinary inputs. |
| Text is missing or oddly spaced | Text resides in nested nodes, or HTML structure does not imply whitespace in the parsed text. | Use a descendant-text expression such as XPath normalize-space(.) and normalize whitespace according to the data contract. |
| MechanicalSoup warns about parser choice | Parser configuration is absent or nonstandard. | Pass soup_config={"features": "lxml"} or the intended installed parser explicitly. |
| Scrapy crawl extracts nothing | Spider selectors do not match the response, callback did not follow the expected links, or content is client-rendered. | Inspect the response body and selector matches; correct selectors and link traversal based on returned HTML. |
Or skip the browser setup
If the task is to capture a rendered page, use ScreenshotNeo’s one-call API instead of installing and maintaining a browser capture stack. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers indicate the page verdict and billing outcome.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, no card required.
Frequently asked questions
Is lxml a drop-in replacement for BeautifulSoup?
No. Both can parse HTML, but their APIs and tree representations differ. Keep your extraction behavior explicit while translating selectors.
Which option is best if I need XPath?
Use lxml for direct XPath parsing, or Parsel if you want a selector API with both CSS and XPath. Scrapy selectors are a fit when those selectors sit inside a crawler.
Should I learn Scrapy before replacing BeautifulSoup?
Only if you need crawl management and spiders. For parsing one downloaded document, a parser or Parsel selector is simpler.
Can these tools scrape JavaScript-rendered content?
They parse the response content they receive. If the desired content is created only after browser-side JavaScript runs, a plain HTTP fetch may not contain it; use an appropriate rendering workflow or a site API.
Is MechanicalSoup fully browser-based?
No. It provides a stateful requests-backed browsing interface and BeautifulSoup parsing; it is useful for sessions and forms, but does not provide a full JavaScript-rendering browser.
