ScreenshotNeo

BlogHow-to

How to Convert HTML to Text in Python

Convert HTML into readable plain text with Python, from a quick Beautiful Soup call to a dependency-free parser, with practical handling for whitespace, entities, and malformed markup.

By the ScreenshotNeo team30 September 202611 min read

How to Convert HTML to Text in Python

Use Beautiful Soup’s get_text() for the quickest practical conversion. Parse the HTML with an explicit parser, then choose a separator and whitespace behavior for your output:

from bs4 import BeautifulSoup

html = "<p>Hello <b>world</b>.</p><p>Next paragraph.</p>"
soup = BeautifulSoup(html, "html.parser")
text = soup.get_text(" ", strip=True)
print(text)
# Hello world. Next paragraph.

This processes HTML you already have as a string or decoded response body; it does not fetch a website or run JavaScript. If you need visible content injected by a page’s scripts, you need a workflow that obtains rendered page content before converting it.

1. Install and choose a conversion approach

Pick the approach that matches your output and dependency constraints:

Approach Use it when Tradeoff
Beautiful Soup You want convenient parsing and control over extracted text. Requires installing a package.
Python’s html.parser You need to avoid third-party dependencies. You implement text collection and formatting.
html2text You want readable plain text that retains some structure. Its purpose is readable text output; check the package documentation for the behavior you need.

Install Beautiful Soup with:

python -m pip install beautifulsoup4

Or install html2text if that is your chosen output style:

python -m pip install html2text

Name the parser explicitly. Beautiful Soup can use different parsers, and they may build different trees for invalid HTML. Specifying html.parser makes the choice clear and results more reproducible. See the Beautiful Soup documentation.

2. Extract text with Beautiful Soup

get_text() returns the text beneath a parsed document or tag as a Unicode string. Its first argument can separate text fragments, and strip=True trims whitespace around each fragment before joining.

A parser extracts text nodes; paragraph boundaries need an explicit formatting choice.
A parser extracts text nodes; paragraph boundaries need an explicit formatting choice.
from bs4 import BeautifulSoup

html = """
<article>
  <h1>Release notes</h1>
  <p>Version <strong>2.0</strong> is ready.</p>
  <p>Read the <a href="/guide">guide</a>.</p>
</article>
"""

soup = BeautifulSoup(html, "html.parser")
text = soup.get_text(" ", strip=True)
print(text)
# Release notes Version 2.0 is ready. Read the guide.

A space separator avoids accidentally joining adjacent text fragments such as Version and 2.0. It does not preserve paragraph layout: it flattens the extracted text into one string. If layout matters, select block elements and join their text yourself.

Preserve paragraph and heading boundaries

For indexing, summaries, or display where blocks should remain distinct, extract the relevant elements:

from bs4 import BeautifulSoup

html = """
<article>
  <h1>Release notes</h1>
  <p>Version <strong>2.0</strong> is ready.</p>
  <p>Read the <a href="/guide">guide</a>.</p>
</article>
"""

soup = BeautifulSoup(html, "html.parser")
blocks = soup.select("h1, h2, h3, p, li")
lines = [block.get_text(" ", strip=True) for block in blocks]
text = "\n".join(line for line in lines if line)
print(text)

This preserves the selected blocks as separate lines, but it is a policy you define rather than a layout reconstructed by get_text(). Choose selectors to match the content you want; for example, extracting all paragraphs may omit captions or list items.

Use a tag’s stripped_strings for custom joining

When you want to process text fragments individually, Beautiful Soup exposes stripped_strings:

from bs4 import BeautifulSoup

soup = BeautifulSoup("<p> One <b>two</b> <i>three</i> </p>", "html.parser")
parts = list(soup.p.stripped_strings)
print(parts)
# ['One', 'two', 'three']
print(" | ".join(parts))
# One | two | three

This is useful when each fragment needs custom treatment. For ordinary extraction, get_text() is shorter.

3. Remove unwanted elements before extraction

Sometimes the markup contains content you do not want in the result, such as navigation, a footer, or a particular component. Remove the matching nodes before extracting:

from bs4 import BeautifulSoup

html = """
<nav>Home | Products</nav>
<main><p>Article text.</p></main>
<footer>Copyright notice</footer>
"""

soup = BeautifulSoup(html, "html.parser")
for node in soup.select("nav, footer"):
    node.decompose()

main = soup.select_one("main")
text = main.get_text("\n", strip=True) if main else ""
print(text)
# Article text.

decompose() removes selected elements and their contents from the parse tree. Alternatively, select a narrower root, such as main or article, so unrelated page sections never enter the extraction. Selectors depend on the input markup, so handle a missing root explicitly.

Beautiful Soup documents that, with version 4.9.0 and later and the html.parser or lxml parser, contents of script, style, and template elements are generally not considered text. This behavior is qualified by version and parser; if those contents matter to your task, or unwanted content remains, inspect the output and remove the nodes explicitly.

4. Convert HTML without installing a package

The standard library’s HTMLParser can call your handler for text data. This small extractor collects those callbacks and normalizes whitespace:

from html.parser import HTMLParser

class TextExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []

    def handle_data(self, data):
        self.parts.append(data)

html = "<p>Hello <b>world</b>.</p>"
parser = TextExtractor()
parser.feed(html)
parser.close()
text = " ".join(" ".join(parser.parts).split())
print(text)
# Hello world.

handle_data() receives text between markup. The parser does not decide how you want paragraphs, headings, or list items represented, so this basic version collapses whitespace across the document. The Python documentation describes HTMLParser as able to parse invalid markup, but parsing and output formatting remain separate jobs. See the Python html.parser documentation.

Keep block boundaries with the standard library

If the output needs line breaks at paragraph and heading boundaries, track tags and insert separators. This example adds a newline when a block closes, then removes blank lines at the edges:

from html.parser import HTMLParser

BLOCK_TAGS = {"p", "div", "li", "h1", "h2", "h3", "h4", "h5", "h6", "br"}

class BlockTextExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag == "br":
            self.parts.append("\n")

    def handle_endtag(self, tag):
        if tag in BLOCK_TAGS and tag != "br":
            self.parts.append("\n")

    def handle_data(self, data):
        self.parts.append(data)

html = "<h1>Title</h1><p>First <b>paragraph</b>.</p><p>Second paragraph.</p>"
parser = BlockTextExtractor()
parser.feed(html)
parser.close()
lines = [" ".join(line.split()) for line in "".join(parser.parts).splitlines()]
text = "\n".join(line for line in lines if line)
print(text)

This is a formatting example, not a full HTML-to-visible-text renderer. For example, it does not resolve every possible layout convention, suppress hidden elements, or execute scripts. If nested block tags or malformed input produce extra boundaries, normalize the output for your own data and verify representative samples.

5. Decode entities and handle encodings

When parsing HTML, entities such as &amp; and numeric references such as &#169; need to become Unicode characters. Beautiful Soup converts entities while parsing. Python’s html.unescape() converts named and numeric character references according to HTML5 rules, and HTMLParser defaults to converting character references in text callbacks.

from html import unescape

print(unescape("Tom &amp; Ada ©"))
# Tom & Ada ©

The example has an escaped ampersand entity: one call produces &, not another decoding pass. Avoid decoding repeatedly unless the source truly contains text that was escaped multiple times. Repeated unescaping can change literal content unexpectedly. See the Python html module documentation.

If your input starts as bytes, decode it correctly before parsing. For a known file encoding, specify it when opening the file:

from pathlib import Path
from bs4 import BeautifulSoup

html = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
print(soup.get_text(" ", strip=True))

If you are given a response body as bytes and do not know its encoding, do not silently assume UTF-8 will always be correct. Establish the encoding from reliable response metadata or the document, or use a parsing workflow that detects and converts input encodings. Beautiful Soup documents its Unicode conversion and encoding-detection behavior in its documentation.

6. Get HTML from a file or HTTP response

Conversion libraries operate on the HTML you provide. Read a local file as text as shown above, or pass decoded response text from your HTTP client. For example, with Requests:

import requests
from bs4 import BeautifulSoup

response = requests.get("https://example.com", timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
print(soup.get_text(" ", strip=True))

This fetch-and-parse example depends on the server returning useful HTML in the response. It does not execute browser JavaScript, so a client-rendered page may have little content in its initial response. Set a finite timeout and call raise_for_status() so transport failures and HTTP errors are surfaced rather than mistaken for a successful empty conversion.

7. Use html2text for readable plain text

If you want readable output with some structure rather than concatenated text nodes, html2text is another option. Its package description presents it as a converter from HTML into clean, easy-to-read plain ASCII text.

import html2text

html = "<h1>Guide</h1><p>Read the <a href='https://example.com'>documentation</a>.</p>"
converter = html2text.HTML2Text()
text = converter.handle(html)
print(text)

Choose it when the desired result is text for people to read and you want a converter aimed at that format. Choose Beautiful Soup when you need to select parts of the document and control extraction directly. The available research supports this purpose distinction, not a detailed feature-by-feature or speed comparison. See the html2text package page.

8. Or skip the browser setup

If the HTML you need is a live web page and your next step is a screenshot or PDF rather than text, ScreenshotNeo can capture it with one GET request. The API accepts a URL and returns PNG, JPEG, WebP, or PDF. See the API documentation for options.

A clean capture can remove common overlays before taking the screenshot.
A clean capture can remove common overlays before taking the screenshot.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers say which page verdict and billing status applied. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month, with no card.

9. Troubleshooting common conversion problems

Symptom Likely cause Fix
Words run together Fragments were joined with an empty separator, or inline tags split a word boundary. Try get_text(" ", strip=True), then inspect output for cases where spaces should or should not be present.
Paragraphs have no separation get_text() flattens text fragments unless you choose a separator; a single separator does not infer block layout. Select block elements and join their text with newlines, or add block-boundary handling in your HTMLParser subclass.
Script or style text appears Behavior can vary by Beautiful Soup version, parser, and extraction approach; your chosen selector may include these nodes. Inspect the parsed tree and explicitly remove unwanted script, style, or other elements before extracting.
Accented characters are corrupted Bytes were decoded with the wrong character encoding. Determine the source encoding and decode accordingly; avoid guessing when metadata or document encoding is available.
Output is empty or incomplete The selected root may be missing, the source may contain little static markup, or content may be injected by JavaScript. Check the input and selection result. Obtain rendered content through a browser-based workflow when the page builds its content dynamically.
Different output on another machine A different parser was selected or parser behavior changed for invalid markup. Name the parser explicitly and keep the parsing environment consistent.
Entities remain visible The input might be escaped markup text rather than parsed HTML, or entities may have been escaped more than once. Confirm whether the input is HTML or text containing escaped HTML; parse or unescape once as appropriate.

10. Performance, reliability, and cost considerations

For normal application code, choose based on correctness and output requirements rather than an assumed speed ranking: the available sources do not provide comparable benchmarks. The standard-library route avoids installing a third-party package, while Beautiful Soup supplies convenient parsing and extraction controls. If processing large inputs, avoid retaining unnecessary intermediate representations and select only the content you need. Measure with your own representative documents if throughput matters.

Parsing is deterministic only to the extent that your input, parser choice, parser version, and cleanup rules are consistent. Malformed HTML can be interpreted differently by parsers, so pin the parser choice in code and validate outputs against representative documents. Add tests around important formatting rules in your application, such as whether list items and paragraph breaks remain separate.

The conversion libraries discussed here are local software choices; no per-request service cost is required for parsing an HTML string with them. Fetching a URL is a separate network operation with its own latency and failure modes. Use timeouts, check HTTP status, handle decoding, and distinguish a valid empty page from a failed fetch.

11. Frequently asked questions

Does Beautiful Soup convert a URL into text?

No. It parses markup you pass to it. Fetch the page separately, then provide its decoded HTML to Beautiful Soup.

Does HTML-to-text conversion produce exactly what a browser displays?

No. Parsing markup does not reproduce browser layout, visibility rules, or content added by JavaScript. It extracts text from the markup supplied to the parser.

Should I use Beautiful Soup or html2text?

Use Beautiful Soup for direct extraction, selecting elements, and controlling separators. Consider html2text when readable plain text with structure is the target. Verify the output against your document needs.

Text extraction returns text content; it does not automatically append each anchor’s destination. If you need URLs, select a elements and read their href attributes as a separate step.

Is regex a good way to remove tags?

A regular expression can remove simple tag-shaped strings, but HTML has nesting, attributes, entities, and malformed cases. A parser is the practical choice when the input is real HTML.

12. Quick checklist

  • Use Beautiful Soup’s get_text() for convenient extraction.
  • Choose a separator and decide whether the output should preserve block boundaries.
  • Specify the parser explicitly for reproducibility.
  • Remove unwanted nodes and select a meaningful root element.
  • Decode bytes with the correct encoding and avoid unnecessary repeated entity decoding.
  • Remember that static parsing does not execute JavaScript or recreate browser visibility.