Data Parsing: How to Turn Web Data into Structured Data
Learn how to parse HTML tables, page elements, and XML into validated Python data structures, with runnable examples and practical maintenance checks.
Data parsing turns source text or markup into a structure your program can inspect and transform. For web data, choose the parser to match the input: use Beautiful Soup for elements such as headings, links, and repeated records; use pandas read_html() for HTML tables; and use pandas read_xml() for XML with relatively shallow records. Then map the result into an explicit schema and validate it against real pages.
Parsing alone does not establish that extracted data is complete, correct, or stable. This guide builds a small, validated Python workflow and explains how to handle common changes and failures.
1. Identify the input and define the output
Start by inspecting one representative source. Determine whether the values you need appear in a table, repeated HTML elements, attributes such as link destinations, or XML nodes. Also check whether the content is present in the HTML you have or depends on scripts; there is no single parsing method that covers every dynamic page.
Before writing selectors, define the output fields and types. For example:
records = [
{"name": "Example item", "price": 12.5, "url": "https://example.com/item"}
]
Decide how to represent missing values, duplicate records, and inconsistent formats. Keep useful source context, such as the page URL, when it helps identify or audit a record.
2. Choose a parser for the shape of the source
| Input | Starting point | Result and considerations |
|---|---|---|
| HTML elements spread across a page | Beautiful Soup | A navigable parse tree; parser choice can change the tree for malformed markup. |
| HTML table | pandas read_html() |
A list of DataFrames, even if only one table is found. Inspect and select the intended table. |
| XML with shallow repeating records | pandas read_xml() |
A DataFrame from nodes and attributes; deep nesting may need flattening first. |
Compare tools using the input shape, desired output, malformed-markup behavior, dependencies, and how you will detect source changes. Beautiful Soup describes itself as “a Python library for pulling data out of HTML and XML files.” Its documentation discusses lxml, html5lib, and Python’s built-in html.parser; inspect the actual result for your input instead of assuming all parsers build the same tree. See the Beautiful Soup documentation.
3. Parse page elements with Beautiful Soup
Install Beautiful Soup and choose a parser. This example extracts repeated records with a title and link. Replace the example selectors with ones observed in the source page.
python -m pip install beautifulsoup4 requests
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
page_url = "https://example.com/catalog"
response = requests.get(page_url, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
records = []
for card in soup.select("article.product"):
title_node = card.select_one("h2")
link_node = card.select_one("a[href]")
if title_node is None or link_node is None:
continue
records.append({
"title": title_node.get_text(" ", strip=True),
"url": urljoin(page_url, link_node["href"]),
})
if not records:
raise ValueError("No product records found; check the source and selectors")
print(records)
select() and select_one() use CSS selectors. Use get_text(" ", strip=True) to join text with spaces and trim surrounding whitespace. Read attributes directly, such as link_node["href"], and resolve relative links with urljoin(). If markup is malformed, try another documented parser and compare the resulting tree and extracted values.
4. Turn an HTML table into a DataFrame
For actual HTML tables, pandas can handle table parsing directly. Its read_html() accepts HTML strings, files, or URLs and returns a list of DataFrames. Select the table you intend to use and inspect its headers and rows.
python -m pip install pandas lxml
import pandas as pd
page_url = "https://example.com/prices"
tables = pd.read_html(page_url)
if not tables:
raise ValueError("No HTML tables found")
df = tables[0]
print(df.head())
print(df.dtypes)
required = {"Product", "Price"}
missing = required - set(df.columns)
if missing:
raise ValueError(f"Missing expected columns: {sorted(missing)}")
df = df.rename(columns={"Product": "product", "Price": "price"})
df["product"] = df["product"].astype("string").str.strip()
df["price"] = pd.to_numeric(df["price"], errors="coerce")
if df["product"].isna().any() or df["price"].isna().any():
raise ValueError("A required product or price value is missing or invalid")
df.to_csv("prices.csv", index=False)
df.to_json("prices.json", orient="records", indent=2)
There may be multiple tables, and the first is not necessarily the one you want. Check the dimensions, columns, and representative values before using or exporting the result. pandas documents this behavior in its HTML I/O guide.
5. Parse XML into a DataFrame
For shallow repeating XML records, read_xml() can map elements and attributes into a DataFrame. Select the repeating node with xpath when needed.
python -m pip install pandas lxml
from io import StringIO
import pandas as pd
xml = """<catalog>
<product id="p1"><name>Notebook</name><price>4.50</price></product>
<product id="p2"><name>Pen</name><price>1.25</price></product>
</catalog>"""
df = pd.read_xml(StringIO(xml), xpath="./product")
if df is None or df.empty:
raise ValueError("No product records parsed")
required = {"id", "name", "price"}
missing = required - set(df.columns)
if missing:
raise ValueError(f"Missing expected fields: {sorted(missing)}")
df["price"] = pd.to_numeric(df["price"], errors="coerce")
if df["price"].isna().any():
raise ValueError("Invalid or missing price")
print(df.to_dict(orient="records"))
XML has many possible shapes. pandas says read_xml() works best for flatter, shallow structures; deeply nested data may require a stylesheet transformation to flatten it before it fits a tabular schema. Refer to the pandas XML I/O guide.
6. Normalize and validate the result
Parsing produces candidate values. Normalize them deliberately, then validate the contract your downstream code expects.
- Check that required fields exist and names match your schema.
- Convert types explicitly, handling invalid values rather than silently accepting them.
- Trim whitespace and normalize formats such as dates or currency according to your requirements.
- Decide how missing values and duplicate records should be represented.
- Compare several parsed records with their source pages and confirm the expected records were found.
- Keep source identifiers or URLs when they are needed for traceability.
These are workflow checks; the libraries do not automatically know your application schema or whether the extracted values match the source’s meaning.
7. Make recurring extraction maintainable
Page structures change. A selector can stop matching, a table can gain or lose a column, or content can move into a different structure. For recurring jobs, log the source and outcome, and alert on empty results, missing required fields, unexpected record counts, or conversion failures. Keep representative inputs so a selector or parser change can be checked against known cases.
Real pages may include navigation, ads, tracking scripts, and deeply nested elements around the target content. Parser differences on malformed HTML add another source of variation. Treat extraction rules as maintained code: review failures, inspect changed pages, and update selectors or transformations when the source evolves. If personal data is involved, account for privacy in collection, storage, and access.
8. Or skip the browser setup
If the data source is rendered in a web page and a visual capture is useful for review, archiving, or a downstream vision workflow, ScreenshotNeo can return a screenshot or PDF with one GET request. It captures a page image; it does not replace HTML or XML parsing when you need structured fields.
See the ScreenshotNeo API documentation for options. This cURL example saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/catalog \
-o catalog.webp
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
9. Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| No records or an empty DataFrame | Selector or XPath no longer matches, wrong page, or content is not in the input markup. | Inspect the exact response or XML, confirm the target structure, and update the selector or extraction method. |
| Different result with another parser | Malformed HTML is repaired differently by different parsers. | Choose a parser based on the tree it produces for representative inputs; add a regression check. |
read_html() returns a list |
That is the documented return shape, even for one table. | Inspect list length and select the intended DataFrame, for example tables[0]. |
| Wrong table selected | A page contains multiple tables or layout tables. | Inspect every table’s columns, shape, and sample rows before choosing. |
| Missing or unexpected columns | Source headers or structure changed. | Fail validation clearly, inspect the page, then update the field mapping. |
| Values remain strings or convert to missing values | Formatting, symbols, separators, or inconsistent source values do not match the assumed type. | Normalize the source representation explicitly and reject or report invalid values. |
| Nested XML does not map cleanly | The records are too deeply nested for a flat table mapping. | Transform or flatten the XML into the target record shape, then parse and validate it. |
| Request fails before parsing | Network failure, HTTP error, timeout, or an unexpected response. | Use a finite timeout, check the HTTP status, inspect the returned content, and handle retry policy in the surrounding job. |
10. Performance, reliability, and cost
The sources do not provide a controlled benchmark that supports universal speed rankings among these parsers. Keep the workflow proportional to the input: parse only the pages or tables needed, avoid unnecessary repeated requests, and measure your actual workload if throughput matters. Network time and source behavior can also affect a recurring extraction job.
Reliability comes from checking the output contract and monitoring changes, not from the act of parsing alone. These libraries are software; operational costs depend on your own infrastructure and workload. Respect applicable access rules and handle personal data carefully. A screenshot API can help capture a rendered page, but an image is not a structured dataset and may require separate interpretation.
FAQ
Should I use Beautiful Soup or pandas read_html()?
Use Beautiful Soup when you need selected elements or attributes from a page tree. Use read_html() when the target is an HTML table you want as a DataFrame.
Does parsing guarantee the data is correct?
No. Validate required fields, types, representative values, and expected records against the source.
Can pandas parse any XML into a flat table?
No. Its XML reader is best suited to flatter, shallow structures; deeply nested XML may need transformation first.
Why did my extraction stop working?
The source structure or content may have changed, or the parser may build a different tree from malformed markup. Inspect current input and validate selectors and fields.


