ScreenshotNeo

BlogHow-to

HTML Table Capture with Python

Capture HTML tables with Python using pandas or Beautiful Soup, handle dynamic pages, clean results, and troubleshoot extraction failures.

By the ScreenshotNeo team29 September 20268 min read

HTML Table Capture with Python

To read an ordinary HTML table into Python, start with pandas.read_html(). It returns a list of DataFrames, so inspect that list and select the table you actually want. Use Beautiful Soup when you need custom table selection, links, attributes, or cell-by-cell rules.

This guide covers static pages, multiple tables, malformed markup, JavaScript-rendered tables, parser choices, cleanup, reliability, and a browser-free capture option.

1. Choose the right extraction path

Situation Recommended tool Why
Conventional table, DataFrame output pandas.read_html Reads table markup directly and handles headers, spans, numbers, and missing cells.
Several tables on one page read_html with match or attrs Filters by distinctive text or attributes such as an id.
Need links, data attributes, or custom row rules Beautiful Soup You control every element and can preserve attributes.
Table appears only after JavaScript runs Render the page first, then parse its HTML A plain HTTP response may contain no rows at all.

The pandas API describes the core operation as reading HTML tables into a list of DataFrame objects. Never assume index zero is correct when a page contains navigation, pricing, or hidden tables.

Use pandas for direct table-to-DataFrame conversion, then verify which table was selected.
Use pandas for direct table-to-DataFrame conversion, then verify which table was selected.

2. Install Python dependencies

python -m pip install pandas lxml beautifulsoup4 html5lib requests

pandas is required for DataFrames. lxml is usually the fastest parser. html5lib is more tolerant of broken markup but slower. Beautiful Soup’s built-in html.parser needs no extra parser package; explicitly naming the parser makes behavior reproducible. Pandas can try lxml and fall back to Beautiful Soup plus html5lib when needed.

3. Read the first table with pandas

import pandas as pd

url = "https://example.com/table-page"
tables = pd.read_html(url)

print(f"Found {len(tables)} table(s)")
for i, table in enumerate(tables):
    print(i, table.shape)
    print(table.head(3))

# Use index 0 only after confirming it is the intended table
df = tables[0]
print(df.to_json(orient="records"))

The URL may also be a local path, HTML string, or file-like object. Keep the list until you have checked table count, column names, and representative values.

Select by visible text with match

import pandas as pd

tables = pd.read_html(
    "https://example.com/prices",
    match="Annual price",
)
prices = tables[0]

match is useful when an identifying phrase appears inside the target table. Choose text that is distinctive but stable; a marketing sentence shared by several tables can still return more than one result.

Select by attributes with attrs

tables = pd.read_html(
    "https://example.com/results",
    attrs={"id": "results-table"},
)
results = tables[0]

The attribute must describe an actual <table> element. If the site changes the id, inspect the HTML and update the selector.

4. Control headers, skipped rows, and number formats

import pandas as pd

df = pd.read_html(
    "https://example.com/report",
    header=0,
    skiprows=[1],
    thousands=",",
    decimal=".",
    na_values=["—", "N/A"],
    converters={"Account": str},
)[0]

print(df.columns.tolist())
print(df.dtypes)

Use header=None when the page has no true header row, then assign names yourself. For multi-row headers, pass a list of header row indexes and verify the resulting MultiIndex. rowspan and colspan can create repeated or hierarchical labels; inspect columns before exporting.

If the page uses a comma decimal mark, set decimal="," and choose the appropriate thousands separator. Locale symbols, currency signs, footnote markers, and non-breaking spaces may still need a cleanup pass.

5. Build a reusable extraction function

from __future__ import annotations
import pandas as pd
from collections.abc import Callable

def read_table(url: str, *, match: str | None = None,
               attrs: dict | None = None,
               parser: str = "lxml",
               clean: Callable[[pd.DataFrame], pd.DataFrame] | None = None
               ) -> pd.DataFrame:
    tables = pd.read_html(url, match=match, attrs=attrs, flavor=parser)
    if not tables:
        raise ValueError("No HTML tables matched the selection")
    if len(tables) > 1:
        raise ValueError(f"Selection matched {len(tables)} tables; narrow match or attrs")
    frame = tables[0].copy()
    if clean:
        frame = clean(frame)
    return frame

def clean_report(frame: pd.DataFrame) -> pd.DataFrame:
    frame.columns = [str(c).strip() for c in frame.columns]
    return frame.dropna(how="all")

report = read_table(
    "https://example.com/report",
    attrs={"id": "results-table"},
    clean=clean_report,
)
print(report.head())

Failing when selection is ambiguous is safer than silently taking the first match in a scheduled job. Pin the parser in production and log the URL, table count, columns, row count, and a sample row.

6. Use Beautiful Soup for custom traversal

from bs4 import BeautifulSoup
import requests

url = "https://example.com/catalog"
html = requests.get(url, timeout=30).text
soup = BeautifulSoup(html, "html.parser")
table = soup.select_one("table#catalog")
if table is None:
    raise ValueError("Target table was not found")

records = []
for tr in table.select("tr"):
    cells = tr.select(":scope > th, :scope > td")
    values = [cell.get_text(" ", strip=True) for cell in cells]
    if values:
        records.append(values)

headers = records[0]
rows = [dict(zip(headers, row)) for row in records[1:]]
print(rows[:3])

Use lxml or html5lib in the constructor when required: BeautifulSoup(html, "lxml"). Invalid HTML can produce different trees with different parsers, so verify the selected table and a few rows against the source page. For links, read the anchor’s href instead of discarding the element.

7. JavaScript-rendered tables and capture choices

requests and read_html parse the response they receive; they do not execute page JavaScript. If the response contains an empty table shell, find the site’s JSON endpoint where permitted, or render the page in a browser and pass the resulting HTML to pandas or Beautiful Soup. Wait for a selector that identifies the rows rather than sleeping for an arbitrary duration.

When you render, account for consent banners, newsletter popups, chat widgets, authentication, lazy-loaded rows, and bot checks. Save the final HTML during debugging so you can distinguish a selector error from a page that never finished loading.

8. Validate and normalize the DataFrame

  • Confirm the intended table id or identifying text.
  • Check column names, row count, and three representative cells.
  • Look for all-null columns caused by header or colspan interpretation.
  • Normalize whitespace and non-breaking spaces.
  • Parse dates and numeric columns explicitly after removing currency or footnote characters.
  • Check missing values and duplicate rows.
  • If links matter, preserve href attributes during a Beautiful Soup pass.
df.columns = [" ".join(str(c).split()) for c in df.columns]
df["amount"] = (
    df["amount"].astype("string")
      .str.replace(r"[^0-9,.-]", "", regex=True)
      .str.replace(",", "", regex=False)
      .astype("Float64")
)
assert len(df) > 0
print(df.shape, df.isna().sum())

Extraction is not validation. A successful parse can still select the wrong table or shift cells because of malformed markup.

9. Troubleshooting common failures

Symptom Likely cause Fix
No tables found Rows are injected by JavaScript, selector is wrong, or response is blocked. Save and inspect response HTML; use an API or browser rendering; verify match/attrs.
Wrong table returned Index zero was assumed or match text is broad. Print every table’s shape and columns; narrow with a table id or stronger text.
Parser error Missing optional dependency or invalid markup. Install lxml/html5lib and try an explicit flavor; compare results.
Headers become Unnamed or NaN Title rows, colspan, or skipped notes were interpreted as headers. Set header/skiprows, then rename columns and inspect.
Numbers remain strings Currency symbols, separators, or locale decimals. Configure thousands/decimal and clean with a converter.
Rows are missing Pagination, lazy loading, or virtualized rendering. Find the next-page endpoint or render and wait for the row selector.
403, 429, or timeout Rate limiting, bot protection, or slow origin. Use a clear User-Agent where allowed, back off, cache responses, and respect site terms.

10. Performance, reliability, and cost

For repeated jobs, download once with requests, store the HTML, and parse locally. This separates network failures from parser failures and avoids fetching unchanged pages. Keep timeouts finite, retry only transient failures with exponential backoff, and cap concurrency so you do not trigger rate limits. Hash the raw HTML and compare row counts or schema to detect silent layout changes.

Parser choice is a trade-off: lxml is generally fast but less predictable on invalid markup; html5lib is lenient but slower; html.parser has no external dependency but can differ on malformed input. Benchmark your own pages if throughput matters; the documentation does not provide a universal speed figure.

Library parsing has no ScreenshotNeo charge, but browser rendering, proxies, and hosted capture services can add cost. Cache stable HTML and avoid recapturing unchanged URLs. Treat extracted data as untrusted input: validate types, limit response sizes, and never execute arbitrary page scripts in your data process.

11. Or skip the browser setup

If the table lives behind consent UI, popups, chat widgets, slow rendering, or bot checks, ScreenshotNeo can produce a clean page image or PDF before downstream review. Its API accepts one GET request; see the ScreenshotNeo documentation for all options.

A rendered capture can remove consent UI before downstream review when raw HTTP HTML is incomplete.
A rendered capture can remove consent UI before downstream review when raw HTTP HTML is incomplete.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, and cache hits are never billed; response headers report the page verdict and whether it was billed. You can capture a selected element, use full-page lazy-image loading, set a viewport or device preset, wait for a selector or network idle, add custom CSS or JavaScript, block resources, provide headers/cookies/authentication, and choose PNG, JPEG, WebP, or PDF. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

12. Short FAQ

Does read_html return one DataFrame?

No. It returns a list. Select deliberately after checking the list length and contents.

Can Beautiful Soup replace pandas?

Yes for custom traversal, but you must design headers, row alignment, type conversion, and missing-value handling yourself.

Which parser should I deploy?

Pin lxml for speed on valid markup, html5lib for tolerant parsing, or html.parser when avoiding external dependencies. Verify output on your pages.

Why did my script find no rows?

The rows may be generated after JavaScript executes, paginated, or blocked. Inspect the raw response before changing selectors.

How do I avoid silently ingesting a changed table?

Assert the expected columns and minimum row count, log a sample, and alert when the schema or HTML hash changes.