ScreenshotNeo

BlogGuides

Data Extraction in Python: Files, APIs, HTML, XML, and Web Pages

Learn a practical Python workflow for extracting data from CSV, JSON, HTML, XML, APIs, and web pages, with runnable code and troubleshooting.

By the ScreenshotNeo team29 September 20268 min read

Data Extraction in Python: Files, APIs, HTML, XML, and Web Pages

How do you extract data in Python? Start by identifying the source and its format, retrieve remote content when necessary, validate the response, parse it with a format-appropriate tool, normalize the fields, and then save or analyze the result. Use the standard library for simple CSV, JSON, HTML, and XML work; use Requests for HTTP; Beautiful Soup for flexible HTML parsing; and pandas when the result should become a DataFrame.

This separation matters. Retrieval gets bytes from a file or server. Parsing turns those bytes into Python objects. Downstream analysis cleans, validates, joins, aggregates, or exports the parsed data. Keeping those steps separate makes failures easier to diagnose and lets you change one part without rewriting the whole pipeline.

This guide covers local files, APIs, HTML and XML pages, large datasets, browser-rendered pages, and a managed screenshot option from ScreenshotNeo.

1. Choose an extraction method

Input Good starting point Choose it when
CSV or fixed-width text csv or pandas.read_csv()/read_fwf() You need rows and columns, either as lists/dicts or a DataFrame.
JSON file or API json, Requests, or pandas.read_json() The source exposes structured objects or arrays.
HTML/XML markup html.parser, xml.etree.ElementTree, or Beautiful Soup You need elements, attributes, links, or text from markup.
Remote API or page Requests You need HTTP status handling, headers, cookies, timeouts, or decoded content.
Tabular analysis pandas readers The next step is filtering, joining, grouping, or exporting DataFrames.

Python includes HTML and XML processing interfaces in its standard library, so a third-party dependency is not required for every markup task. See the Python structured markup documentation. pandas documents readers for CSV, fixed-width text, JSON, HTML, XML, and Excel in its I/O tools guide.

Choose the parser after identifying the source and its format.
Choose the parser after identifying the source and its format.

2. A reliable extraction workflow

  1. Identify the source and format. Confirm whether you have a local file, API response, HTML page, XML document, or browser-rendered application.
  2. Retrieve remote data. Set a timeout and provide required headers or authentication.
  3. Validate the response. Check the HTTP status before decoding JSON or parsing markup.
  4. Parse the actual format. Do not parse an HTML error page as JSON, or assume every table is a CSV.
  5. Normalize fields. Convert dates, numbers, missing values, and nested records deliberately.
  6. Validate the result. Check required columns, row counts, types, and duplicate keys.
  7. Persist or analyze. Write a clean JSON/CSV file or pass a DataFrame to the next operation.

3. Extract data from CSV and fixed-width files

Standard library CSV

import csv
from pathlib import Path

rows = []
with Path("sales.csv").open("r", encoding="utf-8", newline="") as f:
    reader = csv.DictReader(f)
    required = {"order_id", "amount"}
    if not required.issubset(reader.fieldnames or []):
        raise ValueError(f"Missing columns: {required - set(reader.fieldnames or [])}")
    for row in reader:
        row["amount"] = float(row["amount"])
        rows.append(row)

print(rows[:3])

csv.DictReader keeps the dependency footprint small and lets you stream rows instead of loading a whole file. Treat every CSV field as text until you convert it; empty strings, decimal separators, and date formats vary between producers.

pandas for tabular work

import pandas as pd

df = pd.read_csv("sales.csv", usecols=["order_id", "amount", "created_at"])
df["amount"] = pd.to_numeric(df["amount"], errors="coerce")
df["created_at"] = pd.to_datetime(df["created_at"], errors="coerce", utc=True)
df = df.dropna(subset=["order_id", "amount"])
print(df.head())

Use usecols, dtype, and chunksize to control memory. For fixed-width input, use pd.read_fwf() and provide column widths or a schema.

4. Extract JSON from files and APIs

Local JSON

import json
from pathlib import Path

with Path("records.json").open(encoding="utf-8") as f:
    payload = json.load(f)

records = payload["records"] if isinstance(payload, dict) else payload
for record in records:
    print(record.get("id"), record.get("name"))

HTTP JSON with Requests

import requests

url = "https://api.example.com/v1/items"
r = requests.get(url, params={"limit": 100}, timeout=30)
r.raise_for_status()                 # Check HTTP success first
try:
    data = r.json()
except ValueError as exc:
    raise RuntimeError("The response was not valid JSON") from exc

items = data["items"]
print(len(items))

Requests provides connection pooling, automatic content decoding, and timeout support; consult its official documentation. A successful call to r.json() only means the body was valid JSON. It does not prove that the request succeeded, so call raise_for_status() or inspect r.status_code first. The Requests quickstart describes this distinction.

Pagination and retries

import time
import requests

session = requests.Session()
all_items = []
url = "https://api.example.com/v1/items"
for page in range(1, 11):
    response = session.get(url, params={"page": page, "per_page": 100}, timeout=30)
    if response.status_code == 429:
        time.sleep(int(response.headers.get("Retry-After", "5")))
        continue
    response.raise_for_status()
    body = response.json()
    all_items.extend(body.get("items", []))
    if not body.get("next_page"):
        break

Bound retries, respect documented rate limits, and make writes idempotent if a page can be fetched twice. Store the last successful cursor or page so a long extraction can resume.

5. Parse HTML and XML

Beautiful Soup for HTML

from bs4 import BeautifulSoup
import requests

r = requests.get("https://example.com/news", timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
articles = []
for node in soup.select("article"):
    title = node.select_one("h2")
    link = node.select_one("a[href]")
    if title and link:
        articles.append({"title": title.get_text(" ", strip=True),
                         "url": link["href"]})
print(articles)

Beautiful Soup parses HTML and XML. Always specify a parser such as html.parser so behavior is reproducible across machines; the Beautiful Soup documentation explains parser choices. CSS selectors are convenient, but selectors tied to presentation classes can break when a site redesigns. Prefer stable attributes or semantic structure where available.

Standard-library XML

import xml.etree.ElementTree as ET

root = ET.parse("catalog.xml").getroot()
for product in root.findall(".//product"):
    sku = product.findtext("sku")
    price = product.findtext("price")
    print(sku, price)

For very large XML files, pandas documents memory-efficient iterparse approaches. Streaming is preferable when the complete tree would exceed available memory.

6. Extract HTML tables with pandas

import pandas as pd

frames = pd.read_html("https://example.com/table-page")
if not frames:
    raise ValueError("No HTML tables found")
df = frames[0]
print(df.columns)
print(df.dropna(how="all"))

read_html() may require an HTML parsing dependency, and pages that build tables with JavaScript may return no table in the initial response. In that case, locate the underlying API, obtain the rendered HTML with a browser, or use a screenshot only when the goal is visual output rather than structured values.

7. JavaScript-rendered pages and browser automation

Requests downloads the server response; it does not execute page JavaScript. If the data appears only after scripts run, identify an XHR/fetch endpoint in the browser’s network panel and call that endpoint directly when permitted. If no usable endpoint exists, a browser automation tool can wait for a selector, click controls, set cookies, and extract the rendered DOM. Keep waits explicit and capture the final HTML for debugging.

Web extraction rules depend on the target, its terms, the data, and your jurisdiction. Check the site’s terms and applicable requirements before collecting data, and avoid bypassing access controls.

8. Or skip the browser setup

When your deliverable is a dependable visual capture of a page, ScreenshotNeo’s API handles the capture request directly:

A capture pipeline can remove common overlays before producing the image.
A capture pipeline can remove common overlays before producing the image.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether it was billed. The MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Free accounts include 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

9. Options that affect extraction quality

  • Encoding: pass an explicit encoding for local files and inspect response.encoding for HTTP.
  • Authentication: use API keys, custom headers, cookies, or an Authorization header as required by the source.
  • Selectors: validate that a selector returns the expected number of nodes before parsing fields.
  • Missing data: distinguish absent fields, empty strings, null JSON values, and conversion failures.
  • Locale: normalize decimal separators, dates, timezone, and currency before analysis.
  • Size: stream CSV/XML, request bounded pages, select only needed columns, and avoid loading giant responses into memory.
  • Reproducibility: pin dependencies and explicitly select parsers; parser availability can otherwise change results.

10. Troubleshooting

Symptom Likely cause Fix
JSONDecodeError HTML error page, empty body, or invalid JSON. Check status, content type, and a short body sample before .json().
HTTP 401/403 Missing credentials or permission. Use the documented authentication method; do not assume a browser cookie grants API access.
HTTP 429 Rate limit. Honor Retry-After, reduce concurrency, and resume from a cursor.
No HTML elements found Wrong selector or JavaScript-rendered content. Inspect the raw response, verify the selector, and locate the underlying API or rendered DOM.
Parser mismatch Different parser installed on another machine. Specify html.parser, lxml, or another intentional parser and pin dependencies.
Out-of-memory Whole-file load or huge XML tree. Stream rows, use chunks, select columns, or process XML incrementally.
Wrong numeric/date values Locale or mixed formats. Normalize explicitly and use coercion only when invalid values are handled and counted.

11. Performance, reliability, and cost

Reuse a Requests Session for many HTTP calls, set connect and read timeouts, and cap concurrency to the source’s limits. Cache immutable responses locally with a source and timestamp. Record URL, status, parser version, row count, and extraction time so a later discrepancy is explainable.

For pandas, narrow columns early and use chunks for files that do not fit comfortably in memory. For APIs, pagination and retries dominate runtime more often than JSON parsing. For browser-rendered pages, waits, assets, and JavaScript execution are the expensive steps; target a stable endpoint when possible.

ScreenshotNeo supports caching with a TTL you choose, bulk capture of up to 100 URLs per call, asynchronous jobs with signed webhooks, custom CSS and JavaScript, device presets, full-page or element capture, PDFs, and signed links. Its billing counts only clean shots; cache hits and failed captures are identified in response headers and cost nothing. Choose the capture mode and cache policy that match your refresh frequency.

12. FAQ

Should I use pandas or the standard library?

Use the standard library when you need a small dependency-free script or streaming row-by-row logic. Use pandas when the result is a DataFrame and you need analysis operations.

Can Requests scrape any website?

Requests can retrieve an HTTP response when the server permits it. It does not execute JavaScript, solve bot challenges, or override a site’s access rules.

Why did JSON parsing succeed on an error?

An error response can still contain valid JSON. Check the status code or call raise_for_status() before decoding.

How do I make HTML extraction stable?

Specify the parser, prefer semantic or stable attributes, validate selectors, pin dependencies, and keep fixtures of representative pages for regression checks.

When is a screenshot useful?

A screenshot is useful for visual archives, previews, PDF workflows, and page-state verification. It is not a substitute for structured extraction when you need values for analysis.