ScreenshotNeo

BlogHow-to

How to Extract Structured Data From Web Pages

Extract JSON-LD, Microdata, RDFa, CSS and XPath data reliably, including JavaScript-rendered pages, with validation and provenance.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: fetch the page, preserve the raw response, extract semantic formats (JSON-LD, Microdata and RDFa), then use CSS or XPath selectors for fields that are only present in the document structure. If the data appears after JavaScript runs, render the page in a headless browser or use a hosted screenshot/rendering service. Normalize every value, validate it against visible content and your target schema, and store field-level provenance.

What “structured data” means

Structured data has two layers:

  • Vocabulary: the meaning of entities and properties, often provided by Schema.org.
  • Encoding: the way that vocabulary is embedded in a page: JSON-LD, Microdata or RDFa.

CSS and XPath selectors read document structure. Semantic formats describe entities and relationships, so they are usually a better first choice for products, articles, events and people. Always check that extracted values agree with what a visitor can see.

Choose the extraction method

Method Use it when Strengths Limitations
JSON-LD The page exposes <script type="application/ld+json"> Semantic, easy to parse, supports graphs and relationships Can be duplicated, malformed or stale
Microdata Properties are marked with itemscope, itemtype and itemprop Associates values directly with elements Nested items require careful traversal
RDFa Semantic attributes such as typeof, property and resource are present Rich relationships and links More complex context handling
CSS selectors Values exist in stable classes, IDs or elements Readable and convenient Breaks when presentation markup changes
XPath You need ancestors, siblings or exact text nodes Precise structural navigation Harder to maintain and read
Browser rendering JavaScript injects the required data Sees the rendered DOM and can perform interactions Slower and more operationally complex

Step 1: Fetch and classify the response

Record the URL, retrieval time, status, content type and raw bytes before parsing. Treat HTML, XML, JSON, JavaScript, images and PDFs differently. A successful HTTP status only proves that a response arrived; it does not prove that the desired data is present.

import requests

url = "https://example.com/article"
r = requests.get(
    url,
    headers={"User-Agent": "structured-data-extractor/1.0"},
    timeout=30,
)
r.raise_for_status()
content_type = r.headers.get("content-type", "").lower()
raw = r.content

print({
    "url": r.url,
    "status": r.status_code,
    "content_type": content_type,
    "bytes": len(raw),
})

if "json" in content_type:
    document = r.json()
elif "html" in content_type or "xml" in content_type or not content_type:
    document = raw.decode(r.encoding or "utf-8", errors="replace")
else:
    raise ValueError(f"Unsupported content type: {content_type}")

Step 2: Extract JSON-LD first

JSON-LD is often the most stable source for semantic entities. A block may contain one object, an array or an object with an @graph. Parse all blocks, retain their context and type, and do not assume that the first block is the right entity.

import json
from bs4 import BeautifulSoup

soup = BeautifulSoup(document, "html.parser")
jsonld = []
for node in soup.select('script[type="application/ld+json"]'):
    text = node.string or node.get_text()
    try:
        value = json.loads(text)
    except json.JSONDecodeError as exc:
        print("Skipping malformed JSON-LD:", exc)
        continue
    if isinstance(value, list):
        jsonld.extend(value)
    elif isinstance(value, dict) and isinstance(value.get("@graph"), list):
        jsonld.extend(value["@graph"])
    else:
        jsonld.append(value)

for entity in jsonld:
    print(entity.get("@type"), entity.get("@id"), entity)

For production use, select entities by @type and stable identifiers instead of position. Handle @type arrays, language maps, nested objects and relative URLs.

Step 3: Extract Microdata and RDFa

Microdata places properties on ordinary elements. RDFa uses attributes that can express subjects, predicates and objects. Extract every available graph before falling back to presentation selectors. The W3C RDFa API specification describes programmatic extraction of RDFa information.

from urllib.parse import urljoin

# Simple Microdata traversal for one item
item = soup.select_one('[itemscope]')
if item:
    item_type = item.get("itemtype")
    fields = {}
    for node in item.select('[itemprop]'):
        name = node.get("itemprop")
        if node.has_attr("content"):
            value = node["content"]
        elif node.name in {"meta"}:
            value = node.get("content", "")
        elif node.name in {"img", "audio", "video", "source"}:
            value = node.get("src", "")
        elif node.name == "a":
            value = node.get("href", "")
        elif node.name == "time":
            value = node.get("datetime") or node.get_text(" ", strip=True)
        else:
            value = node.get_text(" ", strip=True)
        fields.setdefault(name, []).append(value)
    print({"type": item_type, "properties": fields})

# Basic RDFa attributes
for node in soup.select('[property]'):
    predicate = node.get("property")
    value = node.get("content") or node.get("resource") or node.get_text(" ", strip=True)
    subject = node.get("about")
    print({"subject": subject, "predicate": predicate, "value": value})

For nested Microdata, recurse into child itemscope nodes and attach them to the property named by their nearest itemprop. For RDFa, resolve about, resource and vocabulary prefixes into absolute identifiers.

Step 4: Use CSS selectors and XPath as fallbacks

Selectors are appropriate when the required value is visible but has no semantic annotation. BeautifulSoup is tolerant of imperfect markup; lxml provides a fast ElementTree-style HTML/XML API and XPath support.

from lxml import html

root = html.fromstring(document)

title_css = root.cssselect("h1")
title = title_css[0].text_content().strip() if title_css else None

price_nodes = root.xpath("//*[contains(@class, 'price')]
")
price = price_nodes[0].text_content().strip() if price_nodes else None

print({"title": title, "price": price})

Keep selectors narrow and explain why each one exists. Prefer an ID or a stable data attribute over a generated CSS class. Use XPath when you need relationships such as “the price beside this product name” or an exact text node.

Step 5: Extract data created by JavaScript

There are three common cases:

  1. The initial HTML contains a JSON state object. Search script contents for serialized state and parse it safely.
  2. The page calls a JSON endpoint. Inspect the browser’s network requests and call that endpoint directly when permitted.
  3. The page creates markup only after JavaScript executes or after an interaction. Use a headless browser, wait for a selector or network idle, then parse the rendered DOM.
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/products", wait_until="networkidle", timeout=60_000)
    page.wait_for_selector("[data-product]", timeout=15_000)
    rendered_html = page.content()
    products = page.locator("[data-product]").evaluate_all(
        "els => els.map(e => ({name: e.innerText.trim(), id: e.dataset.product}))"
    )
    print(products)
    browser.close()

Do not use arbitrary long sleeps as your only wait condition. Prefer a required selector, a specific response, or network idle, with a maximum timeout and a diagnostic capture when it fails.

Step 6: Normalize, validate and preserve provenance

Convert extracted values into a typed record. Normalize dates to one timezone-aware representation, numbers to numeric types, and URLs to absolute URLs. Deduplicate repeated entities using stable IDs. Validate required properties, syntax, expected types and conflicts between semantic markup and visible text.

from datetime import datetime, timezone
from urllib.parse import urljoin
from decimal import Decimal

record = {
    "url": url,
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "title": title,
    "price": Decimal("19.99") if price else None,
    "source": {
        "title": {"method": "css", "selector": "h1", "raw": title},
        "price": {"method": "xpath", "selector": "//*[contains(@class, 'price')]", "raw": price},
    },
}

errors = []
if not record["title"]:
    errors.append("missing title")
if record["price"] is not None and record["price"] < 0:
    errors.append("negative price")
record["validation_errors"] = errors

Store the source URL, retrieval time, selector or JSON path, original value, normalized value and parser version for every field. This makes template changes diagnosable instead of mysterious.

Build a resilient extraction pipeline

  1. Save the raw response and metadata.
  2. Classify the content type.
  3. Parse JSON-LD, Microdata and RDFa in a defined order.
  4. Merge entities by stable IDs, retaining conflicts for review.
  5. Apply CSS and XPath fallbacks.
  6. Render with a browser only when required fields are still missing.
  7. Normalize and validate against your target schema.
  8. Emit typed records, provenance and validation errors.
  9. Keep regression fixtures for representative page templates.
  10. Monitor extraction completeness and alert on sudden changes.

Common errors and fixes

Error Likely cause Fix
No JSON-LD found Data is Microdata, RDFa or rendered later Inspect all three formats, then render the page or locate its JSON endpoint.
JSON decode error Malformed block, HTML comment or multiple objects Log the block, skip it safely, and parse each valid object separately.
Empty selector result Wrong template, changed class or JavaScript-only markup Check the raw HTML, use stable attributes, or render before selecting.
Duplicate entities Several JSON-LD blocks describe the same item Merge by @id or another stable key and retain source locations.
Values disagree Stale annotations, hidden offers or multiple variants Compare with visible text, record the conflict and apply an explicit precedence rule.
Relative URLs href or src is not absolute Resolve with urljoin(page_url, value).
Timeout in browser Slow resource, bot check or never-ending request Set bounded timeouts, wait for a specific selector, block unnecessary resources and save diagnostics.
403 or CAPTCHA Access controls or automated-traffic detection Respect the site’s rules, authenticate where authorized, slow requests and do not attempt to bypass protections.

Performance, reliability and cost

  • Direct HTTP parsing is usually faster and cheaper than launching a browser. Try it first.
  • Parse only the formats and selectors required for your schema, but retain the raw response for replay.
  • Reuse browser processes, limit concurrency and cache immutable pages.
  • Use bounded connect, read and navigation timeouts with retries only for transient failures. Avoid retry storms.
  • Record status, content type, parser version and validation counts so regressions are visible.
  • Rendering adds compute and operational cost. A hosted API can be simpler when pages require JavaScript, interaction, cookie handling or many URLs.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. Its capture flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

Use the same rendered page as an inspection artifact while your extractor reads JSON-LD, Microdata, RDFa or visible elements:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for the 63 options, including full-page capture, CSS element capture, custom JavaScript and CSS, waits, headers, cookies, resource blocking, caching, signed links, asynchronous jobs, bulk capture and usage reporting. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Should I choose CSS, XPath or JSON-LD?

Use JSON-LD, Microdata or RDFa for semantic entities, then CSS or XPath for fields that are not annotated. XPath is useful for relationships and precise text selection.

Can I extract data without a browser?

Yes, when the required data is in the server response or a directly accessible JSON endpoint. Render the page when JavaScript creates the data or interaction is required.

How do I know an extraction is correct?

Validate types and required properties, compare semantic values with visible text, detect conflicts and retain field-level provenance.

What should I save for debugging?

Save the raw response, URL, retrieval time, content type, parser version, selector or JSON path, normalized record and validation errors.