ScreenshotNeo

BlogHow-to

How to Extract Structured JSON Data from Websites

A practical guide to extracting JSON, JSON-LD, API responses, and rendered data reliably, with validation, provenance, and browser examples.

By the ScreenshotNeo team29 September 20268 min read

How to Extract Structured JSON Data from Websites

Structured data extraction works best as a decision sequence: use an official API when one exists, inspect the initial HTML for embedded JSON or JSON-LD, observe network responses on JavaScript-rendered pages, and use DOM selectors only as a fallback. Always validate the result and preserve enough provenance to reproduce it.

This guide shows each method with runnable Python, cURL, and Node.js examples, then covers JSON-LD, Schema.org, dynamic pages, pagination, authentication, validation, troubleshooting, performance, and cost.

1. Choose the extraction method

Method Use it when Strength Main risk
Official API The site documents an endpoint Stable fields, authentication, pagination, and error contracts Access limits or paid credentials
Embedded JSON or JSON-LD Data is present in the downloaded HTML Simple and fast; no browser required Markup can contain multiple shapes or stale values
Network observation Content appears after JavaScript runs Captures the application’s actual JSON payload Private endpoints can change and may have access rules
DOM extraction No usable API or payload exists Works for visible semantic content Selectors depend on presentation markup

Before writing code

  1. Check the site’s developer documentation and terms for an official API.
  2. Define the output schema: required fields, types, identifiers, and date format.
  3. Decide whether you need one record, every page of a collection, or a live rendered state.
  4. Record the source URL, retrieval time, extraction method, and a hash of the raw payload.
  5. Respect robots rules, authentication requirements, rate limits, and applicable law.
A practical extraction pipeline: API first, embedded data or network payload next, and DOM parsing as the fallback.
A practical extraction pipeline: API first, embedded data or network payload next, and DOM parsing as the fallback.

2. Fetch and inspect the initial HTML

Start with a normal HTTP request. Confirm the status and final URL before parsing; a login page or server error is not a data record.

Python: download HTML and locate structured blocks

import hashlib
import json
from datetime import datetime, timezone
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

url = "https://example.com/products/42"
response = requests.get(
    url,
    headers={"User-Agent": "data-extractor/1.0"},
    timeout=30,
    allow_redirects=True,
)
response.raise_for_status()

html = response.text
soup = BeautifulSoup(html, "html.parser")
blocks = []
for script in soup.select('script[type="application/ld+json"]'):
    try:
        blocks.append(json.loads(script.string or script.get_text()))
    except json.JSONDecodeError as exc:
        print(f"Invalid JSON-LD block: {exc}")

result = {
    "source_url": response.url,
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "method": "embedded-json-ld",
    "raw_sha256": hashlib.sha256(html.encode()).hexdigest(),
    "json_ld": blocks,
}
print(json.dumps(result, indent=2, ensure_ascii=False))

cURL: save the response for inspection

curl --fail --location --compressed \
  --user-agent 'data-extractor/1.0' \
  'https://example.com/products/42' \
  --output page.html

Node.js: fetch and parse JSON-LD

import * as cheerio from 'cheerio';

const url = 'https://example.com/products/42';
const res = await fetch(url, {
  headers: { 'user-agent': 'data-extractor/1.0' }
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
const $ = cheerio.load(html);
const blocks = [];
$('script[type="application/ld+json"]').each((_, node) => {
  try { blocks.push(JSON.parse($(node).text())); }
  catch (err) { console.error('Invalid JSON-LD:', err.message); }
});
console.log(JSON.stringify({ source_url: res.url, json_ld: blocks }, null, 2));

Inspect ordinary script tags as well. Many frameworks serialize application state under identifiers such as __NEXT_DATA__ or a state-specific data attribute. Treat those names as site-specific implementation details and add fixtures before depending on them.

3. Extract and normalize JSON-LD

JSON-LD is a JSON-based format for Linked Data designed to work with interoperable web programming environments. The W3C defines its syntax and processing algorithms in the JSON-LD 1.1 specification and JSON-LD 1.1 Processing Algorithms and API. Schema.org publishes machine-readable vocabulary definitions and a JSON-LD context at its developer documentation.

A block can be an object, an array, or an object containing @graph. Do not assume the first block is the product or article you need. Filter by @type, stable identifiers, and required properties.

def iter_jsonld(value):
    """Yield every node from object, array, and @graph forms."""
    if isinstance(value, list):
        for item in value:
            yield from iter_jsonld(item)
    elif isinstance(value, dict):
        graph = value.get("@graph")
        if isinstance(graph, list):
            for item in graph:
                yield from iter_jsonld(item)
        else:
            yield value

def first_product(blocks):
    for block in blocks:
        for node in iter_jsonld(block):
            types = node.get("@type", [])
            if isinstance(types, str):
                types = [types]
            if "Product" in types:
                return {
                    "id": node.get("@id"),
                    "name": node.get("name"),
                    "url": node.get("url"),
                    "image": node.get("image"),
                    "offers": node.get("offers"),
                }
    return None

Keep unknown properties until normalization is complete. A missing property, explicit null, and empty array have different meanings. If linked-data semantics matter across vocabularies, use a JSON-LD processor for expansion or compaction instead of manually rewriting context terms.

4. Find JSON on JavaScript-rendered pages

When the HTML has no useful payload, open the page in a browser and observe requests and responses. Playwright documents request, response, requestfinished, and requestfailed lifecycle events; these events help locate the JSON or GraphQL response that supplies the rendered view.

Python Playwright: capture JSON responses

import json
from playwright.sync_api import sync_playwright

url = "https://example.com/app"
with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    matches = []

    def on_response(response):
        content_type = response.headers.get("content-type", "")
        if "json" in content_type:
            try:
                matches.append({
                    "url": response.url,
                    "status": response.status,
                    "body": response.json(),
                })
            except Exception:
                pass

    page.on("response", on_response)
    page.goto(url, wait_until="networkidle", timeout=60_000)
    page.wait_for_timeout(1_000)
    browser.close()

print(json.dumps(matches, indent=2))

In production, narrow the listener by URL pattern, response status, and expected keys. Once you identify a documented and permitted endpoint, replaying it directly is usually faster and less brittle than scraping rendered text. Keep a browser fallback because private endpoint shapes can change.

5. DOM extraction as a fallback

Use semantic elements such as headings, links, time, and data attributes. Normalize whitespace and locale-specific numbers, but retain the original text when a conversion is uncertain.

from bs4 import BeautifulSoup
from decimal import Decimal
import re

soup = BeautifulSoup(html, "html.parser")
name = soup.select_one("h1")
price = soup.select_one("[data-price], .price")
raw_price = price.get("data-price") if price and price.has_attr("data-price") else (price.get_text(" ", strip=True) if price else None)

def parse_decimal(value):
    if not value:
        return None
    cleaned = re.sub(r"[^0-9,.-]", "", value)
    # Apply a site-specific locale rule; this example assumes a decimal point.
    try:
        return str(Decimal(cleaned.replace(",", "")))
    except Exception:
        return None

record = {"name": name.get_text(" ", strip=True) if name else None,
          "price": parse_decimal(raw_price),
          "price_raw": raw_price}

Add regression fixtures for representative pages. Presentation markup changes more often than a documented API, so alert on missing selectors rather than emitting silently incomplete records.

6. Pagination, authentication, and normalization

  • Pagination: follow the API’s cursor or page token until it is absent; stop on repeated cursors; deduplicate by a stable identifier.
  • Authentication: use the documented scheme, keep secrets out of logs, and never copy browser cookies into a shared worker without an explicit security design.
  • Types: validate required strings, numbers, booleans, arrays, and ISO dates before writing output.
  • Redirects: store both the requested and final URL, and reject unexpected domains when that matters.
  • Provenance: store source URL, retrieval timestamp, method, selector or request URL, status, and a raw payload hash.
from jsonschema import validate

schema = {
  "type": "object",
  "required": ["id", "name"],
  "properties": {
    "id": {"type": "string"},
    "name": {"type": "string"},
    "price": {"type": ["string", "null"]}
  },
  "additionalProperties": True
}
validate(instance=record, schema=schema)

7. Or skip the browser setup

ScreenshotNeo provides a website capture API when your workflow needs a rendered page artifact alongside extracted data. Its clean capture steps accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

Clean capture removes common consent and overlay elements before producing the page image.
Clean capture removes common consent and overlay elements before producing the page image.

See the ScreenshotNeo API documentation for all options. A one-call capture looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Use it when you need full-page or element captures, lazy images loaded, device presets, dark mode, custom viewport and retina scale, PDF output, custom CSS or JavaScript, clicks and waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture, or usage data. An MCP server also lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.

8. Troubleshooting

Symptom Likely cause Fix
JSON decode error HTML error page, truncated body, or malformed script Check status and content type, log a bounded body sample, and handle each block independently.
No JSON-LD found Data is client-rendered or uses Microdata/RDFa Inspect network responses, then parse semantic attributes as a fallback.
Only partial records Pagination or lazy loading was missed Follow cursors, wait for the required selector, and detect duplicate IDs.
Browser sees a challenge Bot check, login wall, or consent flow Use an official API or permitted authenticated session; do not attempt to bypass access controls.
Numbers are wrong Locale separators or currency text Parse with an explicit locale rule and preserve the raw value.
Selectors suddenly fail Presentation markup changed Prefer APIs or embedded data, add fixtures, and alert on missing required fields.

9. Performance, reliability, and cost

  • Prefer direct API calls: they avoid browser startup and transfer less data.
  • Cache immutable pages and use conditional requests where supported. Hash raw payloads to skip unchanged normalization.
  • Bound every timeout, retry only transient failures with exponential backoff, and cap concurrency to the site’s documented limits.
  • Separate fetch, parse, validate, and persist stages so a parser bug does not require refetching source pages.
  • Measure status distribution, parse failures, missing required fields, duplicate rates, pagination completion, and response sizes. No general accuracy or throughput benchmark should be assumed without a source-specific test.
  • For browser jobs, reuse a browser process, create isolated contexts, block unnecessary resources when permitted, and wait for a meaningful selector rather than an arbitrary long delay.

10. Short FAQ

Is JSON-LD always the same as the data visible on the page?

No. It can be incomplete, stale, or intended for search engines. Compare it with the rendered content and validate fields your application actually needs.

Should I expand JSON-LD?

Use JSON-LD expansion or compaction when linked-data context and identifiers matter across vocabularies. For a small application schema, deliberate field mapping is often simpler.

Can I scrape a private XHR endpoint?

Only when your access and the site’s terms permit it. Prefer documented endpoints and expect private contracts to change.

How do I make extraction auditable?

Persist the source and final URLs, retrieval time, method, request or selector, status, raw payload hash, parser version, and validation errors.

When should I use a screenshot API?

Use one when you need a reliable rendered artifact, visual verification, or browser features without maintaining Playwright infrastructure. ScreenshotNeo’s clean shots and verdict headers make billing and capture outcomes explicit.