How to Extract Structured JSON Data from Websites
A practical guide to extracting JSON, JSON-LD, API responses, and rendered data reliably, with validation, provenance, and browser examples.

Structured data extraction works best as a decision sequence: use an official API when one exists, inspect the initial HTML for embedded JSON or JSON-LD, observe network responses on JavaScript-rendered pages, and use DOM selectors only as a fallback. Always validate the result and preserve enough provenance to reproduce it.
This guide shows each method with runnable Python, cURL, and Node.js examples, then covers JSON-LD, Schema.org, dynamic pages, pagination, authentication, validation, troubleshooting, performance, and cost.
1. Choose the extraction method
| Method | Use it when | Strength | Main risk |
|---|---|---|---|
| Official API | The site documents an endpoint | Stable fields, authentication, pagination, and error contracts | Access limits or paid credentials |
| Embedded JSON or JSON-LD | Data is present in the downloaded HTML | Simple and fast; no browser required | Markup can contain multiple shapes or stale values |
| Network observation | Content appears after JavaScript runs | Captures the application’s actual JSON payload | Private endpoints can change and may have access rules |
| DOM extraction | No usable API or payload exists | Works for visible semantic content | Selectors depend on presentation markup |
Before writing code
- Check the site’s developer documentation and terms for an official API.
- Define the output schema: required fields, types, identifiers, and date format.
- Decide whether you need one record, every page of a collection, or a live rendered state.
- Record the source URL, retrieval time, extraction method, and a hash of the raw payload.
- Respect robots rules, authentication requirements, rate limits, and applicable law.

2. Fetch and inspect the initial HTML
Start with a normal HTTP request. Confirm the status and final URL before parsing; a login page or server error is not a data record.
Python: download HTML and locate structured blocks
import hashlib
import json
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
url = "https://example.com/products/42"
response = requests.get(
url,
headers={"User-Agent": "data-extractor/1.0"},
timeout=30,
allow_redirects=True,
)
response.raise_for_status()
html = response.text
soup = BeautifulSoup(html, "html.parser")
blocks = []
for script in soup.select('script[type="application/ld+json"]'):
try:
blocks.append(json.loads(script.string or script.get_text()))
except json.JSONDecodeError as exc:
print(f"Invalid JSON-LD block: {exc}")
result = {
"source_url": response.url,
"retrieved_at": datetime.now(timezone.utc).isoformat(),
"method": "embedded-json-ld",
"raw_sha256": hashlib.sha256(html.encode()).hexdigest(),
"json_ld": blocks,
}
print(json.dumps(result, indent=2, ensure_ascii=False))
cURL: save the response for inspection
curl --fail --location --compressed \
--user-agent 'data-extractor/1.0' \
'https://example.com/products/42' \
--output page.html
Node.js: fetch and parse JSON-LD
import * as cheerio from 'cheerio';
const url = 'https://example.com/products/42';
const res = await fetch(url, {
headers: { 'user-agent': 'data-extractor/1.0' }
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
const $ = cheerio.load(html);
const blocks = [];
$('script[type="application/ld+json"]').each((_, node) => {
try { blocks.push(JSON.parse($(node).text())); }
catch (err) { console.error('Invalid JSON-LD:', err.message); }
});
console.log(JSON.stringify({ source_url: res.url, json_ld: blocks }, null, 2));
Inspect ordinary script tags as well. Many frameworks serialize application state under identifiers such as __NEXT_DATA__ or a state-specific data attribute. Treat those names as site-specific implementation details and add fixtures before depending on them.
3. Extract and normalize JSON-LD
JSON-LD is a JSON-based format for Linked Data designed to work with interoperable web programming environments. The W3C defines its syntax and processing algorithms in the JSON-LD 1.1 specification and JSON-LD 1.1 Processing Algorithms and API. Schema.org publishes machine-readable vocabulary definitions and a JSON-LD context at its developer documentation.
A block can be an object, an array, or an object containing @graph. Do not assume the first block is the product or article you need. Filter by @type, stable identifiers, and required properties.
def iter_jsonld(value):
"""Yield every node from object, array, and @graph forms."""
if isinstance(value, list):
for item in value:
yield from iter_jsonld(item)
elif isinstance(value, dict):
graph = value.get("@graph")
if isinstance(graph, list):
for item in graph:
yield from iter_jsonld(item)
else:
yield value
def first_product(blocks):
for block in blocks:
for node in iter_jsonld(block):
types = node.get("@type", [])
if isinstance(types, str):
types = [types]
if "Product" in types:
return {
"id": node.get("@id"),
"name": node.get("name"),
"url": node.get("url"),
"image": node.get("image"),
"offers": node.get("offers"),
}
return None
Keep unknown properties until normalization is complete. A missing property, explicit null, and empty array have different meanings. If linked-data semantics matter across vocabularies, use a JSON-LD processor for expansion or compaction instead of manually rewriting context terms.
4. Find JSON on JavaScript-rendered pages
When the HTML has no useful payload, open the page in a browser and observe requests and responses. Playwright documents request, response, requestfinished, and requestfailed lifecycle events; these events help locate the JSON or GraphQL response that supplies the rendered view.
Python Playwright: capture JSON responses
import json
from playwright.sync_api import sync_playwright
url = "https://example.com/app"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
matches = []
def on_response(response):
content_type = response.headers.get("content-type", "")
if "json" in content_type:
try:
matches.append({
"url": response.url,
"status": response.status,
"body": response.json(),
})
except Exception:
pass
page.on("response", on_response)
page.goto(url, wait_until="networkidle", timeout=60_000)
page.wait_for_timeout(1_000)
browser.close()
print(json.dumps(matches, indent=2))
In production, narrow the listener by URL pattern, response status, and expected keys. Once you identify a documented and permitted endpoint, replaying it directly is usually faster and less brittle than scraping rendered text. Keep a browser fallback because private endpoint shapes can change.
5. DOM extraction as a fallback
Use semantic elements such as headings, links, time, and data attributes. Normalize whitespace and locale-specific numbers, but retain the original text when a conversion is uncertain.
from bs4 import BeautifulSoup
from decimal import Decimal
import re
soup = BeautifulSoup(html, "html.parser")
name = soup.select_one("h1")
price = soup.select_one("[data-price], .price")
raw_price = price.get("data-price") if price and price.has_attr("data-price") else (price.get_text(" ", strip=True) if price else None)
def parse_decimal(value):
if not value:
return None
cleaned = re.sub(r"[^0-9,.-]", "", value)
# Apply a site-specific locale rule; this example assumes a decimal point.
try:
return str(Decimal(cleaned.replace(",", "")))
except Exception:
return None
record = {"name": name.get_text(" ", strip=True) if name else None,
"price": parse_decimal(raw_price),
"price_raw": raw_price}
Add regression fixtures for representative pages. Presentation markup changes more often than a documented API, so alert on missing selectors rather than emitting silently incomplete records.
6. Pagination, authentication, and normalization
- Pagination: follow the API’s cursor or page token until it is absent; stop on repeated cursors; deduplicate by a stable identifier.
- Authentication: use the documented scheme, keep secrets out of logs, and never copy browser cookies into a shared worker without an explicit security design.
- Types: validate required strings, numbers, booleans, arrays, and ISO dates before writing output.
- Redirects: store both the requested and final URL, and reject unexpected domains when that matters.
- Provenance: store source URL, retrieval timestamp, method, selector or request URL, status, and a raw payload hash.
from jsonschema import validate
schema = {
"type": "object",
"required": ["id", "name"],
"properties": {
"id": {"type": "string"},
"name": {"type": "string"},
"price": {"type": ["string", "null"]}
},
"additionalProperties": True
}
validate(instance=record, schema=schema)
7. Or skip the browser setup
ScreenshotNeo provides a website capture API when your workflow needs a rendered page artifact alongside extracted data. Its clean capture steps accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

See the ScreenshotNeo API documentation for all options. A one-call capture looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Use it when you need full-page or element captures, lazy images loaded, device presets, dark mode, custom viewport and retina scale, PDF output, custom CSS or JavaScript, clicks and waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture, or usage data. An MCP server also lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| JSON decode error | HTML error page, truncated body, or malformed script | Check status and content type, log a bounded body sample, and handle each block independently. |
| No JSON-LD found | Data is client-rendered or uses Microdata/RDFa | Inspect network responses, then parse semantic attributes as a fallback. |
| Only partial records | Pagination or lazy loading was missed | Follow cursors, wait for the required selector, and detect duplicate IDs. |
| Browser sees a challenge | Bot check, login wall, or consent flow | Use an official API or permitted authenticated session; do not attempt to bypass access controls. |
| Numbers are wrong | Locale separators or currency text | Parse with an explicit locale rule and preserve the raw value. |
| Selectors suddenly fail | Presentation markup changed | Prefer APIs or embedded data, add fixtures, and alert on missing required fields. |
9. Performance, reliability, and cost
- Prefer direct API calls: they avoid browser startup and transfer less data.
- Cache immutable pages and use conditional requests where supported. Hash raw payloads to skip unchanged normalization.
- Bound every timeout, retry only transient failures with exponential backoff, and cap concurrency to the site’s documented limits.
- Separate fetch, parse, validate, and persist stages so a parser bug does not require refetching source pages.
- Measure status distribution, parse failures, missing required fields, duplicate rates, pagination completion, and response sizes. No general accuracy or throughput benchmark should be assumed without a source-specific test.
- For browser jobs, reuse a browser process, create isolated contexts, block unnecessary resources when permitted, and wait for a meaningful selector rather than an arbitrary long delay.
10. Short FAQ
Is JSON-LD always the same as the data visible on the page?
No. It can be incomplete, stale, or intended for search engines. Compare it with the rendered content and validate fields your application actually needs.
Should I expand JSON-LD?
Use JSON-LD expansion or compaction when linked-data context and identifiers matter across vocabularies. For a small application schema, deliberate field mapping is often simpler.
Can I scrape a private XHR endpoint?
Only when your access and the site’s terms permit it. Prefer documented endpoints and expect private contracts to change.
How do I make extraction auditable?
Persist the source and final URLs, retrieval time, method, request or selector, status, raw payload hash, parser version, and validation errors.
When should I use a screenshot API?
Use one when you need a reliable rendered artifact, visual verification, or browser features without maintaining Playwright infrastructure. ScreenshotNeo’s clean shots and verdict headers make billing and capture outcomes explicit.


