How to Extract React Props When Scraping a Website with Python
Learn how to inspect server-rendered HTML, find serialized React state, validate it safely, and handle pages whose data loads only in JavaScript.
Short answer: React does not expose one universal public “props” object for scrapers. Request the page, inspect the raw HTML, find script or data elements that contain serialized state, parse valid JSON, and validate the structure before using it. If the required data appears only after client-side JavaScript runs, use an authorized data endpoint or a browser-capable workflow.
Server-rendered React can place initial data in the response used for hydration, but that HTML is not automatically the application’s complete runtime state. Frameworks, routes, rendering modes, sessions, and versions all affect where data appears.
1. Understand what you are extracting
“React props” is often shorthand for framework-serialized page data. React’s server APIs render components into HTML and hydration later makes the result interactive; React does not promise a scraper-facing props object. See the React server API documentation.
A response can contain several different kinds of information:
- Rendered markup: HTML already produced on the server.
- Serialized route data: JSON or another payload embedded in a script element for hydration.
- Client-only data: values fetched after JavaScript executes.
- Session-specific data: content that depends on cookies, authentication, locale, or other request context.
Extract only data you are authorized to access, and follow the target site’s terms and access rules.
2. Inspect the raw response before parsing
Always retain the status code, final URL, headers, and response body. A scraper may receive a bot challenge, login page, redirect, or error document instead of the page you expected.
import requests
url = "https://example.com/page"
response = requests.get(
url,
headers={"User-Agent": "Mozilla/5.0 (compatible; research client)"},
timeout=20,
allow_redirects=True,
)
print("status:", response.status_code)
print("final URL:", response.url)
print("content type:", response.headers.get("content-type"))
print(response.text[:500])
response.raise_for_status()
content_type = response.headers.get("content-type", "").lower()
if "html" not in content_type:
raise ValueError(f"Expected HTML, received {content_type!r}")
Check for signs that the response is not the intended page:
- A status such as 401, 403, 429, or 5xx.
- A title or body saying that you must sign in, enable JavaScript, or complete a challenge.
- A redirect to a different hostname or route.
- An unusually small document or a generic error template.
3. Parse script elements as elements
Use an HTML parser instead of regular expressions over the entire document. Beautiful Soup supports element lookup; its get_text() helper is intended for human-readable text and generally does not include script contents. Read a script element’s contents directly.
import json
import requests
from bs4 import BeautifulSoup
url = "https://example.com/page"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for script in soup.find_all("script"):
script_id = script.get("id")
script_type = script.get("type")
raw = script.string or script.decode_contents()
if raw and (script_id or script_type):
sample = raw.strip().replace("\\n", " ")[:160]
print({"id": script_id, "type": script_type, "sample": sample})
Do not assume an identifier such as __NEXT_DATA__ exists on every React site. Inspect the actual response first. A framework can change its payload format between versions or rendering modes.
4. Parse and validate a candidate payload
Only pass a candidate to json.loads when it is valid JSON. Some scripts contain JavaScript assignments, wrappers, escaped strings, or a non-JSON serialization. Never execute scraped script content.
import json
import requests
from bs4 import BeautifulSoup
def parse_json_script(soup, script_id=None, script_type=None):
"""Return the first valid JSON object matching the supplied hints."""
for script in soup.find_all("script"):
if script_id is not None and script.get("id") != script_id:
continue
if script_type is not None and script.get("type") != script_type:
continue
raw = script.string or script.decode_contents()
if not raw or not raw.strip():
continue
try:
value = json.loads(raw)
except json.JSONDecodeError:
continue
if isinstance(value, (dict, list)):
return value
return None
url = "https://example.com/page"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
# Replace this with an id or type observed in the target response.
state = parse_json_script(soup, script_id="REPLACE_WITH_OBSERVED_ID")
if state is None:
raise ValueError("No valid JSON state script was found")
if not isinstance(state, dict):
raise ValueError(f"Unexpected state type: {type(state).__name__}")
# Validate the fields your application actually needs.
props = state.get("props")
if props is not None and not isinstance(props, dict):
raise ValueError("The props field has an unexpected type")
print(state.keys())
The placeholder identifier is deliberate: the correct selector must be confirmed against the target response. Some parsed elements expose their content through a child string, which is why the example checks both script.string and decode_contents().
5. Find framework-serialized data safely
Next.js Pages Router
For a Next.js Pages Router page, inspect the returned document for the framework’s data payload and verify its shape for the specific route and version. Next.js documents getServerSideProps as part of its server-side rendering workflow, but that documentation does not establish one universal payload identifier or structure for every Next.js generation.
import json
import requests
from bs4 import BeautifulSoup
url = "https://example.com/page"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
candidate = soup.find("script", id="__NEXT_DATA__")
if candidate is None:
raise LookupError("The expected Next.js script was not present in this response")
raw = candidate.string or candidate.decode_contents()
try:
payload = json.loads(raw)
except json.JSONDecodeError as exc:
raise ValueError("The observed script is not plain JSON") from exc
if not isinstance(payload, dict):
raise ValueError("Unexpected Next.js payload type")
# Inspect first, then depend only on fields confirmed for this route.
print(payload.keys())
page_props = payload.get("props", {}).get("pageProps")
if page_props is not None and not isinstance(page_props, dict):
raise ValueError("Unexpected pageProps shape")
Other React frameworks
React itself does not dictate where a framework stores loader data. Search the actual response for JSON script tags, framework-specific identifiers, or data attributes, then document the observed format in your scraper. Treat every internal format as changeable.
6. When the data is not in the initial HTML
If the browser shows information that the response does not contain, compare the raw response with the browser’s rendered result. The page may fetch data after hydration, require an interaction, or be using streaming and Suspense.
React documents that renderToString does not wait for suspended content; a suspended component can produce its nearest fallback instead. Streaming server rendering is a separate approach. See the renderToString documentation.
Use this decision sequence:
- Look for an authorized, documented endpoint that returns the needed data.
- Reproduce the required request context, such as a permitted cookie, locale, or authorization header.
- If JavaScript execution is essential, use a suitable browser automation stack and wait for a specific selector or network condition.
- Capture the final DOM or call the application’s documented data layer, rather than treating internal React objects as a stable API.
7. Embedded state is untrusted input
Do not evaluate a script with exec, a JavaScript runtime, or an equivalent mechanism merely to extract data. Parse data formats explicitly and validate types, sizes, and expected keys.
Serialization can also create security problems for the site producing the page. TanStack Query’s SSR guide warns that plain JSON.stringify does not escape script-sensitive content by default in custom SSR. As a scraper, assume embedded values may contain unexpected strings and handle them as data.
8. A reusable extraction function
from __future__ import annotations
import json
from typing import Any
import requests
from bs4 import BeautifulSoup
def fetch_html(url: str) -> tuple[requests.Response, BeautifulSoup]:
response = requests.get(url, timeout=20, allow_redirects=True)
response.raise_for_status()
content_type = response.headers.get("content-type", "").lower()
if "html" not in content_type:
raise ValueError(f"Expected HTML, received {content_type!r}")
return response, BeautifulSoup(response.text, "html.parser")
def json_script_candidates(soup: BeautifulSoup) -> list[dict[str, Any]]:
candidates: list[dict[str, Any]] = []
for script in soup.find_all("script"):
raw = script.string or script.decode_contents()
if not raw or not raw.strip():
continue
try:
value = json.loads(raw)
except json.JSONDecodeError:
continue
if isinstance(value, (dict, list)):
candidates.append({
"id": script.get("id"),
"type": script.get("type"),
"value": value,
})
return candidates
response, soup = fetch_html("https://example.com/page")
print("status:", response.status_code, "final URL:", response.url)
candidates = json_script_candidates(soup)
for item in candidates:
value = item["value"]
print("id:", item["id"], "type:", item["type"], "top-level:", type(value).__name__)
if not candidates:
raise LookupError("No valid JSON script candidates found; inspect for client-side loading")
# Select a candidate only after confirming its identity and shape.
state = candidates[0]["value"]
if isinstance(state, dict):
print("keys:", list(state)[:20])
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Response is 403, 429, or a challenge page | Access controls, rate limits, or bot protection | Respect the site’s rules, slow requests, use an authorized endpoint, or obtain permission. Do not attempt to bypass a challenge. |
| No matching script element | The identifier is different, the route is client-rendered, or the payload is not embedded | List script ids and types, inspect samples, then compare with the browser response. |
json.loads raises an error |
The script contains an assignment, wrapper, escaped value, or another format | Confirm the format before writing a parser. Do not execute the script. |
| JSON exists but expected keys are absent | Wrong candidate, route variant, session state, or a changed framework schema | Log the candidate’s top-level keys and validate the route-specific shape. |
| HTML has a loading shell or Suspense fallback | Content is suspended or fetched after hydration | Use a documented endpoint or a browser workflow that waits for the required content. |
| Browser and requests output differ | Cookies, headers, viewport, locale, authentication, or JavaScript affect the response | Compare request context and determine which differences you are authorized to reproduce. |
| Values change between requests | Personalization, experiments, time-sensitive data, or cache variation | Record headers and session context; validate fields instead of assuming a permanent schema. |
10. Performance, reliability, and cost
- Prefer one response per page: inspect and parse locally before adding follow-up requests.
- Set explicit timeouts: a hung origin should not block the whole crawl.
- Bound document and payload sizes: reject unexpectedly large responses before parsing them.
- Cache carefully: cache only when the site’s rules and the data’s freshness requirements allow it.
- Retry selectively: transient network failures may be retried with backoff; repeating authorization or challenge failures usually will not help.
- Measure the right stage: distinguish request time, HTML parsing time, browser execution time, and downstream validation.
- Browser execution costs more operationally: it requires a browser process, waits, concurrency limits, and cleanup. Use it only when the initial response or an authorized endpoint cannot provide the data.
11. Or skip the browser setup
If your workflow needs a rendered page image while you investigate what a React route displays, ScreenshotNeo provides a website screenshot API. It does not return React props; it handles the browser capture step for visual output.
One GET request returns an image or PDF. The API documentation is at screenshotneo.com/docs/.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/page"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
import { writeFile } from "node:fs/promises";
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/page' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile("shot.webp", Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
12. Checklist
- Record status, final URL, headers, and raw HTML.
- Confirm the response is the intended page, not a challenge or login document.
- Inspect script elements and read their contents directly.
- Parse only confirmed JSON formats.
- Validate keys, types, and required fields.
- Never execute scraped script content.
- Compare the response with the browser when data is missing.
- Prefer documented and authorized endpoints.
- Use browser automation only when client-side execution is required.
- Log schema changes and handle missing fields.
13. FAQ
Can I read React component props directly from HTML?
Usually no. HTML may contain rendered values or serialized framework data, but React component instances and their live props are not a stable public scraping interface.
Does every Next.js page expose the same JSON object?
No. Route type, Next.js generation, rendering mode, and application code affect the response. Inspect and validate the actual document.
Why does Beautiful Soup return no text for a script?
Script contents are not ordinary human-visible text. Read the script element’s string or decoded contents directly.
Should I use a headless browser for every React site?
No. Start with the initial response and an authorized endpoint. Add browser execution only when the required data appears after JavaScript, interaction, or a wait condition.
Is embedded JSON safe to trust?
Treat it as untrusted input. Validate its structure and never evaluate it as code.


