How to Scrape Data from React, Vue, and Angular Websites
Learn how to find where a JavaScript site’s data comes from, choose HTTP or browser rendering, and extract and validate it with runnable examples.
If an HTTP scraper returns empty HTML from a React, Vue, or Angular site, first check whether the data is absent, embedded in the response, or fetched by a separate request. The framework name alone does not determine the method. Prefer the underlying JSON or HTML response when practical; use a headless browser such as Playwright when the required content only appears after JavaScript runs or browser state is needed. [Scrapy: dynamic content]
1. Diagnose where the data comes from
- Fetch the page without rendering. Save the response body and search it for a target value. Inspect script elements for embedded JSON or other structured data. “View source” shows the response; the live DOM in developer tools may include content inserted later.
- Inspect the browser’s Network panel. Reload the page, filter for Fetch/XHR, and inspect likely responses. Find whether the target records arrive as JSON, are embedded in the HTML, or are loaded from a JavaScript resource. Scrapy recommends identifying the source and reproducing the relevant request where possible. [Scrapy documentation]
- Try the simplest permitted source. Parse HTML or embedded data if it already contains what you need. If a stable JSON request returns the records, request and parse that response. Do not assume an endpoint is stable or authorized just because it is visible in the browser.
- Render only when needed. Use browser automation if page execution, interaction, or browser-specific state is required, or if recreating the data request is impractical.
Server-side rendering or pre-rendering can put content in the initial response. An app-shell page can instead return a shell and render its content in the browser. Google describes these as different patterns; its crawling, rendering, and indexing process applies to Google Search and is not a guarantee about every scraper. [Google Search Central: JavaScript SEO basics]
| What you observe | Good starting method |
|---|---|
| Target text is in the initial response HTML | HTTP client and HTML selectors |
| Records are embedded in a script as structured data | Extract and parse that representation |
| A request response contains the records as JSON | Reproduce that request when appropriate and parse JSON |
| Content exists only after scripts or interactions run | Headless browser and a content-based readiness condition |
| A crawl needs orchestration with occasional browser rendering | Use a crawler with a browser integration |
2. Try the HTTP and JSON path first
This runnable Python example fetches a page and searches its response for a marker. Replace the URL and marker with values from a page you are permitted to access. It is a diagnostic: a missing marker does not prove that the data is unavailable; it may come from a separate request or require rendering.
import requests
from bs4 import BeautifulSoup
url = "https://example.com/catalog"
response = requests.get(url, timeout=30)
response.raise_for_status()
html = response.text
marker = "Example item"
print("Marker in response:", marker in html)
soup = BeautifulSoup(html, "html.parser")
for script in soup.find_all("script"):
text = script.string or script.get_text()
if text and marker in text:
print("Marker found in script data")
for heading in soup.select("h1, h2"):
print(heading.get_text(" ", strip=True))
Install dependencies with python -m pip install requests beautifulsoup4. If the browser Network panel shows a suitable JSON request, request that URL and parse its response instead of parsing the rendered page:
import requests
api_url = "https://example.com/api/catalog"
response = requests.get(api_url, timeout=30)
response.raise_for_status()
data = response.json()
for record in data["items"]:
print(record.get("name"), record.get("price"))
The API URL and fields above are placeholders. Use the request and response structure you actually observe, and check whether the site permits this access.
3. Render the page with Playwright when necessary
When the needed content only appears after JavaScript runs, Playwright can launch a browser, navigate to the page, wait for a target condition, and read the DOM. The example below waits for a result element rather than treating a fixed delay as proof that the page is ready. Playwright documents locator-based waiting through its Page API. [Playwright Page API]
Install Playwright for Python and its browser binaries:
python -m pip install playwright
python -m playwright install chromium
Save this as scrape_page.py, replace the URL and CSS selector, then run python scrape_page.py:
import asyncio
from playwright.async_api import async_playwright
URL = "https://example.com/catalog"
ITEM_SELECTOR = "[data-testid='catalog-item']"
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
response = await page.goto(URL, wait_until="domcontentloaded", timeout=60000)
if response:
print("HTTP status:", response.status)
# Wait for the content this scraper actually needs.
await page.locator(ITEM_SELECTOR).first.wait_for(state="visible", timeout=20000)
items = await page.locator(ITEM_SELECTOR).evaluate_all("elements => elements.map(el => ({text: el.innerText.trim()}))")
if not items:
raise RuntimeError("The page loaded but no catalog items matched the selector")
for item in items:
print(item["text"])
await browser.close()
asyncio.run(main())
Use the page’s actual selectors. A selector copied from a temporary class name may break after a redesign; prefer stable attributes or semantic structure where available. If a page progressively adds records, waiting for the first record is not enough: wait for a known count, a next-page state, or another condition that matches the data you need.
4. Choose a readiness condition that matches the page
- Target element appears: wait for the result container or a representative record.
- List is populated: wait for a minimum count or a known end condition. Recheck counts after navigation or filter changes.
- Page performs an interaction: click the relevant control, then wait for the resulting content or state to appear.
- Network activity settles: this can help on pages with several requests, but persistent analytics or polling may prevent a network-idle condition. Prefer a specific content condition when possible.
- Fixed delay: use only as a diagnostic fallback. It can be too short on a slow run and waste time on a fast one.
Some browser-rendering APIs also expose selector-based waits; Cloudflare’s Browser Rendering reference documents a selector wait option. [Cloudflare Browser Rendering API reference]
5. Extract, validate, and handle changing pages
Getting a non-empty DOM is not the same as getting correct records. Validate the result before using or storing it.
- Check that required fields exist and have the expected types.
- Compare a small sample with what the page visibly displays.
- Detect zero records, duplicate records, error messages, and unexpected item counts.
- Account for pagination, client-side route changes, lazy-loaded content, and lists that update after filters.
- Record the URL, timestamp, response status, and a useful error reason so failures can be diagnosed.
Selectors and request assumptions can break when a site changes. Keep extraction logic small, report validation failures clearly, and revisit the observed request or DOM when the page structure changes.
6. Respect crawl boundaries
Check the site’s terms, access controls, and applicable legal requirements before collecting data. RFC 9309 defines robots.txt as a protocol for crawler access preferences; it explicitly says, “These rules are not a form of access authorization.” An allowed path in robots.txt does not grant access to protected content. [IETF RFC 9309]
Keep requests bounded and avoid collecting personal or restricted data without an appropriate basis. This guide covers diagnosing data delivery and extraction, not bypassing access controls or defeating bot checks.
7. Performance, reliability, and cost
A direct request to a structured response usually avoids the extra browser startup and page-rendering work, but the actual speed and resource use depend on the site, network, and workload. Browser rendering adds browser processes and timing coordination; use it only for pages whose required data depends on it. There is no universal performance or cost figure for these approaches.
- For reliability: use explicit timeouts, content-based waits, status checks, and field validation. Retry only transient failures with a limit and backoff; repeating a deterministic selector failure will not fix it.
- For throughput: control concurrency to avoid overloading the target or exhausting local resources. Reuse a browser process for a bounded batch where appropriate, and close pages and browsers when done.
- For operating cost: count network requests, browser runtime, compute, storage, and maintenance. Direct JSON extraction may reduce browser work, while changes to undocumented request patterns can increase maintenance.
- For long crawls: checkpoint progress, make writes idempotent, and keep enough logs to distinguish a page with no records from a failed load.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its one-call API can capture a page after rendering it in a browser; screenshots are useful for visual inspection and monitoring, but they do not replace structured JSON or DOM extraction when you need record fields.
See the ScreenshotNeo API docs. This example saves a screenshot response; set your API key and use an authorized target URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers say which page verdict occurred and whether it was billed.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
Common problems and fixes
| Symptom | Likely cause | What to check or change |
|---|---|---|
| HTTP response has no visible records | Records are fetched later or embedded in a script | Inspect scripts and Network requests; try the structured response before rendering. |
| Browser opens but selector times out | Wrong selector, navigation error, delayed content, or a different page state | Check status and URL, inspect the live DOM, verify the selector, then wait for a meaningful condition. |
| Some records are missing | Pagination, lazy loading, filters, or partial list rendering | Handle pagination or scrolling as required and validate the final count or end state. |
| Network-idle wait never finishes | Persistent polling, streaming, or analytics requests | Wait for the target content instead of all network activity to stop. |
| JSON parsing fails | Response is HTML, an error object, or a changed schema | Check status, content type, and a sample body before parsing; handle error responses separately. |
| Scrape works locally but fails in a batch | Concurrency, resource pressure, transient network failures, or an unhandled timeout | Reduce concurrency, set bounded timeouts, log failures, and retry transient errors with limits. |
| Records suddenly disappear after a site update | Selector or request assumptions changed | Reinspect the page and network flow, update extraction logic, and retain validation alerts. |
Frequently asked questions
Do React, Vue, and Angular need separate scraping tools?
No. Choose based on where the target data is delivered and whether browser execution or interaction is required.
Does an empty initial response mean the site has no data?
No. The data may arrive in a later request, be embedded in a script, or appear after JavaScript renders the page.
Should I use Playwright or call the API?
Use the structured request when it is practical and permitted. Use Playwright when the required result depends on rendered DOM, interaction, or browser state.
Can Google’s rendering behavior predict what my scraper sees?
No. Google documents its own Search crawling and rendering process. Other clients can behave differently.
Further reading
For broader background, Web Scraping with Python, 3rd Edition by Ryan Mitchell is an optional general reference covering JavaScript scraping and crawling through APIs. O’Reilly lists the book as published in February 2024, with 352 pages, for intermediate to advanced readers.


