Web Scraping Dynamic Websites: What Actually Works
Learn when to call a hidden data endpoint, when to parse embedded state, and when a headless browser is justified for JavaScript-heavy sites.

Direct answer: Do not begin by launching a browser for every JavaScript website. Fetch the page normally, inspect the HTML and embedded state, then inspect the browser’s network requests. If the data comes from a reproducible JSON or HTML request, call that request directly and parse its response. Use Playwright or another headless browser only when reproducing the request is difficult, interaction is required, or the output exists only in the rendered DOM or as a browser screenshot.
This sequence usually gives you less code, less network transfer, and a simpler failure model. Scrapy’s documentation recommends finding the data source and extracting it directly when content is loaded dynamically: Selecting dynamically-loaded content. A browser remains valuable, but it should be a deliberate fallback.
1. Decide whether you need a browser
| What you find | Best approach | Why |
|---|---|---|
| Data is in the initial HTML | HTTP request plus selectors | Fastest and easiest to operate |
| Data is in a script tag or serialized state | Fetch HTML, extract and parse the state | Avoids rendering while preserving structured data |
| Browser calls a JSON, GraphQL or HTML endpoint | Reproduce that request | Returns the payload directly instead of a rendered page |
| Request needs difficult tokens, interaction or browser-only behavior | Playwright or another browser | Recreates navigation and user actions |
| You need a screenshot or rendered DOM | Headless browser or screenshot API | The rendered output is the product you need |
Use this checklist before writing browser code:
- Can an ordinary GET response contain the fields you need?
- Do script elements contain JSON or a JavaScript object with the records?
- Which request in DevTools Network returns the records?
- What method, URL, query, body, cookies and headers does that request require?
- Does the response have a stable JSON, HTML or XML format?
- Is an interaction such as clicking “load more,” selecting a filter or scrolling required?
- Are you collecting data in a way allowed by the site’s instructions and access terms?
2. Fetch the page without rendering
Start with a normal request and inspect the response you actually received. A browser’s Elements panel shows a post-render DOM; it is not proof that the server sent those values in the original response.

import requests
from bs4 import BeautifulSoup
url = "https://example.com/catalog"
r = requests.get(
url,
headers={"User-Agent": "catalog-research/1.0"},
timeout=30,
)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
for card in soup.select("article.product"):
name = card.select_one(".name")
price = card.select_one(".price")
print({
"name": name.get_text(" ", strip=True) if name else None,
"price": price.get_text(" ", strip=True) if price else None,
})
Inspect the raw response when selectors return nothing. Save a sample, search for a known title or ID, and check the response’s Content-Type. A page can contain an empty container that JavaScript fills later; in that case, normal selectors cannot recover records that were never sent.
Embedded state in script elements
Many applications place initial data in a script tag for the front-end to hydrate. Look for JSON-LD, a script with an application-specific ID, or a JavaScript assignment. Parse strict JSON with json.loads whenever possible. Treat JavaScript object literals as a separate format: they may contain trailing commas, unquoted keys or expressions that are unsafe to evaluate.
import json
import requests
from bs4 import BeautifulSoup
html = requests.get("https://example.com", timeout=30).text
soup = BeautifulSoup(html, "html.parser")
node = soup.select_one("script#__NEXT_DATA__")
if node and node.string:
state = json.loads(node.string)
for item in state.get("props", {}).get("pageProps", {}).get("items", []):
print(item.get("id"), item.get("name"))
3. Find and reproduce the data request
Open the page in a browser, open Developer Tools, select Network, reload, and filter by fetch or xhr. Trigger the action that reveals the records, such as a search or “next page.” Select the request that contains the data and record:
- HTTP method and full URL
- Query parameters or form/JSON body
- Headers that affect the response, such as
Authorization,Refereror a CSRF token - Cookies and whether they are session-specific
- Response content type and pagination fields
Use “Copy as cURL” as a starting point, then remove headers that are not required. Keep secrets out of source control. If the endpoint returns JSON, parse JSON directly; if it returns HTML or XML, use selectors appropriate to that format.
curl 'https://example.com/api/products?page=2&limit=50' \
-H 'Accept: application/json' \
-H 'User-Agent: catalog-research/1.0'
import requests
params = {"page": 2, "limit": 50}
r = requests.get(
"https://example.com/api/products",
params=params,
headers={"Accept": "application/json"},
timeout=30,
)
r.raise_for_status()
payload = r.json()
for product in payload.get("items", []):
print(product["id"], product["name"])
const params = new URLSearchParams({ page: '2', limit: '50' });
const res = await fetch(`https://example.com/api/products?${params}`, {
headers: { Accept: 'application/json' }
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const payload = await res.json();
for (const product of payload.items ?? []) {
console.log(product.id, product.name);
}
Pagination is part of the contract. Follow the API’s cursor or next URL when supplied instead of guessing page numbers. Put a maximum page count and a record limit in your crawler so a malformed “next” link cannot create an unbounded job.
4. Render with Playwright when rendering is required
Use a browser when the request cannot be reproduced reliably, when a workflow requires clicks or scrolling, or when the desired artifact is the rendered DOM or a screenshot. Playwright exposes navigation and page-event APIs in Python and other languages; its Page API reference documents the available controls.
Install the Python package and browser once:
python -m pip install playwright
playwright install chromium
This complete example waits for a selector, extracts rendered cards, and saves the final HTML:
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page(
viewport={"width": 1440, "height": 900},
user_agent="catalog-research/1.0",
)
await page.goto("https://example.com/catalog", wait_until="domcontentloaded", timeout=60_000)
await page.locator("article.product").first.wait_for(timeout=30_000)
await page.locator("button.load-more").click()
await page.wait_for_timeout(500)
rows = await page.locator("article.product").evaluate_all("""cards => cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null
}))""")
print(rows)
await page.screenshot(path="catalog.png", full_page=True)
await browser.close()
asyncio.run(main())
Prefer explicit readiness conditions such as wait_for_selector or a known response over a large fixed sleep. Use networkidle only when the site eventually becomes idle; analytics, streams and polling can keep a page busy indefinitely. For multiple pages, reuse a browser process and create isolated contexts rather than launching a new browser for every URL.
Scrapy integration
Scrapy is useful for queues, deduplication, retries and item pipelines. The scrapy-playwright project lets selected requests use Playwright while the rest use Scrapy’s normal downloader. Account for its response behavior: rendered content is serialized into the response body, so a JSON document may appear inside a pre element. Parse what the integration returns instead of assuming the original server content type.
5. Parse the payload you actually received
- JSON: call
response.json()orjson.loads, validate required keys, and preserve the raw response for debugging. - HTML/XML: use stable attributes and semantic selectors; avoid selectors based only on generated class names.
- Embedded scripts: isolate the smallest script, parse strict JSON, and version your extraction path when the site changes.
- Images and PDFs: download the binary response and inspect its headers and magic bytes; do not decode it as text.
- Rendered DOM: wait for the specific content, then extract text or attributes from locators.
6. Robots rules, reliability and safe crawl behavior
Scrapy provides robots.txt middleware and a ROBOTSTXT_OBEY setting. Enable it when your project requires obedience, and set a deliberate user agent because robots rules can vary by agent. See the Scrapy robots middleware documentation. This is technical guidance; access terms and other obligations depend on the target and your use.
For reliable jobs:
- Set connect and read timeouts separately where your HTTP client supports them.
- Retry transient 429, 502, 503 and 504 responses with exponential backoff and a maximum attempt count.
- Honor
Retry-Afterwhen supplied. - Cache successful responses during development to avoid needless traffic.
- Record URL, status, elapsed time, response size and parser errors for every item.
- Detect empty results and schema changes as failures instead of silently exporting blanks.
- Bound concurrency per host and keep crawl rate predictable.
7. Performance and cost tradeoffs
Direct endpoint extraction normally transfers a smaller payload and avoids browser startup, JavaScript execution and image layout. It is often the right choice for high-volume jobs. Browser rendering consumes more CPU and memory and can be slower, but it handles interactions and browser-only output that an HTTP client cannot.
Keep the two stages separate in your design: retrieval obtains bytes, and parsing interprets them. This makes it possible to replay a saved response when a parser changes. For browser jobs, block unnecessary media only when it cannot affect the data, reuse contexts, and capture screenshots only for records that need visual evidence. Do not claim a speedup without measuring your own target and workload; the reviewed Scrapy and Playwright documentation does not publish a universal benchmark.
8. Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
| HTML has an empty app shell | Records arrive from a later request | Inspect Network and reproduce the JSON/HTML request |
| Selector works manually but not in code | You inspected post-render DOM | Save the initial response or use Playwright |
| 401 or 403 from an endpoint | Missing session, token or required header | Identify the minimum required authentication flow; never hard-code a personal secret |
| 429 Too Many Requests | Concurrency or rate is too high | Reduce concurrency, back off, and honor server guidance |
| JSON parsing fails after scrapy-playwright | Rendered DOM serialization wrapped the document | Inspect the response body and extract the actual JSON text |
| Browser times out | Wrong readiness condition, slow resource or perpetual polling | Wait for a specific selector/response, increase timeout narrowly, and capture diagnostics |
| Results are intermittently empty | Target overload, banning or a race with rendering | Log status and timing, retry transient failures, and wait for a data-specific condition |
| “Load more” repeats records | Cursor is not advanced or deduplication is absent | Follow the returned cursor and deduplicate by a stable ID |
9. Or skip the browser setup
If your goal is a clean screenshot or PDF rather than a structured data feed, ScreenshotNeo provides a single GET request. Its capture pipeline accepts cookie and consent banners, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each step off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed; response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for all options. A minimal call:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options cover full-page captures with lazy images, CSS-element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page ranges, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. The parameter names used by other screenshot APIs also work, which simplifies migration.
ScreenshotNeo also includes an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
10. FAQ
Do I always need Playwright for JavaScript sites?
No. A JavaScript front end may call a normal endpoint that you can reproduce directly. Use a browser when interaction, difficult browser state or rendered output is genuinely required.
How do I know which request contains the data?
Reload with DevTools Network open, filter to Fetch/XHR, trigger the action that reveals the data, and inspect response previews. Confirm the request by changing a search term or page and checking that the payload changes.
Can I parse the DOM returned by a browser integration as JSON?
Not automatically. Some integrations serialize rendered DOM into the response body, even when the original document was JSON. Inspect the returned body and parse the representation your integration provides.
Should I scrape the HTML or an API?
Prefer the structured request when it is stable and permitted. Prefer HTML when the records are server-rendered or the endpoint is unavailable. Keep a browser fallback for interactions and visual output.
How should I handle a site redesign?
Keep retrieval and parsing separate, save representative fixtures, validate required fields, and alert on empty results or schema changes. This turns a silent extraction failure into a visible maintenance task.


