ScreenshotNeo

BlogHow-to

How to Scrape Dynamic Web Pages

Learn how to find where dynamic page data comes from, reproduce its request, parse the response, and use browser automation when the task needs it.

By the ScreenshotNeo team4 October 202611 min read

A basic HTTP scraper may receive HTML that does not contain the information you see in a browser. That does not automatically mean you need a browser to scrape the page. First inspect the initial response and the requests the page makes. If the data comes from a reproducible request, fetch and parse that response directly; use browser automation when the task depends on browser interaction or rendered output, or when reproducing the request is impractical. Scrapy’s guidance follows this same workflow: locate the data source, reproduce its request where feasible, and use a browser when the task calls for it. Scrapy: Dynamic content

1. Check what the initial response contains

Start by separating what the server returned from what the browser eventually displayed. A browser can make additional requests and run scripts after the first HTML response arrives. Your scraper may be looking only at that first response.

  1. Request the page with your ordinary HTTP client.
  2. Record the status code and inspect the response body, not just the browser’s Elements panel.
  3. Search the body for a distinctive phrase or value you expected to scrape.
  4. If it is missing, inspect script elements for embedded state and then inspect the browser’s network activity.

For example, a site may include the data in a JSON script element, or fetch it from a separate JSON endpoint. In either case, the browser DOM is not necessarily the best source to scrape.

curl -i -L 'https://example.com/products' -o response.html

# Search the saved response for a phrase that should appear on the page.
rg -i 'expected product name' response.html

Replace the example URL and phrase with values from the site you are examining. The command saves the body to a file; use the response headers and status line shown by -i to help diagnose what the server returned.

If a browser receives different content from your HTTP client, compare the requests before concluding that JavaScript rendering is required. Scrapy recommends checking the response with an HTTP client and comparing request details such as the user agent when behavior differs. A difference can come from request construction or server behavior. Scrapy’s dynamic content guide

2. Find where the data comes from

When a value is absent from the initial HTML, check two places: embedded JavaScript and the browser’s network activity.

Inspect embedded state

Search the response body for the field names or values you expect. Some pages serialize initial application state into a script element. If the data is there, parse that script payload rather than rendering the page. The exact script format varies by site, so inspect the actual response before writing a parser.

Inspect network requests

Open the page in browser developer tools and use the Network panel. Reload the page, then look at requests that appear as the content loads. Find a response containing the target data. Record the request method, URL, query parameters, body, and headers; also note whether cookies or a prior request appear necessary. Playwright’s documentation describes observing and handling browser network traffic in automation. Playwright: Network

For a request you want to understand, check:

  • Method and URL: Is it a GET, POST, or another method? Are filters encoded in the query string?
  • Request body: Does it send JSON or form fields?
  • Headers: Does it require a content type, authorization, or other request headers?
  • Response format: Is the result HTML, JSON, XML, or another format?
  • Sequence: Does the page first obtain a token or cookie that the data request uses?

Inspecting a request helps you understand how the page obtains its data; it does not establish permission to collect or reuse that data. Follow the site’s applicable terms and rules.

3. Reproduce the data request

If the request is practical to reproduce, send it directly and parse its response in its native format. The method and URL may be enough. If not, include the body, headers, and parameters you observed. Scrapy documents reproducing the request that supplies dynamic data and parsing responses according to their format. Scrapy: Dynamic content

cURL: inspect a JSON response

curl -sS -X GET 'https://example.com/api/products?category=books' \
  -H 'Accept: application/json' \
  -o products.json

python -m json.tool products.json

Use the URL, method, and headers observed for the actual site. If the browser sent a request body, reproduce that too. For example, a JSON POST body can be sent with -H 'Content-Type: application/json' and --data '{"category":"books"}'. Do not copy secrets, session tokens, or personal data into source code or logs.

Python: request JSON and extract fields

import requests

url = "https://example.com/api/products"
params = {"category": "books"}
headers = {"Accept": "application/json"}

response = requests.get(url, params=params, headers=headers, timeout=30)
response.raise_for_status()
data = response.json()

for product in data["products"]:
    print(product.get("name"), product.get("price"))

This is runnable after replacing the example endpoint and adapting the JSON field names to the response you observed. The timeout bounds how long the client waits for a response; raise_for_status() makes HTTP error responses visible instead of treating their bodies as successful data.

Node.js: request JSON and extract fields

const url = new URL('https://example.com/api/products');
url.searchParams.set('category', 'books');

const response = await fetch(url, {
  headers: { Accept: 'application/json' },
  signal: AbortSignal.timeout(30_000),
});

if (!response.ok) {
  throw new Error(`Request failed: ${response.status} ${response.statusText}`);
}

const data = await response.json();
for (const product of data.products) {
  console.log(product.name, product.price);
}

This uses the built-in fetch available in current Node.js releases. Change the endpoint, parameters, and response field names to match the site. For a POST request, set the method, content type, and body to match the observed request.

Parse according to the response

  • JSON: Use a JSON parser, then validate the expected fields and types.
  • HTML or XML: Use selectors suited to the document, and account for missing or repeated elements.
  • Embedded script data: Identify the serialization format and parse it deliberately; do not assume every script element contains valid JSON.
  • Image-based content: Text may not be present in document markup. Choose an extraction method appropriate to the actual image or document format.

Scrapy’s documentation covers selectors and JSON as well as JavaScript and image-based extraction approaches. Scrapy: Dynamic content

4. Use browser automation when the task needs a browser

Use a browser when the result depends on interaction, when the browser must construct the output you need, or when reproducing the underlying request is unusually difficult. Examples include clicking through a multi-step interface, waiting for a browser-rendered view, or capturing a screenshot. A browser is also useful during diagnosis because it lets you observe page behavior and network requests. It adds a browser runtime and page-loading work, so use it where those capabilities matter.

Python example with Playwright

Install Playwright and its Chromium browser using its official Python setup instructions, then save this as scrape_page.py. The example waits for a product selector, reads its text, and prints it. Replace the URL and selector with ones that match the page.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/products", wait_until="domcontentloaded")
    page.locator(".product-name").first.wait_for(timeout=15_000)
    names = page.locator(".product-name").all_text_contents()
    for name in names:
        print(name.strip())
    browser.close()

Playwright documents page navigation, locators, and waiting for page conditions. Playwright for Python: Getting started · Playwright Python: Page API

domcontentloaded waits for the document to be parsed, not for every later request or visual update. Waiting for the specific target locator makes the example more tied to the content it needs. Choose a wait condition that matches the page; a fixed delay may work for a known short delay, but it can waste time or still be too short when load times vary.

5. Choose the lightest method that returns the needed data

Method Use it when Trade-off
Initial HTML response The response already contains the content. Fast to inspect and parse; does not include data fetched later.
Embedded page state The initial response includes serialized data in a script. Avoids rendering, but the embedded format can be site-specific.
Reproduced data request A browser request returns the desired structured data and can be recreated. Often gives structured data with less parsing and transfer work; requires matching the request.
Browser automation The task requires browser interaction, browser-produced output, or request reproduction is impractical. Runs a browser and must handle page waits and browser behavior.

Scrapy describes reproducing data requests as a way to obtain structured data with minimum parsing time and network transfer, while browser rendering helps when the request is difficult to reproduce or browser output is required. These are method trade-offs, not a guaranteed speed or cost result for every site. Scrapy: Dynamic content

6. Handle edge cases and inconsistent responses

  • Pagination: Determine whether the next page is represented by a URL, request parameter, or cursor in the response. Follow the observed pattern and stop when the response indicates there are no more results.
  • Filters and sorting: Reproduce the parameters the browser sends; do not assume the visible control changes only the page locally.
  • Authentication or session state: A request may depend on a legitimate authenticated session. Keep credentials secure and follow the site’s access rules.
  • Changing response shape: Validate required fields and handle absent values. A parser that assumes every response has the same structure can fail when the site changes.
  • Different results across clients: Compare status, response body, and request details such as user agent. Do not assume that a browser is required until you know what differs.
  • Intermittent results: Record the request and response for each attempt. Scrapy notes that inconsistent expected responses can reflect a target server that is buggy, overloaded, or banning requests; investigate the evidence rather than assuming a crawler defect or a particular cause. Scrapy: Dynamic content

7. Troubleshooting

Symptom Likely cause to investigate What to do
Expected text is missing from the response The data is embedded in a script or loaded by a later request. Search the initial HTML, inspect script data, and inspect browser network activity.
Browser shows data but HTTP client does not The requests differ, or the page fetches data after its initial response. Compare status, body, method, URL, headers, and follow-up requests.
JSON parsing fails The response is not JSON, is an error page, or has a different content type or shape. Check the status and body before parsing; verify the endpoint and expected schema.
Request returns an HTTP error The URL, method, parameters, body, headers, or access state may not match the observed request. Compare the request details with the browser request and inspect the error response.
Browser script times out waiting for content The selector may be wrong, the content may not load, or the chosen wait condition may not fit the page. Confirm the selector in the rendered page, inspect network activity, and wait for a meaningful page condition.
Results are inconsistent The target may respond inconsistently or request construction may vary. Log request and response details across attempts; distinguish observed facts from possible causes.

8. Respect crawler guidance and access boundaries

Check the target site’s crawler guidance and applicable terms before collecting data. RFC 9309 defines the Robots Exclusion Protocol and explains that robots.txt rules are crawler guidance, not access authorization: “These rules are not a form of access authorization.” That means a robots.txt rule neither grants permission nor settles legal questions. Whether collection or reuse is allowed depends on the specific site, purpose, data, terms, and applicable rules. IETF RFC 9309: Robots Exclusion Protocol

9. Performance, reliability, and cost

For structured data, a direct request is usually the simpler path when the data endpoint is reproducible: it avoids launching a browser and parsing a rendered page. Browser automation has additional runtime and page-loading work, but it is the appropriate choice when interaction or browser output is part of the requirement. Actual time, reliability, and resource use depend on the site and workload; do not infer them from the method alone.

  • Request only the data needed and parse the response format directly.
  • Set timeouts and surface HTTP errors so failures are visible.
  • Validate response fields before using them, and preserve enough request context to diagnose failures.
  • For browser work, wait for the relevant element or state instead of relying on an arbitrary delay.
  • Respect the site’s crawler guidance and applicable access rules; do not treat a robots.txt entry as authorization.

Or skip the browser setup

If your task is to capture what a page looks like rather than extract its underlying structured data, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. Responses report the page verdict and billing status in headers.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 shots, and every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Does every JavaScript website need a headless browser?

No. First look for data in the initial HTML, embedded scripts, or a reproducible network request. Use a browser when the task depends on interaction or rendered output, or the request is difficult to reproduce.

Can I scrape a page by copying its API request?

You can reproduce a request when it is appropriate to do so and you can match the necessary method, URL, body, and headers. A visible request does not by itself establish permission to collect or reuse its response.

Does robots.txt give permission to scrape?

No. RFC 9309 says robots.txt rules are not a form of access authorization. Check the site’s terms and the rules that apply to your situation.

When is a screenshot useful?

Use a screenshot when the output you need is the page’s visual appearance. If you need structured fields, find and parse the response that contains those fields where practical.