ScreenshotNeo

BlogGuides

Scraping Network Requests: A Guide to Efficient Data Extraction

Find the Fetch/XHR calls behind dynamic pages, reproduce authorized requests, handle pagination and rate limits, and know when browser automation is necessary.

By the ScreenshotNeo team1 October 202612 min read

To find the data behind a JavaScript-rendered page, open Chrome DevTools, select the Network panel, enable the Fetch/XHR filter, and reload the page or repeat the interaction that displays the data. Inspect the matching request and response. If you are authorized to use the endpoint, reproduce its method, URL, parameters, body, required headers or session state, and pagination fields with an HTTP client. Use browser automation only when the data depends on rendered output or interaction that you cannot reproduce reliably with a direct request.

This workflow is for public or otherwise authorized data access. A request visible in your browser is not automatically permitted for automated reuse. Check the site’s terms and applicable rules, avoid accessing protected data, and pace requests responsibly.

1. Find the request that supplies the data

  1. Open the page in Chrome and open DevTools with F12 or Ctrl+Shift+I (Cmd+Option+I on macOS).
  2. Select Network. Keep DevTools open while reloading the page; the panel records network activity while it is open.
  3. Select Fetch/XHR to narrow the list to common JavaScript data requests. If nothing relevant appears, clear the filter and check other request types or the page’s document response.
  4. Repeat the action that reveals the data: submit a search, change a filter, click “load more,” or scroll if the page loads more results on scroll.
  5. Choose a likely request and inspect Headers, Payload, Preview, Response, Initiator, Timing, and Cookies. Search responses for a distinctive value visible on the page.
  6. Record the request method, full URL, query parameters, request body, relevant headers and cookies, response format, and any page number, cursor, offset, or continuation token.

The Network panel can record activity, inspect and filter requests, search headers and responses, change loading behavior, block requests, and export request data. See Chrome DevTools Network documentation.

If the request list is noisy, clear it immediately before repeating the interaction. Use the domain and status filters, and inspect Initiator to see what triggered a request. The timing view can help identify whether a request is waiting on connection setup, the server, or response transfer; it does not by itself prove why a remote service is slow.

2. Decide whether direct requests are suitable

A JSON endpoint can be simpler to parse than rendered HTML, but only if it is stable enough for your use and your use is permitted. Compare these approaches before building a scraper:

Approach Use it when Tradeoffs
Direct HTTP request to a structured endpoint The response contains the needed records and the request can be reproduced with authorized credentials and parameters. Usually straightforward to parse and save. Internal endpoints may change without notice and can depend on session state.
HTML request and parser The needed content is present in the server-rendered HTML. Avoids browser rendering; markup changes can break selectors. It will not reveal data loaded only after JavaScript runs.
Browser automation The page requires rendering, interaction, or browser state that you cannot reproduce with a direct request. More setup and resource use; browser, timing, and page changes can affect reliability.

Do not treat an endpoint as a supported public API merely because DevTools exposes it. Prefer documented APIs when available. Do not copy credentials into source control, share HAR files containing session cookies, or attempt to bypass access controls.

3. Reproduce an authorized request in Python

Here is a runnable template for an endpoint that accepts a GET request and returns JSON. Replace the example URL and parameters with the values you observed and are permitted to use. Install the dependency with python -m pip install requests.

import json
import os
from pathlib import Path

import requests

# Set DATA_URL to the endpoint you are authorized to call.
url = os.environ["DATA_URL"]
params = {
    "q": "example",
    "page": 1,
}
headers = {
    "Accept": "application/json",
    # Add only headers the endpoint actually requires.
    # "Authorization": f"Bearer {os.environ['API_TOKEN']}",
}

response = requests.get(
    url,
    params=params,
    headers=headers,
    timeout=(5, 30),  # connection timeout, read timeout in seconds
)
response.raise_for_status()

# Preserve a fixture so parser changes can be checked without another request.
Path("response.json").write_text(response.text, encoding="utf-8")
data = response.json()
print(json.dumps(data, indent=2, ensure_ascii=False))

For a POST request, use the observed body format. For JSON, pass json=payload; for form data, pass data=payload. Do not change the method or body encoding just because another pattern seems more convenient.

payload = {"query": "example", "page": 1}
response = requests.post(
    url,
    json=payload,
    headers=headers,
    timeout=(5, 30),
)
response.raise_for_status()
data = response.json()

Only add cookies or authentication when your authorized workflow requires them. Prefer environment variables or a secret manager; do not paste live session cookies into scripts that may be logged or committed.

4. Reproduce the request with cURL

For a GET request, encode query parameters rather than assembling an unescaped URL by hand:

curl --fail --show-error --silent \
  --get "$DATA_URL" \
  --data-urlencode 'q=example' \
  --data-urlencode 'page=1' \
  --header 'Accept: application/json' \
  --output response.json

For a JSON POST request:

curl --fail --show-error --silent \
  "$DATA_URL" \
  --request POST \
  --header 'Accept: application/json' \
  --header 'Content-Type: application/json' \
  --data '{"query":"example","page":1}' \
  --output response.json

Chrome DevTools can copy a request as cURL from its context menu. Treat copied commands as sensitive: they may include cookies, authorization headers, or other session data. Remove anything unnecessary and keep secrets out of shell history and shared logs.

5. Reproduce the request with Node.js

Recent Node.js versions provide fetch globally. Set DATA_URL in the environment, then run this as an ES module or in a Node.js environment that supports top-level await:

const endpoint = new URL(process.env.DATA_URL);
endpoint.searchParams.set('q', 'example');
endpoint.searchParams.set('page', '1');

const response = await fetch(endpoint, {
  headers: { Accept: 'application/json' },
  signal: AbortSignal.timeout(30_000),
});

if (!response.ok) {
  throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}

const data = await response.json();
console.log(JSON.stringify(data, null, 2));

For an observed POST request, preserve its content type and payload format:

const response = await fetch(process.env.DATA_URL, {
  method: 'POST',
  headers: {
    Accept: 'application/json',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ query: 'example', page: 1 }),
  signal: AbortSignal.timeout(30_000),
});

if (!response.ok) {
  throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
const data = await response.json();

6. Handle pagination without losing or duplicating records

Pagination is defined by the response and request you observed. Common patterns include incrementing a page number, requesting an offset and limit, following a continuation URL, or sending a cursor returned by the previous response. Do not guess that the last page is an empty array if the service provides an explicit completion field.

  1. Identify the request field that advances the result set and the response field that supplies its next value.
  2. Stop when the documented or observed completion signal says there are no more results. Add a maximum page or record limit as a safety stop.
  3. Deduplicate using a stable record identifier. If no stable key exists, determine a suitable composite key from the data.
  4. Persist the next page or cursor and completed record keys so a long run can resume after interruption.
  5. Log the page or cursor, status, record count, and elapsed time. Avoid logging credentials or full sensitive payloads.

Example for a page-number response with a items array and has_more flag:

page = 1
seen_ids = set()
records = []
max_pages = 100  # safety bound; choose an appropriate value

for _ in range(max_pages):
    response = requests.get(
        url,
        params={"q": "example", "page": page},
        headers=headers,
        timeout=(5, 30),
    )
    response.raise_for_status()
    result = response.json()

    for item in result["items"]:
        record_id = item["id"]
        if record_id not in seen_ids:
            seen_ids.add(record_id)
            records.append(item)

    if not result.get("has_more", False):
        break
    page += 1
else:
    raise RuntimeError("Stopped at max_pages; verify the pagination stop condition")

Adapt the fields to the actual response. Cursor pagination usually requires sending the returned cursor unchanged; offset pagination may shift if records are added or removed during a crawl, so use a stable sort or snapshot mechanism if the service provides one.

7. Pace requests and handle retries

Start with low, bounded concurrency. Cache responses where allowed, avoid repeating identical requests, and use a clear stop condition. If a response includes Retry-After, honor it before retrying: RFC 9110 permits either an HTTP date or a delay in seconds and describes the field as the server’s requested wait before a follow-up request. See RFC 9110, Retry-After.

A simple Python helper for parsing either form:

from datetime import datetime, timezone
from email.utils import parsedate_to_datetime


def retry_after_seconds(value, now=None):
    """Return a non-negative delay, or None when the value is invalid."""
    if value is None:
        return None
    value = value.strip()
    try:
        return max(0.0, float(value))
    except ValueError:
        pass

    try:
        retry_at = parsedate_to_datetime(value)
        if retry_at.tzinfo is None:
            retry_at = retry_at.replace(tzinfo=timezone.utc)
        current = now or datetime.now(timezone.utc)
        return max(0.0, (retry_at - current).total_seconds())
    except (TypeError, ValueError, OverflowError):
        return None

For transient network errors and selected server errors, use bounded exponential backoff with jitter. Do not blindly retry every status: a 400-series response often needs a corrected request, and repeated authentication failures will not improve by retrying. A 429 or 503 may carry Retry-After; when present, treat that value as the pacing instruction. Set a retry limit and stop when the service continues failing.

Concurrency multiplies request load. Increase it only when you have a clear reason and the service permits it. Avoid parallel pagination when each page depends on the preceding cursor.

8. Check robots.txt and access constraints

Before crawling a host, retrieve its /robots.txt, identify the user-agent group that applies to your crawler, and follow the most-specific matching Allow or Disallow rule. RFC 9309 describes the Robots Exclusion Protocol and says that robots rules are not access authorization. Read RFC 9309.

A robots.txt rule does not grant permission to access a protected endpoint, override a site’s terms, or authorize collection of personal or otherwise restricted data. Likewise, a browser session that can view a page does not automatically authorize unattended collection. When in doubt, use a documented API or obtain permission from the site operator.

9. Diagnose dependencies with DevTools

Use throttling and request blocking locally to learn whether a page depends on a resource and how it behaves when that resource is slow or unavailable. For example, block an image or analytics domain to see whether the data request still completes. These are diagnostic controls; they do not grant permission to increase traffic to a remote service. Chrome documents request blocking and loading behavior in its Network panel guide.

10. Troubleshooting common failures

Symptom Likely cause What to check or change
No matching request appears DevTools opened after the request, wrong filter, data was already cached, or the page uses a different request type. Open Network first, enable Preserve log if navigation occurs, clear the log, reload, repeat the interaction, and inspect other request types.
Request returns 401 or 403 Missing or expired authorized credentials, wrong session context, or access is not permitted. Inspect the original request’s auth and cookie requirements. Use only credentials you are authorized to use; do not try to bypass access controls.
Request returns 400 or 422 A required parameter, body field, content type, or encoding differs from the browser request. Compare method, query, payload, and content type. Check URL encoding and required fields.
Request returns 404 Endpoint path or API version is wrong, or an internal endpoint changed. Copy the complete URL from a fresh authorized browser session and verify the method and host.
Request returns 429 Requests are arriving too quickly or a quota has been reached. Stop or reduce traffic, honor Retry-After if present, and review the service’s published limits.
Request returns 5xx or times out Remote service or network failure, overloaded endpoint, or a timeout shorter than the response time. Use bounded retries for transient failures, sensible connect/read timeouts, and a stop limit. Do not retry aggressively.
Response is HTML instead of JSON Redirect, login page, bot check, error page, or incorrect endpoint. Check final URL, status, content type, and response body before parsing. Do not assume every successful HTTP status contains the expected JSON.
Only the first batch is collected Pagination parameter or cursor is missing, or the stop condition is wrong. Compare consecutive browser requests and inspect the response for next-page fields. Persist and send the returned cursor correctly.
Duplicate or missing records Unstable ordering, shifting offsets, overlapping pages, or a faulty deduplication key. Use a stable sort or snapshot if available, deduplicate by record ID, and compare page boundaries.
Works in browser, fails from script Browser session state, required headers, cookies, or client-side challenge differs. Inspect the request details and reproduce only necessary authorized state. If the flow genuinely requires browser interaction, use browser automation where permitted.

11. Performance, reliability, and cost

  • Reduce work first: request only needed fields or pages if the endpoint supports it, avoid duplicate calls, and cache results when allowed.
  • Keep resource use bounded: use timeouts, bounded concurrency, a maximum page count, and retry limits. Save progress so failures do not require starting over.
  • Make runs reproducible: retain sanitized response fixtures and parser tests or checks. A fixture helps catch parser changes without repeatedly contacting the site.
  • Expect endpoint changes: an undocumented endpoint may change its schema, authentication, or pagination behavior. Validate required fields and fail visibly instead of silently saving malformed output.
  • Measure your own workload: log request counts and elapsed time for the permitted task. No universal speed or cost advantage can be inferred without measuring the target and workload.
  • Account for operational cost: direct HTTP requests avoid maintaining a browser when structured data is sufficient. Browser automation can require more compute and maintenance, but may be necessary when rendering or interactions are essential.

12. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. If the task is to capture a page rather than extract structured records, a single request returns an image or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Responses include page-verdict and billing headers.
  • An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
  • The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

Frequently asked questions

Can I copy a Fetch/XHR request into Python?

Yes, when you are authorized to call it. Copy the method, URL, parameters, body, and required headers or session state, then translate them into a Python HTTP request. Remove unnecessary secrets and do not assume an internal endpoint is a supported API.

Is robots.txt enough permission to scrape a site?

No. It communicates crawler preferences; RFC 9309 explicitly says the rules are not access authorization. Check the site’s terms and applicable permissions separately.

Should I use an API request or browser automation?

Use an authorized structured endpoint when it reliably contains the data. Use HTML parsing for server-rendered content, and browser automation when rendering or interaction is necessary. Consider maintenance, session complexity, request volume, and permission for each option.

How do I know when pagination is finished?

Use the response’s explicit completion field, missing next cursor, or documented page boundary. Add a maximum limit and log progress so a malformed stop condition cannot create an unbounded run.

What should I do when the server sends Retry-After?

Wait for the indicated date or number of seconds before the next request. Do not shorten the delay to make a crawl finish sooner.