ScreenshotNeo

BlogHow-to

How to Debug Web Scraping API Requests

A practical workflow for diagnosing authentication, status codes, timeouts, retries, pagination, and parsing bugs in web scraping API requests.

By the ScreenshotNeo team1 October 20269 min read

Debug a scraping API request in layers: first prove what was sent, then classify the HTTP response, then inspect redirects and timing, and only after that validate pagination and parsing. A 401 points to authentication, a 429 points to rate limiting, a timeout means the client stopped waiting, and a 200 response can still contain incomplete or invalid data.

1. Record the exact request before changing code

Most debugging gets harder when the failing request cannot be reproduced. Capture these fields for one failing attempt:

  • HTTP method and final URL
  • Query parameters and request body, with secrets redacted
  • Authentication method and the key or token scope (never the secret value)
  • Relevant headers, including Accept, Content-Type, User-Agent, and request IDs
  • Client timeout, elapsed time, redirect history, and retry count
  • HTTP status, response headers, response content type, and a bounded body sample
  • For list endpoints, requested limit or cursor, echoed pagination fields, and item count

Requests recommends setting an explicit timeout; without one, a request can wait indefinitely. See the Requests timeout documentation.

2. Make a minimal request outside your scraper

Remove concurrency, parsing, database writes, and browser automation. Reproduce the call with the smallest possible request. This separates an API or network problem from an application problem.

curl --verbose --max-time 30 \
  -H "Authorization: Bearer $SCRAPER_API_KEY" \
  -H "Accept: application/json" \
  "https://api.example.com/v1/items?limit=10"

Keep the verbose output in a private log. It can contain credentials in request headers. If the minimal call fails, fix transport or API usage first. If it succeeds, compare it byte for byte with the request produced by your application.

3. Read the status code and structured error together

Do not diagnose from the status line alone. Parse the response body as JSON when possible and retain its error type, message, field name, request ID, and retry information. A common managed scraping API mapping is:

Status Typical meaning First action
400 Malformed request or invalid parameter Check required fields, types, limits, cursors, and JSON syntax.
401 Missing, expired, or invalid credentials Check the authorization header, key value, environment, and scope.
402 Insufficient credits or plan allowance Check account balance, quota, and billing state.
403 Credentials are known but access is forbidden Check project permissions, endpoint access, IP policy, and target authorization.
404 Unknown endpoint or resource Check the base URL, API version, resource ID, and account ownership.
409 State conflict, often duplicate or already-running work Inspect the error details and use the documented idempotency or job status flow.
429 Rate limit exceeded Honor Retry-After when supplied and apply bounded backoff.
500 Provider-side internal error Retry safe requests, save the request ID, and contact the provider if it persists.

401 versus 403

A 401 normally means the service could not authenticate the request. Confirm that the header is present, uses the expected scheme (often Bearer), and is read from the environment you think is running. A 403 means authentication succeeded but the principal is not allowed to perform that operation or access that resource. Rotating a key will not fix a missing permission.

400 and pagination validation

Check parameter names and types before changing the target URL. Common causes include a string where an integer limit is required, a limit outside the documented range, an expired cursor, mutually exclusive filters, or a JSON body that does not match Content-Type. Log the server’s field-level error instead of replacing it with a generic exception.

4. Inspect redirects, transport errors, and response content

Transport success is separate from extraction success. Follow or disable redirects deliberately and inspect every hop. A redirect can send an API call to a login page, a different host, or an HTTP endpoint that strips authorization.

import requests

url = "https://api.example.com/v1/items"
headers = {"Authorization": "Bearer " + YOUR_API_KEY, "Accept": "application/json"}
try:
    response = requests.get(url, headers=headers, params={"limit": 10}, timeout=(5, 30), allow_redirects=True)
    print("status:", response.status_code)
    print("url:", response.url)
    print("redirects:", [(r.status_code, r.headers.get("Location")) for r in response.history])
    print("content-type:", response.headers.get("Content-Type"))
    print("request-id:", response.headers.get("X-Request-ID"))
    print("body:", response.text[:2000])
    response.raise_for_status()
except requests.exceptions.Timeout:
    print("The client stopped waiting; this does not prove the provider returned no data.")
except requests.exceptions.ConnectionError as exc:
    print("DNS, TCP, TLS, or proxy failure:", exc)
except requests.exceptions.HTTPError as exc:
    print("HTTP failure:", exc)

Use separate connect and read timeouts when the client supports them. A connection error occurs before an HTTP response exists; an HTTP error means the server did respond. Treat those paths differently in metrics and retries.

5. Validate the payload before parsing business fields

A 200 response can be an HTML login page, a provider error encoded as JSON, truncated output, or a valid envelope containing zero items. Check:

  • Content type and a bounded prefix of the body
  • Required top-level fields and their types
  • Item count against the requested page size
  • Pagination metadata such as next_cursor, has_more, or an equivalent field
  • Duplicate IDs and ordering across pages
  • Whether an empty result is valid for the query
payload = response.json()
if not isinstance(payload, dict):
    raise ValueError("Expected a JSON object")
items = payload.get("items")
if not isinstance(items, list):
    raise ValueError("Missing or invalid items array")
print("items:", len(items), "next_cursor:", payload.get("next_cursor"))

6. Debug pagination as its own subsystem

  1. Start with the smallest valid limit and one known filter.
  2. Save the first response, including the echoed cursor or page number.
  3. Request the next page using only the server-provided continuation value.
  4. Stop when the API says there is no next page; do not infer completion from a short page unless documented.
  5. Track IDs to detect repeated pages or cursor loops.
  6. Compare the total collected count with the provider’s total when one is returned.
cursor = None
seen = set()
all_items = []
while True:
    params = {"limit": 100}
    if cursor:
        params["cursor"] = cursor
    data = requests.get(url, headers=headers, params=params, timeout=30).json()
    page = data.get("items", [])
    ids = [item.get("id") for item in page if isinstance(item, dict)]
    if any(item_id in seen for item_id in ids if item_id is not None):
        raise RuntimeError("Pagination repeated an item; stop to avoid a loop")
    seen.update(item_id for item_id in ids if item_id is not None)
    all_items.extend(page)
    next_cursor = data.get("next_cursor")
    if not next_cursor:
        break
    cursor = next_cursor
print("collected:", len(all_items))

7. Retry only transient failures

Retry idempotent GET and HEAD requests. Retry a POST only when the API documents idempotency and you send its idempotency key. Use a maximum attempt count and a maximum elapsed time. Exponential backoff with jitter prevents every worker from retrying simultaneously.

import random, time, requests

TRANSIENT = {429, 500, 502, 503, 504}
for attempt in range(1, 5):
    try:
        r = requests.get(url, headers=headers, timeout=30)
        if r.status_code not in TRANSIENT:
            r.raise_for_status()
            data = r.json()
            break
        retry_after = r.headers.get("Retry-After")
    except (requests.exceptions.Timeout, requests.exceptions.ConnectionError):
        retry_after = None
    if attempt == 4:
        raise RuntimeError("request failed after bounded retries")
    delay = float(retry_after) if retry_after and retry_after.isdigit() else min(30, 2 ** (attempt - 1))
    time.sleep(delay + random.uniform(0, 0.25))

Never retry a 400, 401, 403, or 404 in a tight loop. Fix the request or credentials first. A retry loop should emit one log event per attempt with status, latency, and the next delay.

8. Use equivalent diagnostics in Node.js

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 30000);
try {
  const res = await fetch('https://api.example.com/v1/items?limit=10', {
    headers: { Authorization: `Bearer ${process.env.SCRAPER_API_KEY}`, Accept: 'application/json' },
    signal: controller.signal,
  });
  const text = await res.text();
  console.log({ status: res.status, contentType: res.headers.get('content-type'), requestId: res.headers.get('x-request-id'), body: text.slice(0, 2000) });
  if (!res.ok) throw new Error(`HTTP ${res.status}`);
  const data = JSON.parse(text);
  console.log('items:', data.items?.length, 'next:', data.next_cursor);
} catch (err) {
  console.error(err.name === 'AbortError' ? 'request timed out' : err);
} finally {
  clearTimeout(timer);
}

9. Keep diagnostic logs useful and safe

Store a redacted record containing timestamp, endpoint, method, status, latency, retry count, request ID, error type, error message, and a hash or small sample of the payload. Redact authorization headers, cookies, API keys, personal data, and full scraped pages. Hashing a payload lets you detect repeated responses without storing its contents.

10. Common failures and fixes

Symptom Likely cause Fix
401 immediately Wrong header, expired key, wrong environment Print the header name and key fingerprint only; verify scope and deployment secrets.
403 for one project Permission, IP allowlist, or resource ownership Compare a permitted project and inspect the structured error.
404 after changing version Wrong base path or resource identifier Use the provider’s documented version and confirm the resource exists.
429 under load Concurrency or quota exceeded Throttle workers, honor Retry-After, and use bounded backoff.
500 once Transient provider failure Retry the safe request and preserve the request ID.
Timeouts with no status Slow target, proxy, DNS, or connect/read timeout Separate connect/read timeouts, inspect network errors, and reduce concurrency.
200 but empty list Valid zero-result query, wrong filter, or wrong page Run a known-good query and validate pagination fields.
JSON parser error HTML, truncated body, or wrong content type Log content type and a bounded prefix before calling the parser.
Missing records Stopped after one page or cursor loop Follow continuation metadata and track IDs across pages.
Works in curl, fails in app Different headers, proxy, URL encoding, or environment secret Compare the fully rendered request and redirect chain.

11. Performance, reliability, and cost

  • Measure phases: record DNS/connect, time to first byte, total latency, response bytes, and parsing time when the client exposes them.
  • Control concurrency: start low, watch 429s and latency, then increase gradually. A large worker pool can reduce throughput when it triggers throttling.
  • Bound work: set request and overall job deadlines, cap retries, and persist a checkpoint after each page.
  • Prefer incremental collection: use cursors or modified-since filters when available instead of repeatedly downloading the full dataset.
  • Budget requests: count initial calls, pagination calls, retries, and asynchronous polling calls. A low per-call price can still become expensive when retries or page counts grow.
  • Make reruns safe: use stable item IDs, deduplicate writes, and idempotency keys for supported POST operations.

12. Or skip the browser setup with ScreenshotNeo

For workflows that need a rendered page image or PDF rather than extracted records, ScreenshotNeo provides a single GET request. The API documentation is at screenshotneo.com/docs/.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. Short FAQ

Should I retry a 401?

No. Verify the credential, header format, environment, and scope. Retry only after those inputs change.

Does a timeout mean the scrape failed?

It means your client stopped waiting. The remote job may still be running or may have completed, so use a job-status endpoint when the provider offers asynchronous runs.

Why can a 200 response be wrong?

HTTP status reports transport-level success. The body can still be an empty result, partial page, login document, or schema change.

How much response should I log?

Log status, headers needed for diagnosis, structured error fields, and a small redacted prefix or payload hash. Avoid full pages and secrets.

When should I use asynchronous scraping?

Use it when targets are slow, datasets are large, or a synchronous timeout would be too short. Poll with a deadline and persist the job ID so a worker restart can resume.