ScreenshotNeo

BlogGuides

Data Extraction Troubleshooting: Find and Fix the First Failure

Diagnose missing records, empty fields, 429s, parsing errors, and stuck queues by finding the earliest failed stage and fixing it with evidence.

By the ScreenshotNeo team29 September 202611 min read

Data Extraction Troubleshooting: Find and Fix the First Failure

When an extractor stops returning useful data, locate the earliest stage that fails: request and access, transport and limits, rendering, selection and parsing, pagination, queue execution, or validation and storage. Record the response and compare expected with actual output before changing selectors or adding retries. A successful HTTP status does not prove that the intended records were extracted.

This guide provides a stage-by-stage repair process for API clients, browser-based scrapers, CSV imports, and scheduled extraction jobs. It includes a bounded Python retry example, equivalent cURL and Node.js diagnostics, and a checklist for missing, duplicated, or malformed records.

1. Trace the extraction pipeline in order

Investigate stages in sequence. Fixing a downstream parser while the request is receiving a login page wastes time. For each run, retain the exact input, response metadata, and output counts so you can compare a failing run with a known-good one.

Trace the pipeline from request to storage and investigate the first stage that fails.
Trace the pipeline from request to storage and investigate the first stage that fails.
Stage Evidence to inspect Typical failure
Request and access URL, method, parameters, API version, authentication, permissions, headers Wrong endpoint, expired token, missing scope
Transport and limits Status, response body, request ID, timing, rate-limit headers 429, timeout, service error
Rendering Raw HTML versus browser DOM; relevant network calls JavaScript shell or delayed content
Selection and parsing Selector/path matches, delimiter, quoting, encoding, types Empty fields or malformed rows
Completeness Page/cursor sequence, stop condition, unique keys, counts Missing pages or duplicate records
Queue and schedule State, next execution time, retry history, external jobs Waiting item mistaken for a hung job
Validation and storage Null rates, schema, write errors, expected and observed counts Parsed data dropped or coerced during load

Start with a reproducible request

  1. Save the exact URL, HTTP method, query/body parameters, and API version.
  2. Verify credentials are present, current, and authorized for this resource. Confirm required headers and content type.
  3. Capture the status code, response body, request ID, relevant headers, and elapsed time. Redact secrets before sharing logs.
  4. Repeat once with a minimal request. Remove optional parameters and reduce page size to isolate the failing input.
  5. Check the response content type and a short, sanitized body sample. A 200 response can still contain an error document, login page, or empty application shell.

GitHub’s REST troubleshooting guide describes several easily misread cases: a private resource can return 404 when authentication is missing, a wrong method or URL can also produce 404, and missing or invalid parameters can produce 422. Check the endpoint documentation and permissions before treating the status code as the whole explanation. See GitHub’s REST API troubleshooting guide.

2. Diagnose 429s and temporary transport failures

A 429 means the service rejected the request under a limit, but the limit might be request rate, token or resource consumption, exhausted credits, or a spending/usage cap. Read the error body and code before retrying. OpenAI’s guidance, for example, distinguishes temporary rate limits from exhausted credit and enforced usage or spending limits; retrying a billing or quota failure will not restore access. Consult the relevant API provider’s limit documentation because limits and reset rules are service-specific. OpenAI’s 429 troubleshooting guidance.

Retry only when the failure can recover

  • Honor Retry-After when provided. For APIs with reset headers, wait until the stated reset time.
  • If no delay is specified and the error is plausibly temporary, use exponential backoff with random jitter.
  • Set a maximum attempt count and total retry duration. Do not retry authentication, validation, billing, or permission errors unchanged.
  • Reduce concurrency or request bursts. A steady pace can avoid short-window limits that an average-per-minute calculation hides.
  • Do not run two independent retry layers without accounting for their combined attempts.

GitHub advises waiting according to retry-after, then the rate-limit reset timestamp when remaining quota is zero, and otherwise at least a minute; continuing while rate-limited can lead to an integration ban. Other APIs specify their own behavior, so use the provider’s headers and documentation rather than copying one service’s exact wait policy. GitHub rate-limit instructions.

Python example: bounded GET with diagnostics

This example is for an API endpoint that returns JSON. Set the endpoint and authentication header to match its documentation. It honors a numeric Retry-After, uses bounded exponential backoff with jitter when absent, and stops on non-retryable HTTP errors. Do not blindly replay non-idempotent requests such as a POST that creates a record.

import random
import time
import requests

URL = "https://api.example.com/v1/records"
HEADERS = {"Authorization": "Bearer YOUR_TOKEN", "Accept": "application/json"}
MAX_ATTEMPTS = 5

with requests.Session() as session:
    for attempt in range(MAX_ATTEMPTS):
        try:
            response = session.get(URL, headers=HEADERS, timeout=(5, 30))
        except requests.RequestException as exc:
            if attempt == MAX_ATTEMPTS - 1:
                raise
            delay = min(30, 2 ** attempt) + random.uniform(0, 0.5)
            print(f"Transport failure: {exc}; retrying in {delay:.1f}s")
            time.sleep(delay)
            continue

        print("status:", response.status_code)
        print("request-id:", response.headers.get("x-request-id"))
        print("retry-after:", response.headers.get("retry-after"))
        print("body sample:", response.text[:500])

        if response.status_code == 429 or response.status_code >= 500:
            if attempt == MAX_ATTEMPTS - 1:
                response.raise_for_status()
            header = response.headers.get("Retry-After")
            try:
                delay = float(header) if header else min(30, 2 ** attempt)
            except ValueError:
                delay = min(30, 2 ** attempt)
            delay += random.uniform(0, 0.5)
            time.sleep(delay)
            continue

        response.raise_for_status()
        payload = response.json()
        print("top-level type:", type(payload).__name__)
        break

Some APIs return an HTTP-date rather than a number in Retry-After; adapt the parser to that service’s documented format. The example uses a fixed timeout and capped delay as safe starting points, not universal values. Add a total wall-clock deadline in production if the surrounding job has a strict completion window.

cURL: inspect status, headers, and body

curl --include --max-time 40 \
  --header "Authorization: Bearer YOUR_TOKEN" \
  --header "Accept: application/json" \
  "https://api.example.com/v1/records?limit=10"

--include prints response headers with the body; --max-time bounds the request duration. Avoid pasting a real credential into shell history or a shared ticket. cURL is useful for reproducing the server response outside your parser.

Node.js: capture response evidence

const url = 'https://api.example.com/v1/records?limit=10';
const res = await fetch(url, {
  headers: {
    Authorization: `Bearer ${process.env.API_TOKEN}`,
    Accept: 'application/json'
  },
  signal: AbortSignal.timeout(30000)
});

const body = await res.text();
console.log({
  status: res.status,
  requestId: res.headers.get('x-request-id'),
  retryAfter: res.headers.get('retry-after'),
  bodySample: body.slice(0, 500)
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${body.slice(0, 500)}`);
const data = JSON.parse(body);
console.log('top-level type:', Array.isArray(data) ? 'array' : typeof data);

3. Fix pages that load but yield empty fields

Compare the raw response your extractor receives with the page as rendered in a browser. If raw HTML has a small shell but the browser shows records, the data may be inserted by JavaScript after load. Inspect the browser’s network activity for a documented or permitted JSON endpoint. If no suitable endpoint exists, use browser automation and wait for a meaningful selector or page state rather than scraping before rendering completes. The Web Scraping with Python reference discusses rendered-page extraction using Selenium with Beautiful Soup.

Compare raw HTML with the rendered page before changing selectors or parsers.
Compare raw HTML with the rendered page before changing selectors or parsers.

When the rendered page contains the data but your output is empty, validate selection independently:

  1. Check that the selector matches at least one element on the current page.
  2. Log a small sanitized sample of the matched text or attribute before transformation.
  3. Confirm whether the value is in visible text, an attribute, embedded JSON, or a child element.
  4. Check whether the page has multiple similar elements and whether your selector targets the intended one.
  5. Wait for the specific content to appear. A fixed delay may hide a timing issue, but an explicit selector wait is easier to reason about.

Do not treat a blank screenshot or a successful page navigation as proof that extraction succeeded. A browser may show a consent prompt, bot check, login, or an empty state instead of the target content.

4. Repair parsing and CSV/schema errors

Reduce the input to the smallest sample that still fails. Inspect delimiters, quote escaping, embedded newlines, character encoding, header names, null representation, required columns, and inferred types. A valid CSV record can contain a newline inside a quoted field; splitting each physical line as a separate record will corrupt it. Likewise, a value that changes from numeric to text can conflict with a strict downstream schema.

Symptom Check Repair
Columns shift after a quoted value Delimiter and escaped quotes Use a CSV parser; preserve quote rules and reproduce with one row
Rows split unexpectedly Newlines inside quoted fields Parse records with a standards-aware reader, not line splitting
Accented characters are corrupted Input encoding and byte-order mark Set the documented encoding explicitly and normalize consistently
Load fails on one column Mixed numeric/text or date formats Inspect failing values; define an explicit schema and coercion policy
Missing values become literal text Null sentinel and empty-string handling Define how blank, null, and sentinel strings map to storage values

SAP’s data-services guidance lists invalid characters, unescaped quotes, embedded line breaks, and column type conflicts among CSV ingestion failure causes. Use the smallest failing sample to distinguish malformed CSV from schema rejection. SAP Data Services documentation.

5. Find missing or duplicate records

Pagination bugs often look like successful extraction with mysteriously incomplete data. Log every page number or cursor, the first and last key on each page, and the count returned. Verify that the next cursor changes, the terminal condition is correct, and the exporter follows all pages. Compare total records with an independent count if the source exposes one. GitHub notes that list endpoints commonly paginate results and that access permissions can also restrict which records are visible. GitHub’s missing-results guidance.

For duplicates, choose a stable source key and record where deduplication happens. Repeated pages, retries after a timeout, and overlapping time windows can all reintroduce records. Make storage writes idempotent where possible: use an upsert or unique constraint keyed by the source identifier, and retain an extraction run ID to trace each write. Do not discard duplicates silently until you know whether the source itself permits repeated keys.

6. Decide whether a queue is actually stuck

Inspect state transitions, scheduled execution time, retry reason, and the last external job or file response. A queued item with a future timestamp may be waiting by design. An item marked in progress can be waiting on an initial or range load, a remote job, or file arrival. Check worker health and the external dependency before manually replaying it; replaying an active or delayed item can create duplicate work.

SAP notes that queue entries may appear stuck when waiting for the next execution time or an external response. Compare the execution timestamp with the current time and timezone, then look for a retry or job dependency. SAP Data Services documentation.

7. Verify data quality and keep useful evidence

Before calling a repair complete, compare expected and observed row counts and review null rates, field lengths, duplicate-key counts, date parsing, and write failures. Keep one representative failing record and the parser/schema version. Add alerts for meaningful changes—such as a sudden drop in extracted rows or a rise in nulls—rather than only alerting on process crashes.

Attach this checklist to a repair ticket:

  • Exact endpoint, URL, method, parameters, API version, and authentication mode (never the secret itself).
  • Relevant request and response headers, HTTP status, error code/body sample, request ID, and timestamp with timezone.
  • Retry history, timing, cursor/page sequence, and queue state with next execution time.
  • Parser and schema versions, expected and actual row counts, null and duplicate counts.
  • A minimal sanitized input or failing record that reproduces the problem.

This evidence helps another engineer reproduce the failure and separates source changes from request, parser, scheduler, and storage changes.

8. Or skip the browser setup

If the failing stage is JavaScript rendering or you need a visual check of what a page actually presents, ScreenshotNeo is a website screenshot API and MCP server. Its GET request can return a PNG, JPEG, WebP, or PDF. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. A screenshot is useful for diagnosing page state; it does not replace structured data extraction.

One-call cURL example (replace the target URL and API key):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for configuration. Create a free account for 1,000 screenshots a month with no card.

9. Common troubleshooting cases

Error or symptom Likely cause First fix
401 or authentication error Missing, expired, malformed credential Check auth scheme, token expiry, and environment used by the job
404 for known resource Typo, wrong method/version, or private resource hidden from unauthorized caller Verify exact endpoint and method; confirm token access
403 or 429 Permission or rate/usage limit Read error code and headers; fix permissions or honor reset guidance
422 validation error Missing field, invalid type, or malformed parameter Compare types and required fields against endpoint documentation
200 with empty fields Wrong selector/path, login/error page, or JavaScript shell Inspect response and rendered DOM; test selector matches
Timeout Slow source, oversized request, stalled browser/network Bound request time; reduce page size; inspect service status and timing
CSV import rejection Quotes, embedded newline, encoding, or type conflict Minimize sample and validate syntax/schema independently
Queue appears frozen Future retry time or external dependency pending Inspect next execution timestamp and external job/file state

10. Performance, reliability, and cost

For API extraction, throughput is bounded by provider limits, source latency, parser capacity, and storage. Measure each separately. Use pagination and batch operations where supported, but avoid increasing concurrency as the first reaction to a backlog: burst traffic can trigger throttling and repeated failed calls consume time and quota. Bound retries and use jitter so multiple workers do not retry together.

For browser rendering, wait for a specific content condition rather than an unnecessarily long fixed delay. Capture only the data or page region required when the tool supports it, and avoid reloading identical pages when a safe cache is available. Cache validity depends on how quickly the source changes; a stale cached result can look like an extraction bug. Keep a timestamp and cache policy with each run.

Cost is not only API fees. Include retries, browser compute, queue delay, storage, and time spent repairing silent data corruption. Track cost per successful, validated record rather than requests attempted. Preserve failed-run evidence so a transient access issue does not lead to rerunning a large historical extraction without a reason.

FAQ

Why does my scraper return data yesterday but empty fields today?

First compare the raw response and rendered page, then test whether the selector or JSON path still matches. The source may have changed markup, begun rendering client-side, or returned a login or challenge page.

Should I retry every failed extraction?

No. Retry temporary transport failures and documented rate limits with bounded delays. Fix credentials, permissions, validation, or billing errors first; repeating the same request does not repair them.

How do I tell a slow queue from a stuck queue?

Compare its next execution timestamp and retry state with the current time, then check any external job or file it depends on. A future-dated item is often scheduled, not stalled.

Can a screenshot tell me whether structured extraction worked?

It can show the rendered visual state, including whether expected content appears. It cannot verify that your selectors, pagination, types, or stored records are correct; validate those separately.

Sources