ScreenshotNeo

BlogEngineering

How to Extract Web Data with an Asynchronous Crawler API

Submit crawl jobs, poll safely, handle JavaScript pages, validate results, and choose between hosted APIs and Scrapy.

By the ScreenshotNeo team1 October 20267 min read

Short answer: submit a crawl or extraction request, save the returned run ID, poll its status with bounded exponential backoff (or receive a callback), download the result when complete, validate the records, and persist them with the run metadata. Use direct HTTP extraction when the required data is in the server response. Use browser rendering when JavaScript creates the content you need.

1. The asynchronous crawler API workflow

  1. Submit work. Send the target URL and extraction settings to the provider’s asynchronous endpoint.
  2. Persist the run. Store the run ID, URL, requested options, creation time, and an idempotency key generated by your application.
  3. Monitor completion. Poll the documented status endpoint with bounded exponential backoff, or process a provider callback when available.
  4. Retrieve results. Download structured output or dataset items after the run reaches a terminal state.
  5. Validate and persist. Check the schema, required fields, source URL, timestamps, and duplicate keys before writing to your warehouse or application database.
  6. Classify failures. Separate transient network and rate-limit errors from rendering, parsing, and permanent access failures. Retry only idempotent transient operations and retain the original run ID and error payload.
Phase Data to keep Typical failure
Submit Request payload, idempotency key, provider response Authentication, validation, rate limit
Poll Run ID, status history, next-attempt time Timeout, throttling, transient network error
Retrieve Dataset URL or result body, checksum, content type Expired result, partial output
Validate Schema version, source URL, extraction timestamp Missing fields, duplicate records

2. HTTP extraction versus browser rendering

Choose HTTP extraction when the server response already contains the HTML or JSON you need. It is usually simpler and faster because it does not start a browser. Choose browser rendering when a script changes the DOM or fetches data after page load. As Zyte documents, “HTTP responses do not reflect HTML content rendered by a web browser that executes JavaScript code.”

  • HTTP mode: best for server-rendered pages, feeds, and JSON endpoints.
  • Browser mode: required for client-rendered content, interactions, or pages whose data appears only after JavaScript execution.
  • Hybrid strategy: try HTTP first, then route known client-rendered URL patterns to browser workers.

Hosted extraction APIs can bundle authentication, proxy and IP controls, geolocation, cookies, sessions, browser automation, and automatic extraction. Self-managed Scrapy gives you control over spiders, scheduling, parsing, deployment, and data contracts, but your team owns the scheduler, storage, browser and proxy layer, observability, and failure handling.

3. Submit a job and poll it

The exact field names differ by provider. Keep the provider-specific mapping in one adapter so the rest of your pipeline uses a stable internal model.

cURL template

#!/usr/bin/env bash
set -euo pipefail

: "${CRAWLER_API_URL:?Set CRAWLER_API_URL to your provider's async submit endpoint}"
: "${CRAWLER_API_KEY:?Set CRAWLER_API_KEY}"

submit=$(curl --fail-with-body -sS -X POST "$CRAWLER_API_URL" \
  -H "Authorization: Bearer $CRAWLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","extraction":"html"}')

echo "$submit"
# Save the provider's run ID from this response, then call its documented
# status endpoint until the run is complete. Field names vary by API.

Zyte documents an extraction endpoint at https://api.zyte.com/v1/extract. Use that endpoint and its authentication and request schema when implementing a Zyte adapter.

Python: asynchronous client with bounded backoff

import os
import time
from typing import Any

import requests

SUBMIT_URL = os.environ["CRAWLER_SUBMIT_URL"]
STATUS_URL_TEMPLATE = os.environ["CRAWLER_STATUS_URL_TEMPLATE"]
API_KEY = os.environ["CRAWLER_API_KEY"]


def submit(url: str, extraction: str) -> dict[str, Any]:
    response = requests.post(
        SUBMIT_URL,
        headers={"Authorization": f"Bearer {API_KEY}"},
        json={"url": url, "extraction": extraction},
        timeout=30,
    )
    response.raise_for_status()
    return response.json()


def wait_for_completion(run_id: str, max_wait: int = 900) -> dict[str, Any]:
    deadline = time.monotonic() + max_wait
    delay = 1.0
    terminal = {"completed", "failed", "cancelled"}

    while time.monotonic() < deadline:
        status_url = STATUS_URL_TEMPLATE.format(run_id=run_id)
        response = requests.get(
            status_url,
            headers={"Authorization": f"Bearer {API_KEY}"},
            timeout=30,
        )
        if response.status_code == 429:
            time.sleep(min(delay * 2, 60))
            delay = min(delay * 2, 60)
            continue
        response.raise_for_status()
        payload = response.json()
        state = payload.get("status")
        if state in terminal:
            return payload
        time.sleep(delay)
        delay = min(delay * 2, 30)

    raise TimeoutError(f"Run {run_id} did not finish within {max_wait} seconds")


submission = submit("https://example.com", "html")
run_id = submission["run_id"]  # Map this once in your provider adapter.
result = wait_for_completion(run_id)
if result.get("status") != "completed":
    raise RuntimeError(result)
print(result)

Node.js: submit, poll, and stop on terminal state

const submitUrl = process.env.CRAWLER_SUBMIT_URL;
const statusTemplate = process.env.CRAWLER_STATUS_URL_TEMPLATE;
const apiKey = process.env.CRAWLER_API_KEY;

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

async function request(url, options = {}) {
  const response = await fetch(url, {
    ...options,
    headers: {
      Authorization: `Bearer ${apiKey}`,
      'Content-Type': 'application/json',
      ...(options.headers || {}),
    },
  });
  if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
  return response.json();
}

const submission = await request(submitUrl, {
  method: 'POST',
  body: JSON.stringify({ url: 'https://example.com', extraction: 'html' }),
});
const runId = submission.run_id;

let delay = 1000;
for (let attempt = 0; attempt < 12; attempt += 1) {
  const status = await request(statusTemplate.replace('{run_id}', encodeURIComponent(runId)));
  if (['completed', 'failed', 'cancelled'].includes(status.status)) {
    if (status.status !== 'completed') throw new Error(JSON.stringify(status));
    console.log(status);
    break;
  }
  await sleep(delay);
  delay = Math.min(delay * 2, 30000);
}

4. Idempotency, retries, and result storage

  • Create an idempotency key from your logical job ID, target URL, extraction version, and crawl date. Persist it before submission.
  • On a network timeout after submission, retry with the same key instead of creating an unrelated run.
  • Use exponential backoff with a maximum delay and an overall deadline. Honor Retry-After when the provider sends it.
  • Store raw provider output before transforming it. Keep the run ID, request payload, response headers, status history, and error body for audits.
  • Make result writes idempotent with a natural key such as source URL plus source timestamp or provider item ID.

5. JavaScript-rendered pages

  1. Determine whether the required field exists in the initial HTML or only after scripts run.
  2. Use browser rendering for the latter case.
  3. Wait for a meaningful selector or network-idle condition rather than an arbitrary long sleep when the provider supports it.
  4. Validate that the rendered result contains the required field; an HTTP 200 response alone does not prove extraction succeeded.

Browser runs cost more time and resources than direct HTTP requests. Limit rendering to URL classes that need it, and reuse sessions only when the provider documents safe session behavior.

6. Scaling and performance

Technique Why it helps Trade-off
Bounded concurrency Uses available capacity without triggering rate limits Requires queue and back-pressure
HTTP-first routing Avoids browser startup for server-rendered pages Needs a reliable classification rule
Selector waits Finishes as soon as required content exists Breaks when selectors change
Batch submission Reduces client overhead One bad item may complicate retries
Raw-result retention Enables replay and debugging Requires storage and retention policy

Measure queue wait, provider processing time, polling latency, result download time, parse time, retry count, and terminal failure class. Set separate budgets for submission, polling, and result download so one stalled run cannot consume all worker time.

7. Reliability, access, and compliance

  • Review authorization, robots directives, terms of service, rate limits, and personal-data obligations before crawling.
  • Use the minimum concurrency that meets your freshness requirement.
  • Record the effective URL, response timestamp, and extraction version.
  • Handle login, cookies, geolocation, and proxy requirements through documented provider features or your own controlled browser layer.
  • Do not treat retries as a solution to permanent access denials, CAPTCHAs, or policy blocks.

8. Troubleshooting

Symptom Likely cause Fix
401 or 403 on submit Invalid credentials or missing permission Check the key, account permissions, and required authentication scheme.
429 responses Rate limit or excessive concurrency Honor Retry-After, reduce concurrency, and use bounded backoff.
Run never completes Provider queue delay or stalled page Set an overall deadline, record the run ID, and inspect status and provider logs.
Empty fields Wrong extraction type or JavaScript content Verify the schema and switch the URL to browser rendering when needed.
Duplicate records Retry created a second run Use an idempotency key and a database uniqueness constraint.
Parser errors Page layout changed or response is an error document Store the raw response, validate content type, and version your parser.
Successful status but missing data Transport succeeded while extraction validation failed Require fields and value types in post-run validation before publishing data.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It is useful when your asynchronous pipeline needs a visual artifact rather than structured fields. One GET request returns a PNG, JPEG, WebP, or PDF.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get started.

10. FAQ

When should I poll versus use a webhook?

Poll when you need a simple integration or the provider has no callback. Use a signed webhook when runs are long-lived or numerous, and still verify the run status before storing results.

Can an asynchronous API crawl multiple URLs?

Usually yes, through batch requests or multiple run submissions. Keep per-URL status and retry state so one failure does not hide successful items.

Is Scrapy asynchronous?

Scrapy exposes crawl_async() and asyncio-compatible runner classes. It is a code-controlled option when your team wants to own scheduling, parsing, deployment, and storage.

What should a result contract contain?

Define required fields, types, source URL, extraction timestamp, parser version, and a stable deduplication key. Reject or quarantine records that fail validation.