ScreenshotNeo

BlogHow-to

How to Scrape Website Data with an API

Learn how to scrape website data with APIs, handle JavaScript pages, avoid blocks, validate results, and choose hosted or self-hosted crawlers.

By the ScreenshotNeo team1 October 202610 min read

Direct answer: Scrape website data with an API by first using the site’s own documented API, feed, search endpoint, or bulk export when one exists. If no suitable endpoint exists, use a hosted scraping API or a crawler such as Scrapy. Authenticate on the server, respect robots.txt and the site’s terms, throttle requests, handle pagination and retries, validate every record, and store raw responses when you need reproducibility.

An API can mean two different things:

  • A website’s data API: returns JSON, CSV, or another structured representation of the records you need.
  • A scraping API: accepts a URL and extraction instructions, then fetches pages for you. Some services render JavaScript or run browser workflows.

Prefer the first option whenever it covers your use case. Scrapy’s documentation summarizes why: “An API, a bulk export or a search endpoint is both faster for you and cheaper for the website than crawling its pages.”

1. Choose the access path

Option Use it when What you control
Official API The publisher exposes the records or operations you need Parameters, authentication, pagination, fields, rate limits
Feed or bulk export You need many records and the site publishes a stable dump Download schedule, incremental processing, validation
Search endpoint You need results matching a query rather than every page Queries, cursors, filters, deduplication
Hosted scraping API You need managed fetching, rendering, proxy or job infrastructure Request schema, extraction rules, polling, output storage
Self-hosted crawler You need custom parsing, workflows, or infrastructure control Requests, callbacks, concurrency, delays, storage, monitoring

Before writing a crawler, inspect the site’s developer documentation, page source, network requests, RSS or Atom feeds, sitemap, robots.txt, terms of service, authentication requirements, and data-use restrictions. An undocumented internal endpoint can change without notice and may have different access rules from a published API.

2. Confirm permission, scope, and data handling

  1. Read robots.txt. Scrapy advises: “Read the robots.txt file of the website.” Translate crawl-delay or request-rate directives into your downloader settings; Scrapy does not apply those directives automatically.
  2. Read the terms and privacy policy. Confirm that your collection, retention, redistribution, and geographic access are permitted.
  3. Define the smallest scope that answers your question: domains, URL patterns, fields, time range, and maximum pages.
  4. Identify personal, confidential, copyrighted, or regulated data before it enters logs or storage.
  5. Document your user agent and provide a contact address when the site’s policy requests one.

Authorization is still required when you use a scraping API. A service can automate HTTP requests or browser rendering; it does not grant permission to collect restricted data.

3. Call a documented JSON API

Use the endpoint and authentication method published by the site. The following example uses a placeholder endpoint because every API has different paths, fields, and credentials. Replace the URL, key name, and parameters with the target API’s documentation.

cURL

curl --fail-with-body \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Accept: application/json" \
  --get "https://example.com/api/products" \
  --data-urlencode "q=keyboard" \
  --data-urlencode "limit=100" \
  --data-urlencode "cursor=START_CURSOR"

Python

import os
import time
import requests

API_URL = "https://example.com/api/products"
headers = {
    "Authorization": f"Bearer {os.environ['API_TOKEN']}",
    "Accept": "application/json",
}
params = {"q": "keyboard", "limit": 100}
rows = []

with requests.Session() as session:
    while True:
        response = session.get(API_URL, headers=headers, params=params, timeout=30)
        response.raise_for_status()
        payload = response.json()

        page = payload.get("items", [])
        rows.extend(page)
        next_cursor = payload.get("next_cursor")
        if not next_cursor:
            break

        params["cursor"] = next_cursor
        time.sleep(0.25)

for row in rows:
    if not row.get("id"):
        raise ValueError(f"Missing id: {row!r}")

print(f"Fetched {len(rows)} records")

Node.js

const token = process.env.API_TOKEN;
const rows = [];
let cursor;

while (true) {
  const url = new URL('https://example.com/api/products');
  url.searchParams.set('q', 'keyboard');
  url.searchParams.set('limit', '100');
  if (cursor) url.searchParams.set('cursor', cursor);

  const response = await fetch(url, {
    headers: {
      Authorization: `Bearer ${token}`,
      Accept: 'application/json'
    }
  });

  if (!response.ok) {
    throw new Error(`HTTP ${response.status}: ${await response.text()}`);
  }

  const payload = await response.json();
  rows.push(...(payload.items || []));
  cursor = payload.next_cursor;
  if (!cursor) break;

  await new Promise(resolve => setTimeout(resolve, 250));
}

for (const row of rows) {
  if (!row.id) throw new Error(`Missing id: ${JSON.stringify(row)}`);
}
console.log(`Fetched ${rows.length} records`);

Keep API keys in environment variables or a secret manager. Never commit them, place them in browser JavaScript, include them in screenshots, or expose them in public query strings unless the provider explicitly requires that method.

4. Extract fields from HTML with Scrapy

When there is no usable data endpoint, a crawler can download HTML and parse fields. Scrapy requests are downloaded into responses; callbacks extract fields and can yield more requests for pagination or detail pages.

Install and create a spider

python -m venv .venv
. .venv/bin/activate
pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.com

Replace the generated spider with this example. Adjust selectors to the target site’s markup.

import scrapy


class ProductsSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/products"]

    custom_settings = {
        "ROBOTSTXT_OBEY": True,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 2,
        "DOWNLOAD_DELAY": 0.5,
        "AUTOTHROTTLE_ENABLED": True,
        "AUTOTHROTTLE_START_DELAY": 1.0,
        "AUTOTHROTTLE_MAX_DELAY": 10.0,
        "FEEDS": {
            "products.jsonl": {
                "format": "jsonlines",
                "encoding": "utf8",
                "overwrite": True,
            }
        },
    }

    def parse(self, response):
        for card in response.css("article.product"):
            yield {
                "name": card.css("h2::text").get(default="").strip(),
                "price": card.css(".price::text").get(default="").strip(),
                "url": response.urljoin(card.css("a::attr(href)").get()),
                "source_url": response.url,
            }

        next_href = response.css("a[rel='next']::attr(href)").get()
        if next_href:
            yield response.follow(next_href, callback=self.parse)

Run it with:

scrapy crawl products

The Scrapy optimization guidance covers concurrency and broad-crawl behavior. Review the project’s robots.txt settings and configure them explicitly for your crawl.

5. Handle JavaScript-rendered pages

First check whether the browser obtains the data from a documented JSON endpoint. Calling that endpoint is usually cheaper and more reliable than rendering the page. If rendering is required, use a crawler integration or hosted service that explicitly supports browser execution.

Document the consequences of rendering:

  • Higher latency and resource usage than a plain HTTP request.
  • More failures from scripts, consent dialogs, browser limits, and third-party resources.
  • Different terms or authentication requirements.
  • Need for deterministic waits, selectors, and page-state checks.

Wait for a meaningful condition such as a result selector or network idle, rather than sleeping for an arbitrary long duration. Capture the final URL, status, key fields, and a failure reason so a blank page is not mistaken for an empty result.

6. Hosted API versus self-hosted crawler

Question Hosted service Self-hosted Scrapy
Coverage Depends on supported domains, browser modes, and anti-bot handling You choose requests, integrations, and infrastructure
Rendering May provide managed browser execution You operate browser workers and their dependencies
Control Bounded by request schema, selectors, limits, and policies Full control of callbacks, headers, cookies, retries, and schemas
Operations Provider operates capacity, upgrades, and much of the monitoring You own deployments, proxies, browser capacity, alerts, and upgrades
Output May include JSON, CSV, JSONL, webhooks, or warehouse connectors You implement exporters and delivery
Scheduling Some products include schedules and run polling You provide a scheduler and job state
Cost Per request, result, compute unit, or subscription, depending on provider Infrastructure plus engineering and maintenance time

For managed platforms, the normal workflow is tool discovery, request submission, recording a run ID, polling status, and exporting dataset rows. Confirm the provider’s exact request and output schema before coding against it.

7. Pagination, deduplication, and incremental crawls

Pagination

  • Prefer opaque cursors supplied by the API over page numbers.
  • Stop when the provider returns no cursor or an explicit completion flag.
  • Protect against a repeated cursor or a cursor that never advances.
  • Record the final cursor and item count for each run.

Deduplication

Choose a stable key such as the API’s record ID or canonical URL. Normalize URLs before comparing them, and keep the source URL and retrieval timestamp. If the source has no stable ID, hash the normalized fields you consider identifying and keep the raw record for review.

Incremental collection

Use an API’s updated-since filter, change feed, or cursor when available. Otherwise maintain a frontier of URLs and a last-seen timestamp. Re-fetch enough overlap to detect late updates, then deduplicate during loading.

8. Rate limits, blocks, and retries

Start conservatively and increase concurrency only while latency and response codes remain healthy. Rising 429, 503, or ban-page counts indicate that the crawl has exceeded a tolerated rate.

  • 429: honor Retry-After when present, apply exponential backoff with jitter, and reduce concurrency.
  • 401 or 403: verify credentials, scopes, expiration, and permission. Do not attempt to bypass an access control.
  • 5xx or network timeout: retry idempotent GET requests with a bounded budget.
  • POST retries: retry only when the operation is documented as idempotent or protected by an idempotency key.
  • Ban or challenge page: stop or slow the job, inspect the response, and contact the site or provider if access is authorized.

Use structured error types in addition to HTTP status codes. Log request ID, URL or endpoint, attempt number, status, latency, and a redacted error body.

9. Validate and store results

Validation prevents a successful HTTP response from becoming bad data.

  • Check required fields, types, ranges, and timestamps.
  • Check that pagination completed and that counts are plausible.
  • Detect duplicate IDs and canonical URLs.
  • Preserve the source URL, retrieval time, parser version, and run ID.
  • Keep raw responses or content hashes when you need reproducibility.
  • Quarantine malformed records instead of silently dropping them.
  • Separate transient failures from genuinely empty results.

10. Performance, reliability, and cost

Performance

  • Use the official API, bulk export, or search endpoint before crawling HTML.
  • Request only needed fields and use server-side filters.
  • Reuse HTTP connections and compress responses where supported.
  • Throttle per domain rather than applying one global rate.
  • Cache immutable pages and avoid re-downloading unchanged assets.
  • Render JavaScript only for pages that require it.

Reliability

  • Make jobs restartable with checkpoints or cursors.
  • Use bounded retries and a dead-letter queue for persistent failures.
  • Alert on completion rate, validation failures, 429/503 counts, and unexpected schema changes.
  • Store raw evidence for records that affect business decisions.

Cost

Compare API charges, proxy or browser compute, storage, scheduler and monitoring costs, and engineering time. A slower crawl can cost more than a documented endpoint because it transfers more bytes and requires more retries. Do not assume that a hosted service or self-hosting is universally cheaper; measure your request volume and rendering needs.

11. Troubleshooting

Symptom Likely cause Fix
401 Unauthorized Missing, expired, or incorrectly scoped credential Check the documented header or token format and rotate the secret if needed
403 Forbidden Account, IP, region, or resource is not permitted Verify authorization and terms; request access instead of bypassing controls
429 Too Many Requests Rate or concurrency is too high Honor Retry-After, back off with jitter, and lower per-domain concurrency
Empty JSON list Wrong filter, cursor, field, or an actual empty result Log the full redacted request, inspect pagination metadata, and test a known record
HTML contains no data Data is rendered by JavaScript or loaded from another endpoint Find the documented data endpoint or use a renderer that supports the required workflow
Duplicate records Overlapping pages, unstable sorting, or retries Use a stable key, deterministic ordering, and idempotent loading
Parser suddenly returns nulls Markup or API schema changed Version selectors, validate required fields, and alert on schema drift
Timeouts Slow origin, heavy assets, or browser scripts Set bounded timeouts, wait for a specific condition, retry safely, and reduce page scope

12. Or skip the browser setup

If your goal is a visual record of a page rather than structured fields, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options, including full-page capture with lazy images, CSS-element capture, dark mode, device presets, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, usage, and the OpenAPI specification.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. Practical checklist

  • Checked for an official API, feed, search endpoint, or bulk export.
  • Read robots.txt, terms, authentication requirements, and data restrictions.
  • Stored credentials outside source code and browser code.
  • Defined URL scope, fields, pagination, and a stable deduplication key.
  • Set conservative concurrency, delays, timeouts, and retry limits.
  • Handled 401, 403, 429, 5xx, timeouts, and challenge pages.
  • Validated required fields, counts, duplicates, timestamps, and source URLs.
  • Saved raw responses or hashes when reproducibility matters.
  • Added monitoring for rate limits, schema changes, and incomplete runs.

FAQ

Can I scrape any website with an API?

No. An API changes how requests are made; it does not remove authorization, robots.txt, terms, privacy, or copyright obligations.

Should I use an API or scrape HTML?

Use the official API, feed, search endpoint, or export when it contains the data you need. Crawl HTML only when those paths are unavailable or insufficient.

How do I scrape a JavaScript-heavy site?

Identify the documented data endpoint first. If the page truly requires rendering, use a browser-capable crawler or hosted service and add explicit waits, failure checks, and higher resource limits.

How do I avoid getting blocked?

Stay within permission, obey published limits, identify your client, throttle per domain, cache results, use narrow scopes, and stop when 429 or challenge responses rise.

What should I do with failed records?

Keep them separate from valid empty results, record the error and attempt metadata, retry safe requests with a bound, and review persistent failures before rerunning.