Web Scraping API: How to Extract Data with REST, Python, and PHP
Learn how to call a web scraping API with REST, Python, PHP, and Node.js, handle JavaScript pages, pagination, rate limits, retries, and costs.
A web scraping API is an HTTPS service that accepts a target URL or scraping job and returns machine-readable data, rendered HTML, text, Markdown, screenshots, or structured JSON. You call it like any REST endpoint: authenticate, send the URL and options, check the response, parse the result, and retry transient failures.
The safest production pattern is:
- Choose a provider and endpoint that supports the rendering and extraction you need.
- Keep the API key in a server-side environment variable or secret manager.
- Send the key in an
Authorization: Bearerheader when supported. - Set connect and read timeouts.
- Check the HTTP status and content type before parsing.
- Persist pagination checkpoints and retry 429 and transient 5xx responses with bounded exponential backoff.
What a web scraping API does
Most services expose one of two models:
- Synchronous request: your request waits for the page to load and returns the result in the same response.
- Asynchronous job: you submit a job, receive an ID, and poll or receive a webhook when the result is ready. This is better for large crawls and bulk datasets.
Rendering varies by provider. Some fetch only the initial HTML. Others run a real browser so JavaScript-generated content appears in the response. Extraction may return raw HTML, visible text, Markdown, screenshots, or provider-defined JSON fields.
Apify’s REST API documentation describes RESTful HTTP endpoints, JSON responses, bearer authentication, Actors, datasets, and rate limits. ScrapingBee’s documentation covers rendered HTML, text, Markdown, screenshots, structured extraction, and JavaScript execution. Bright Data’s Web Scraper API focuses on prebuilt site datasets and synchronous or asynchronous bulk jobs.
Before you send a request
Check authorization and site rules
API access does not override a site’s terms, robots directives, authentication boundary, copyright restrictions, or applicable law. Use examples and targets you are authorized to access. Do not attempt to defeat login controls, CAPTCHAs, paywalls, or technical restrictions without permission.
Define the output you need
| Need | Useful API capability |
|---|---|
| Simple metadata | HTML or JSON response |
| JavaScript-rendered content | Headless-browser rendering |
| Article text | Text, Markdown, or extraction rules |
| Many URLs | Async jobs, datasets, bulk endpoints, pagination |
| Visual verification | Screenshot or PDF output |
First REST request with cURL
Replace the endpoint and parameter names with those documented by your provider. The example keeps the key out of the URL and sends it in a header.
export SCRAPER_API_KEY='replace-me'
curl --fail-with-body --silent --show-error \
--connect-timeout 10 --max-time 60 \
-H "Authorization: Bearer ${SCRAPER_API_KEY}" \
-H "Accept: application/json" \
--get 'https://api.example.com/v1/scrape' \
--data-urlencode 'url=https://example.com'
Use --data-urlencode so query characters in the target URL are encoded correctly. If the provider requires POST, send a JSON body instead:
curl --fail-with-body --silent --show-error \
--connect-timeout 10 --max-time 60 \
-H "Authorization: Bearer ${SCRAPER_API_KEY}" \
-H 'Content-Type: application/json' \
-d '{"url":"https://example.com","render_js":true}' \
'https://api.example.com/v1/scrape'
Python: requests with timeouts, JSON checks, and retries
Python’s requests library provides query parameters, headers, JSON bodies, status checking, TLS verification, and reusable sessions with connection pooling.
import os
import random
import time
import requests
ENDPOINT = "https://api.example.com/v1/scrape"
API_KEY = os.environ["SCRAPER_API_KEY"]
def scrape(url, attempts=4):
headers = {
"Authorization": f"Bearer {API_KEY}",
"Accept": "application/json",
}
with requests.Session() as session:
for attempt in range(attempts):
try:
response = session.get(
ENDPOINT,
params={"url": url},
headers=headers,
timeout=(10, 60),
)
except requests.RequestException:
if attempt == attempts - 1:
raise
time.sleep(min(30, 2 ** attempt + random.random()))
continue
if response.status_code == 429 or 500 <= response.status_code < 600:
if attempt == attempts - 1:
response.raise_for_status()
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after and retry_after.isdigit() else min(30, 2 ** attempt + random.random())
time.sleep(delay)
continue
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "json" not in content_type.lower():
return {"content_type": content_type, "body": response.text}
return response.json()
raise RuntimeError("request failed after retries")
print(scrape("https://example.com"))
For a POST provider, replace session.get(..., params=...) with session.post(..., json={...}). Keep the same timeout, status handling, and retry policy.
PHP: portable cURL integration
<?php
$target = 'https://example.com';
$endpoint = 'https://api.example.com/v1/scrape?url=' . rawurlencode($target);
$key = getenv('SCRAPER_API_KEY');
$ch = curl_init($endpoint);
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . $key,
'Accept: application/json',
],
CURLOPT_CONNECTTIMEOUT => 10,
CURLOPT_TIMEOUT => 60,
]);
$body = curl_exec($ch);
if ($body === false) {
throw new RuntimeException(curl_error($ch));
}
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$contentType = curl_getinfo($ch, CURLINFO_CONTENT_TYPE) ?: '';
curl_close($ch);
if ($status < 200 || $status >= 300) {
throw new RuntimeException("Scraping API returned HTTP $status: $body");
}
if (stripos($contentType, 'json') !== false) {
$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
var_dump($data);
} else {
echo $body;
}
Apify also documents a PHP client option. A portable cURL integration is useful when you want to avoid a provider-specific SDK. Never place the key in browser JavaScript or commit it to source control.
Node.js: fetch and response validation
const endpoint = 'https://api.example.com/v1/scrape';
const key = process.env.SCRAPER_API_KEY;
const target = 'https://example.com';
const url = new URL(endpoint);
url.searchParams.set('url', target);
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 60_000);
try {
const response = await fetch(url, {
headers: {
Authorization: `Bearer ${key}`,
Accept: 'application/json'
},
signal: controller.signal
});
if (!response.ok) {
throw new Error(`Scraping API returned HTTP ${response.status}: ${await response.text()}`);
}
const contentType = response.headers.get('content-type') || '';
const result = contentType.includes('json')
? await response.json()
: await response.text();
console.log(result);
} finally {
clearTimeout(timer);
}
Authentication and secret handling
Prefer an HTTP authorization header. Apify says it recommends header authentication because it is more secure, and ScrapingBee documents the Bearer authorization header as its recommended method. ScrapingBee marks query-string API keys as deprecated.
- Load keys from environment variables, a cloud secret manager, or a CI secret store.
- Restrict key permissions and rotate keys after suspected exposure.
- Redact authorization headers and full URLs from logs.
- Make API calls from your server, worker, or scheduled job, never from untrusted browser code.
JavaScript pages, proxies, and extraction modes
If the required data is inserted after page load, request browser rendering or JavaScript execution. Rendering costs more time and, with some providers, more credits. ScrapingBee's documented examples list rotating proxy without JavaScript at 1 credit, rotating proxy with JavaScript at 5 credits, premium proxy without JavaScript at 10 credits, premium proxy with JavaScript at 25 credits, and stealth proxy with JavaScript at 75 credits. Verify current pricing before committing to an architecture.
Proxy and anti-bot capability should match your authorized use case. Compare geographic coverage, JavaScript support, proxy tiers, extraction formats, and provider limits. Do not assume a proxy makes an unauthorized crawl permissible.
Pagination and restartable crawls
Providers expose pagination differently: page numbers, cursors, dataset offsets, or job result URLs. Persist the cursor after each successful page so a process restart does not duplicate or lose records.
cursor = None
while True:
params = {"url": "https://example.com/catalog", "limit": 100}
if cursor:
params["cursor"] = cursor
page = scrape_page(params) # provider-specific function
save_records(page["items"])
cursor = page.get("next_cursor")
save_checkpoint(cursor)
if not cursor:
break
Use a stable sort key where the provider supports one. Deduplicate by the source record ID or canonical URL, and record the crawl timestamp.
Rate limits, 429 responses, and retries
HTTP 429 means the provider is throttling you. Respect Retry-After and any documented rate headers. Otherwise use bounded exponential backoff with jitter, such as 1, 2, 4, and 8 seconds plus a random fraction. Retry connection failures and transient 5xx responses; do not blindly retry authentication errors, malformed requests, or a blocked target.
Apify documents a global limit of 250,000 requests per minute and a default per-resource limit of 60 requests per second in its API v2 reference. These limits are provider-specific and can change, so read the current documentation and design for lower effective concurrency.
Async jobs and bulk extraction
Use asynchronous jobs when a crawl can exceed a request timeout, when you need hundreds of URLs, or when the provider offers dataset export. Submit the job, persist its ID, poll with backoff or register a webhook, then download results and mark the job complete only after validation. Bright Data documents synchronous and asynchronous bulk jobs and prebuilt site datasets.
Make job processing idempotent. Store the input URL, provider job ID, attempt count, status, result checksum, and last error. A worker should safely resume after interruption.
Or skip the browser setup
When your goal is a clean screenshot or PDF rather than extracted records, ScreenshotNeo provides a single GET request. Its API accepts a URL and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. The MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and try the API with 1,000 monthly screenshots at no cost.
Performance and reliability checklist
- Reuse HTTP sessions or keep-alive connections.
- Set separate connect and read timeouts.
- Request only the rendering mode and fields you need.
- Cache immutable URLs and use provider caching where available.
- Limit concurrency to the provider's documented quotas.
- Use queues for large jobs and persist checkpoints.
- Capture status, latency, provider request ID, response size, and failure category.
- Validate records before publishing them downstream.
Cost planning
Pricing may be per request, credit, rendered page, proxy tier, browser minute, dataset row, or completed job. JavaScript rendering and premium or stealth proxies commonly consume more credits. Estimate the number of URLs, retry rate, pages per URL, rendering mode, and retention period. Include failed requests in your operational budget even when a provider has a separate policy for billing them.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, expired, or incorrectly formatted key | Check the secret, use the documented Bearer header, and confirm account permissions. |
| 400 | Wrong parameter name or malformed URL | URL-encode the target and compare the payload with the provider schema. |
| 429 | Rate limit exceeded | Honor Retry-After, reduce concurrency, and add exponential backoff. |
| HTML lacks visible content | Page requires JavaScript | Enable browser rendering or JavaScript execution. |
| Timeout | Slow target, heavy assets, or overloaded render job | Increase read timeout within provider limits, reduce scope, or use an async job. |
| JSON parsing error | Error page or HTML returned with a success status | Inspect Content-Type and retain the raw body before calling a JSON parser. |
| Duplicate records | Retry or pagination restart repeated a page | Use checkpoints and deduplicate by a stable source ID or canonical URL. |
| Incomplete crawl | Process stopped without saving its cursor | Persist the cursor after every successful page and resume from it. |
FAQ
Can a scraping API bypass a login?
Only use credentials and authenticated endpoints when you are authorized. An API does not grant permission to cross an authentication boundary.
Should I parse HTML or request JSON?
Request provider-structured JSON when its schema matches your needs. Keep HTML or text when you need custom parsing or the provider's extraction is insufficient.
When should I choose an asynchronous job?
Choose async execution for long pages, large URL sets, bulk datasets, or workflows that must survive worker restarts.
How do I avoid leaking API keys?
Use server-side environment variables or secret managers, authorization headers, redacted logs, restricted permissions, and regular rotation.
Which screenshot API should I try?
ScreenshotNeo is the first option to try for clean screenshots because it removes consent banners and overlays before capture, bills only clean shots, and has a $5 paid plan.


