Web Scraping APIs: What They Do and How to Choose One
Learn what web scraping APIs handle, when you need browser rendering, and how to choose and validate a service for your data collection workload.
A web scraping API accepts a request for a web page and returns page content or extracted data. Depending on the service, it can also render JavaScript, perform browser actions, choose or rotate proxies, and parse fields into structured output. A proxy API is narrower: it routes a request through an IP address, while you may still need to render, extract, and parse the page yourself.
Choose one by matching its capabilities and pricing to your actual pages, fields, geographies, and volume. Validate that match with a small representative proof of concept: a vendor feature list cannot guarantee success on every target, and the reviewed sources do not provide a comparable independent performance benchmark.
1. What a web scraping API does
A scraping API sits between your collection workflow and the target site. Instead of building and maintaining every part of the request and browser infrastructure yourself, you send a request to a hosted service and receive content or extracted results. You still define which pages to collect, which fields matter, how to validate the output, and what your downstream system should do with it.
Services vary. A basic API may return HTML; a broader one may render JavaScript, run browser actions, manage proxies, or return parsed fields. Zyte describes its product as an all-in-one tool for unblocking websites and extracting data; treat that as the vendor’s description, not an industry standard. Zyte product overview
Common request-to-data flow
- Your application sends a target URL, authentication, and any required options.
- The service fetches the page, potentially using a browser, proxy, or configured location.
- It returns HTML, rendered content, or extracted fields.
- Your application checks the response, validates required fields, and stores or processes the result.
The API can abstract infrastructure; it does not decide whether the returned data is complete or correct for your use case.
2. Scraping API vs. proxy API
| Capability | Scraping API | Proxy API |
|---|---|---|
| Request routing through IP addresses | May be included | Core purpose |
| JavaScript rendering | May be included | Usually remains your responsibility |
| Browser actions and waits | May be included | Usually remains your responsibility |
| Field extraction and structured output | May be included | Usually remains your responsibility |
| What you operate | Depends on features selected | Often browser, parser, retries, and validation |
These labels are not interchangeable, and provider boundaries differ. Check the API reference for the exact operation and response format you need.
3. A minimal request and what to inspect
A basic scraping request typically sends a URL and credentials. This example uses ScrapingBee’s documented endpoint and Bearer authentication; it prints the returned status and a short preview of the response body. Store the key in an environment variable rather than committing it to source control. ScrapingBee API documentation
cURL
export SCRAPINGBEE_API_KEY="YOUR-API-KEY"
curl -sS -D response-headers.txt \
"https://app.scrapingbee.com/api/v1?url=https%3A%2F%2Fexample.com" \
-H "Authorization: Bearer $SCRAPINGBEE_API_KEY" \
-o response.html
head -c 500 response.html
Python
import os
import requests
api_key = os.environ["SCRAPINGBEE_API_KEY"]
response = requests.get(
"https://app.scrapingbee.com/api/v1",
params={"url": "https://example.com"},
headers={"Authorization": f"Bearer {api_key}"},
timeout=60,
)
response.raise_for_status()
print("HTTP status:", response.status_code)
print(response.text[:500])
Node.js
const apiKey = process.env.SCRAPINGBEE_API_KEY;
if (!apiKey) throw new Error("Set SCRAPINGBEE_API_KEY");
const endpoint = new URL("https://app.scrapingbee.com/api/v1");
endpoint.searchParams.set("url", "https://example.com");
const response = await fetch(endpoint, {
headers: { Authorization: `Bearer ${apiKey}` },
signal: AbortSignal.timeout(60_000),
});
const body = await response.text();
if (!response.ok) {
throw new Error(`Scraping request failed: ${response.status} ${body.slice(0, 300)}`);
}
console.log(body.slice(0, 500));
Inspect both the HTTP status and the content. A technically successful API response can still contain an unexpected page, an interstitial, incomplete content, or a layout that no longer matches your parser. Validate required fields and record enough response metadata to diagnose changes.
4. How to choose a web scraping API
Start from a sample of the pages and fields you actually need, then compare providers against these criteria. A short proof of concept should cover representative page types, including the hardest pages, rather than relying on a feature checklist alone.
| Decision area | Questions to answer | How to validate |
|---|---|---|
| Target compatibility | Can it handle the intended sites, page types, and geography for this workload? | Request a representative sample and verify the returned content and required fields. |
| Rendering and interaction | Is static HTML enough, or does the needed content appear after JavaScript? Do you need waits, scrolling, clicks, or form interaction? | Compare a non-rendered request with a rendered request; test the specific action and wait condition. |
| Access and geography | Do you need proxy selection, rotation, or geographic targeting? | Confirm supported locations and options in current documentation, then test from the location relevant to the use case. |
| Extraction and output | Do you want raw HTML for your own parser, or provider-side extraction into JSON, CSV, or another structured format? | Check field rules, missing-value behavior, and output shape against known examples. |
| Cost model | What counts as a request or credit? Do rendering, proxy modes, or extraction change the charge? | Estimate cost using your feature mix and expected volume; verify live pricing and plan limits before committing. |
| Throughput and integration | What are the rate or concurrency limits? How are credentials supplied? What errors and response codes are documented? | Read the API reference, then test the intended concurrency and error handling. |
| Operational fit | Who owns parsing, retries, scheduling, monitoring, and maintenance? | Map each operational task to a team or service and estimate the ongoing work. |
Run a representative proof of concept
- List target page types and the fields that must be present.
- Choose samples that include static content, script-rendered content, and any required interaction or geographic variation.
- Request each sample with the smallest feature set that could work, then add rendering, proxy, or interaction features only where needed.
- Compare returned output with the expected content and record missing fields, errors, latency, and feature settings. Do not generalize the result to sites or conditions you did not test.
- Estimate cost using the provider’s current request or credit rules and your expected workload.
- Decide who will own schema validation, retries, monitoring, and parser updates after launch.
5. Provider examples (not a ranking)
These examples illustrate different documented feature mixes. They are not controlled head-to-head results, and this guide does not claim that any provider is universally best.
- ScrapingBee: its documentation describes JavaScript rendering, browser scenarios, proxy modes, extraction rules, and credit costs that vary with request features. That makes it a useful example when optional rendering and access modes affect the cost model. Documentation · Pricing
- Zyte: its product material describes proxy selection and rotation, structured extraction, and usage-based pricing; its API reference documents a synchronous single-URL extraction operation. This is an example of an integrated extraction API with usage-based billing. Product overview · API documentation
- Bright Data: its product page lists rendering, residential proxies, geotargeting, automated proxy management, JSON/CSV parsing, and API or webhook delivery. This is an example of a feature set that includes geographic options and structured delivery. Product overview
Feature availability, pricing, concurrency, and plan limits can change. Recheck each provider’s current documentation and pricing for the exact product you intend to use.
6. Extraction and configuration choices
Raw HTML or structured fields
Raw HTML gives your application control over parsing and schema changes, but it also means you own the parser. Provider-side extraction can reduce parsing work, but verify selector or extraction-rule behavior, output types, and what happens when a field is absent. In either approach, validate the result against your own requirements before saving it.
JavaScript rendering and waits
Use browser rendering when required content is produced only after scripts run. Some pages also require a specific wait condition or action. ScrapingBee documents options for rendering, additional wait time, waiting for a selector, and JavaScript scenarios such as browser actions. Do not add a long fixed delay by default: wait for a condition that corresponds to the content you need where possible, and check the current API reference for accepted values. Rendering and browser scenario options
Proxy and geography options
Proxy selection, rotation, and geotargeting can matter when the workload has geographic requirements or different access conditions. They do not guarantee that every target will be accessible. Test with the target and location that matter, and account for any feature-specific charge.
URL encoding and secrets
Encode the target URL as a query parameter; use a URL or HTTP library that handles query encoding rather than concatenating arbitrary URLs into a query string. Keep API credentials out of public client-side code, logs, and committed files. The ScrapingBee documentation recommends Bearer authentication and marks the API key query parameter as deprecated. Authentication and URL parameter documentation
7. Reliability, performance, and cost
Reliability
- Check response status, content type, and required fields, not just whether the API call completed.
- Use bounded timeouts and retry only transient failures. Avoid retrying permanent configuration or authentication errors unchanged.
- Make downstream writes idempotent where duplicate requests could create duplicate records.
- Log request identifiers and safe diagnostic metadata; do not log API keys or sensitive page data unnecessarily.
- Monitor missing fields and schema changes, since a page redesign can break extraction without making the API request itself fail.
Performance
Browser rendering and interactions add work compared with fetching static HTML, so enable them only when the content requires them. Measure the end-to-end workflow on representative targets, including the time for your own parsing and validation. The reviewed vendor material does not establish comparable independent speed or success-rate benchmarks, so treat any performance conclusion as workload-specific.
Cost
Do not compare headline monthly prices without comparing included work. Depending on the provider, rendering, proxy mode, and other options can change request or credit costs. Estimate the actual combination of URL count and features, include retries and expected failures where charged, and check current pricing and plan limits before launch. Pricing and credit schedules are volatile.
8. Common problems and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| Authentication error | Missing, invalid, or incorrectly supplied API key. | Check the credential and the provider’s required authentication method. Keep the key in a server-side environment variable. |
| Request fails for URLs with query strings | The target URL was concatenated into the API query without encoding. | Pass parameters through an HTTP library’s query encoder or correctly percent-encode the target URL. |
| Content is missing from returned HTML | The content may be injected after JavaScript runs, or the response may arrive before the relevant content is ready. | Test browser rendering and a suitable selector or action wait. Confirm the content exists on the rendered page. |
| Parser returns empty fields | Selectors may no longer match, content may be absent, or the response may be an unexpected page. | Save a safe diagnostic sample, inspect the returned structure, validate selectors, and handle missing values explicitly. |
| Requests time out | The target, rendering, or wait condition may take longer than the client timeout. | Set a bounded timeout appropriate to the workload, simplify unnecessary browser actions, and retry only transient cases with limits. |
| Costs exceed the estimate | Feature options may multiply credits or usage, or retries and volume may be higher than modeled. | Inspect the current pricing schedule and request configuration; estimate cost using the actual feature mix and observed volume. |
| Results differ by location | The page may vary by geography or the request may use a different access route. | Confirm the required location and test it explicitly; do not assume a generic request reproduces a location-specific page. |
9. When the deliverable is a screenshot
A scraping API is appropriate when the output you need is page content or extracted fields. If the deliverable is a visual capture, browser automation or a screenshot API is a more direct fit. For a local do-it-yourself capture, a browser automation library such as Playwright can navigate to the URL, wait for page readiness, and save a screenshot; choose full-page or viewport capture according to what you need to inspect. For a hosted workflow, [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server for developers. It returns PNG, JPEG, WebP, or PDF from one GET request, and supports capture controls such as full-page screenshots, element selection, device presets, waits, custom CSS and JavaScript, and PDF options. ScreenshotNeo API documentation
DIY example with Playwright (Node.js)
Install Playwright and its Chromium browser using the Playwright installation instructions for your environment. This runnable script captures a full-page PNG after navigation reaches the load event. For sites that continue fetching content after load, replace the readiness condition with a selector or application-specific condition.
// Save as capture.mjs; run with: node capture.mjs https://example.com
import { chromium } from "playwright";
const target = process.argv[2];
if (!target) throw new Error("Usage: node capture.mjs https://example.com");
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(target, { waitUntil: "load", timeout: 60_000 });
await page.screenshot({ path: "page.png", fullPage: true });
} finally {
await browser.close();
}
This method means you manage the browser installation, runtime, navigation failures, waiting strategy, and output storage. If content appears only after a user interaction, add that action and wait for the resulting page state; do not assume that the load event means every page-specific task has completed.
10. Or skip the browser setup
ScreenshotNeo lets you capture a page with one API request. See the API documentation for parameters and options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners are accepted like a visitor and removed, along with known newsletter popups and chat widgets, before capture; each step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 shots a month with no card. Paid plans start at $5 for 3,000 shots; all features are on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
11. FAQ
Does a scraping API guarantee access to every website?
No. Compatibility depends on the target and request conditions. Test the actual pages and geographies you need.
Can I use a scraping API without JavaScript rendering?
Yes, when the required content is present in the server-returned HTML. Use rendering only when scripts or browser actions are needed to produce the content.
Is a screenshot API the same as a web scraping API?
No. A screenshot API returns a visual image or PDF; a scraping API returns page content or extracted data. Pick based on the output your workflow needs.
How many URLs should I include in a proof of concept?
There is no universal number. Include enough representative page types to exercise the rendering, access, and extraction requirements that could affect your production workload.


