Web Scraping APIs for Search, Mapping, and Crawling
Choose the right API for search results, page extraction, site maps, and full crawls, with implementation patterns, costs, and troubleshooting.
Direct answer: choose a SERP API when you need ranked search results, a scraping API when you need content from known URLs, a map API when you need a site’s URL structure, and a crawl API when you need to follow links across a domain. Place-search APIs are for local businesses and geographic entities. Many platforms combine several of these operations, but the workload, output, rendering, access controls, and billing model still differ.
Start with a representative set of target pages and define success as correct, complete structured data rather than an HTTP 200 response. Measure JavaScript coverage, field accuracy, success rate, latency, geographic consistency, retry behavior, and effective cost before committing to a provider.
1. Search, scrape, map, crawl, and place APIs compared
| API type | Input | Typical output | Use it when |
|---|---|---|---|
| SERP or search API | Query, engine, location or language | Ranked results, snippets, metadata, links, pagination | You need search-engine results rather than one known page |
| Scraping API | One or more URLs | HTML, rendered HTML, Markdown, or extracted fields | You know which pages to fetch |
| Map API | Domain or starting URL | Discovered URLs and site structure | You need an inventory before selecting pages to scrape |
| Crawl API | Domain, seed URL, limits and filters | Content from pages reached by following links | You need broad, multi-page collection |
| Place-search API | Place query and geographic constraints | Businesses, addresses, coordinates and related metadata | Your records represent local businesses or geographic entities |
Search is ranked retrieval. Scraping is page acquisition. Mapping discovers URLs without necessarily downloading every page. Crawling combines discovery and acquisition. Keeping those jobs separate makes cost and failure analysis easier.
2. A practical selection checklist
- Define the unit of work. Is billing per request, result, page, dataset, or credit?
- Choose the output. Raw HTML, rendered content, Markdown, structured JSON, screenshots, or downloadable CSV/JSON datasets require different products.
- Check rendering needs. Plain HTTP is faster and cheaper for server-rendered pages. Browser execution is needed when JavaScript creates the content you need.
- Check access controls. Verify proxy pools, sessions, geographic targeting, custom headers, cookies, and anti-bot handling.
- Design operations. Confirm synchronous and asynchronous endpoints, retries, concurrency, pagination, webhooks, and result storage.
- Measure a corpus. Use representative pages, not a single easy URL. Record field accuracy, missing pages, latency, retries, and effective cost.
- Recheck limits and prices. Vendor plans change; validate current limits immediately before purchase.
3. Search and SERP APIs
Use a SERP API when the query itself is the input. A useful response normally includes the query, organic results, result links, snippets, ranking position, and pagination metadata. Location and language controls matter when rankings vary by market.
Brave documents a Search API backed by an independent web index and also offers a Place Search API positioned as an alternative to Google Maps. You.com documents Search and Answer APIs with structured metadata and page content; its documentation lists $0.005 per Search API call and $5 per 1,000 Answer API calls. WebScraping.AI documents structured search responses with query, organic-result, pagination, and result-link fields. WebScrapingAPI documents a DuckDuckGo Search API alongside page scraping and browser-backed workflows.
Search implementation pattern
- Normalize the query, locale, language, and device parameters.
- Request one page of results with an explicit page size.
- Persist the provider request ID, query, timestamp, and location settings.
- Validate that result links and ranking positions are present before storing records.
- Paginate only when the use case needs more results; each additional page increases cost and latency.
import os
import requests
SEARCH_ENDPOINT = os.environ["SEARCH_ENDPOINT"]
SEARCH_API_KEY = os.environ["SEARCH_API_KEY"]
params = {
"q": "developer documentation",
"language": "en",
"country": "US",
"page": 1,
}
response = requests.get(
SEARCH_ENDPOINT,
params=params,
headers={"Authorization": f"Bearer {SEARCH_API_KEY}"},
timeout=30,
)
response.raise_for_status()
data = response.json()
for result in data.get("organic", data.get("results", [])):
print(result.get("position"), result.get("title"), result.get("link"))
The endpoint and response field names are provider-specific. Keep the provider adapter isolated so changing vendors does not rewrite your application.
4. Direct scraping APIs
A direct scraping API fetches a URL and returns raw or rendered content, extracted fields, or a normalized document. Browser-backed rendering is useful for JavaScript-heavy pages, but it adds execution time and may change the cost model. Proxy choice, sessions, geographic routing, and anti-bot behavior can determine whether a page succeeds.
WebScrapingAPI documents page scraping and browser-backed workflows, plus marketplace endpoints for Amazon, eBay, and Walmart. Scrapy.io documents API-key HTTP calls, marketplace scrapers, an official Python SDK, synchronous and asynchronous endpoints, and JSON or CSV dataset downloads. Its FAQ exposes a pricePerResult concept, so model costs around the number of accepted records rather than requests alone.
import os
import requests
SCRAPE_ENDPOINT = os.environ["SCRAPE_ENDPOINT"]
API_KEY = os.environ["SCRAPE_API_KEY"]
url = "https://example.com/article"
response = requests.get(
SCRAPE_ENDPOINT,
params={"url": url, "render_js": "true"},
headers={"Authorization": f"Bearer {API_KEY}"},
timeout=90,
)
response.raise_for_status()
open("page.html", "wb").write(response.content)
print("saved", len(response.content), "bytes")
5. Mapping a website before crawling
A map operation discovers or organizes URLs. It is useful for audits, documentation indexing, migration planning, and selecting a bounded crawl set. A map can be cheaper than downloading every page because discovery and content extraction remain separate decisions.
DIY URL mapper in Python
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
import requests
from bs4 import BeautifulSoup
start = "https://example.com/"
origin = urlparse(start).netloc
queue = deque([start])
seen = {start}
while queue:
current = queue.popleft()
try:
response = requests.get(current, timeout=20, headers={"User-Agent": "site-mapper/1.0"})
response.raise_for_status()
except requests.RequestException as exc:
print("ERROR", current, exc)
continue
print(current)
content_type = response.headers.get("content-type", "")
if "text/html" not in content_type:
continue
soup = BeautifulSoup(response.text, "html.parser")
for anchor in soup.select("a[href]"):
absolute, _ = urldefrag(urljoin(current, anchor["href"]))
parsed = urlparse(absolute)
if parsed.scheme in {"http", "https"} and parsed.netloc == origin and absolute not in seen:
seen.add(absolute)
queue.append(absolute)
This simple mapper does not implement robots.txt policy, canonical-link handling, sitemap parsing, authentication, JavaScript navigation, rate limiting, or URL budgets. Add those controls before using it against a production site.
6. Crawling an entire domain
A crawl follows links and usually combines URL discovery, fetching, rendering, extraction, retries, and storage. Define boundaries before starting:
- Allowed hostnames and URL prefixes
- Maximum pages and depth
- Allowed content types and file-size limits
- Concurrency and delay between requests
- Query-parameter normalization and duplicate detection
- Retry policy for timeouts, rate limits, and server errors
- Checkpointing so a failed run can resume
- Output schema and durable storage
Wayfern documents a 5,000-page crawl limit and concurrency capped at 5. Firecrawl documents separate Scrape, Crawl, Map, and Monitor operations: Scrape, Crawl, Map, and Monitor each cost 1 credit per page, while Search costs 2 credits per 10 results. Treat these figures as point-in-time documentation and verify current plans before purchase.
Reliable crawl sequence
- Run a map or parse the sitemap to estimate scope.
- Deduplicate and classify URLs.
- Fetch a small sample with the intended rendering and proxy settings.
- Validate extracted fields, not just HTTP status.
- Start the crawl with bounded concurrency and checkpoints.
- Retry transient failures with exponential backoff and a maximum attempt count.
- Write results incrementally and retain failure reasons.
- Run a completeness report: discovered, fetched, successful, rejected, retried, and permanently failed URLs.
7. Rendering, proxies, and anti-bot behavior
| Requirement | Configuration to look for | Trade-off |
|---|---|---|
| Server-rendered HTML | Plain HTTP fetch | Lower latency and cost |
| JavaScript-generated content | Browser rendering or JavaScript execution | Higher latency and often higher cost |
| Country-specific output | Geographic proxy or location parameter | More consistent regional results, added cost |
| Login-protected pages | Sessions, cookies, headers, or authentication | More state to secure and maintain |
| Rate-limited targets | Concurrency controls, delays, retries | Lower pressure, longer completion time |
Do not classify a response as successful solely because it is HTTP 200. Bot checks, consent pages, empty shells, login redirects, and error templates can all return successful transport status. Validate page identity and required fields.
8. cURL, Python, and Node.js request patterns
curl -G "$SCRAPE_ENDPOINT" \
-H "Authorization: Bearer $SCRAPE_API_KEY" \
--data-urlencode "url=https://example.com/article" \
--data-urlencode "render_js=true" \
-o page.html
import os
import requests
r = requests.get(
os.environ["SCRAPE_ENDPOINT"],
params={"url": "https://example.com/article", "render_js": "true"},
headers={"Authorization": f"Bearer {os.environ['SCRAPE_API_KEY']}"},
timeout=90,
)
r.raise_for_status()
open("page.html", "wb").write(r.content)
const endpoint = process.env.SCRAPE_ENDPOINT;
const key = process.env.SCRAPE_API_KEY;
const q = new URLSearchParams({
url: 'https://example.com/article',
render_js: 'true'
});
const res = await fetch(`${endpoint}?${q}`, {
headers: { Authorization: `Bearer ${key}` }
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.html', body));
9. Or skip the browser setup
If your deliverable is a visual capture rather than extracted HTML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
10. Performance, reliability, and cost
- Concurrency: increase workers only after measuring target-site throttling and provider limits.
- Timeouts: use separate connection and overall deadlines; browser renders need longer limits than plain HTTP.
- Retries: retry timeouts, connection resets, and rate-limit responses with backoff. Do not blindly retry permanent authorization or validation errors.
- Caching: cache stable pages and search responses where allowed. Record the cache key and freshness policy.
- Pagination: stop when the business requirement is met; result pages can dominate search costs.
- Rendering: enable JavaScript only for targets that require it.
- Storage: stream large responses and write checkpoints so a process restart does not discard completed work.
- Cost model: calculate request, result, page, proxy, rendering, storage, and retry costs together.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 200 but empty content | JavaScript shell, consent wall, or bot page | Enable rendering, inspect the returned body, and validate required fields |
| Many 403 or 429 responses | Rate limiting, blocked IP, or missing headers | Lower concurrency, add backoff, use supported proxy or session settings, and send an appropriate user agent |
| Results differ by run | Location, language, device, personalization, or rotating proxies | Pin those parameters and record them with every result |
| Crawl never finishes | Unbounded query parameters, calendars, or duplicate URLs | Normalize URLs, cap depth and pages, and define allowed parameters |
| High bill with little useful data | Unnecessary rendering, retries, pagination, or duplicate fetches | Sample first, deduplicate, limit fields, and enable rendering selectively |
| Missing internal pages | Client-side navigation or links generated after load | Use browser crawling or combine sitemap and map discovery |
| Authentication failures | Expired key, wrong header, or environment variable missing | Check secret loading, authorization format, and provider account status |
| Parser breaks after a redesign | Selectors tied to unstable markup | Prefer semantic attributes, version schemas, and keep fixture pages |
12. Compliance and operational safeguards
Check each target’s robots directives, terms, contracts, privacy obligations, authentication requirements, and jurisdiction-specific rules. This dossier does not provide a universal legal conclusion. Keep credentials server-side, minimize stored personal data, honor deletion requirements, and document why each URL is collected.
13. FAQ
Do I need both a map API and a crawl API?
Not always. Use a map first when you need scope and URL discovery; use a crawl when you also need page content. A combined product can perform both operations, but separating them helps control cost.
Is a SERP API the same as a scraping API?
No. A SERP API queries a search index and returns ranked results. A scraping API fetches a specified page and returns its content or extracted fields.
When should I use a place-search API?
Use one when the primary records are local businesses or geographic entities. General web search is a poorer fit for structured place data.
Should every crawl use a headless browser?
No. Start with plain HTTP and enable browser rendering only for pages whose required content is created by JavaScript.
How should I compare providers?
Run the same representative corpus through each candidate and compare field accuracy, completeness, JavaScript coverage, latency, geographic consistency, retry behavior, and effective cost.


