ScreenshotNeo

BlogGuides

Web Scraping APIs for Search, Mapping, and Crawling

Choose the right API for search results, page extraction, site maps, and full crawls, with implementation patterns, costs, and troubleshooting.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: choose a SERP API when you need ranked search results, a scraping API when you need content from known URLs, a map API when you need a site’s URL structure, and a crawl API when you need to follow links across a domain. Place-search APIs are for local businesses and geographic entities. Many platforms combine several of these operations, but the workload, output, rendering, access controls, and billing model still differ.

Start with a representative set of target pages and define success as correct, complete structured data rather than an HTTP 200 response. Measure JavaScript coverage, field accuracy, success rate, latency, geographic consistency, retry behavior, and effective cost before committing to a provider.

1. Search, scrape, map, crawl, and place APIs compared

API type Input Typical output Use it when
SERP or search API Query, engine, location or language Ranked results, snippets, metadata, links, pagination You need search-engine results rather than one known page
Scraping API One or more URLs HTML, rendered HTML, Markdown, or extracted fields You know which pages to fetch
Map API Domain or starting URL Discovered URLs and site structure You need an inventory before selecting pages to scrape
Crawl API Domain, seed URL, limits and filters Content from pages reached by following links You need broad, multi-page collection
Place-search API Place query and geographic constraints Businesses, addresses, coordinates and related metadata Your records represent local businesses or geographic entities

Search is ranked retrieval. Scraping is page acquisition. Mapping discovers URLs without necessarily downloading every page. Crawling combines discovery and acquisition. Keeping those jobs separate makes cost and failure analysis easier.

2. A practical selection checklist

  1. Define the unit of work. Is billing per request, result, page, dataset, or credit?
  2. Choose the output. Raw HTML, rendered content, Markdown, structured JSON, screenshots, or downloadable CSV/JSON datasets require different products.
  3. Check rendering needs. Plain HTTP is faster and cheaper for server-rendered pages. Browser execution is needed when JavaScript creates the content you need.
  4. Check access controls. Verify proxy pools, sessions, geographic targeting, custom headers, cookies, and anti-bot handling.
  5. Design operations. Confirm synchronous and asynchronous endpoints, retries, concurrency, pagination, webhooks, and result storage.
  6. Measure a corpus. Use representative pages, not a single easy URL. Record field accuracy, missing pages, latency, retries, and effective cost.
  7. Recheck limits and prices. Vendor plans change; validate current limits immediately before purchase.

3. Search and SERP APIs

Use a SERP API when the query itself is the input. A useful response normally includes the query, organic results, result links, snippets, ranking position, and pagination metadata. Location and language controls matter when rankings vary by market.

Brave documents a Search API backed by an independent web index and also offers a Place Search API positioned as an alternative to Google Maps. You.com documents Search and Answer APIs with structured metadata and page content; its documentation lists $0.005 per Search API call and $5 per 1,000 Answer API calls. WebScraping.AI documents structured search responses with query, organic-result, pagination, and result-link fields. WebScrapingAPI documents a DuckDuckGo Search API alongside page scraping and browser-backed workflows.

Search implementation pattern

  1. Normalize the query, locale, language, and device parameters.
  2. Request one page of results with an explicit page size.
  3. Persist the provider request ID, query, timestamp, and location settings.
  4. Validate that result links and ranking positions are present before storing records.
  5. Paginate only when the use case needs more results; each additional page increases cost and latency.
import os
import requests

SEARCH_ENDPOINT = os.environ["SEARCH_ENDPOINT"]
SEARCH_API_KEY = os.environ["SEARCH_API_KEY"]

params = {
    "q": "developer documentation",
    "language": "en",
    "country": "US",
    "page": 1,
}
response = requests.get(
    SEARCH_ENDPOINT,
    params=params,
    headers={"Authorization": f"Bearer {SEARCH_API_KEY}"},
    timeout=30,
)
response.raise_for_status()
data = response.json()
for result in data.get("organic", data.get("results", [])):
    print(result.get("position"), result.get("title"), result.get("link"))

The endpoint and response field names are provider-specific. Keep the provider adapter isolated so changing vendors does not rewrite your application.

4. Direct scraping APIs

A direct scraping API fetches a URL and returns raw or rendered content, extracted fields, or a normalized document. Browser-backed rendering is useful for JavaScript-heavy pages, but it adds execution time and may change the cost model. Proxy choice, sessions, geographic routing, and anti-bot behavior can determine whether a page succeeds.

WebScrapingAPI documents page scraping and browser-backed workflows, plus marketplace endpoints for Amazon, eBay, and Walmart. Scrapy.io documents API-key HTTP calls, marketplace scrapers, an official Python SDK, synchronous and asynchronous endpoints, and JSON or CSV dataset downloads. Its FAQ exposes a pricePerResult concept, so model costs around the number of accepted records rather than requests alone.

import os
import requests

SCRAPE_ENDPOINT = os.environ["SCRAPE_ENDPOINT"]
API_KEY = os.environ["SCRAPE_API_KEY"]
url = "https://example.com/article"

response = requests.get(
    SCRAPE_ENDPOINT,
    params={"url": url, "render_js": "true"},
    headers={"Authorization": f"Bearer {API_KEY}"},
    timeout=90,
)
response.raise_for_status()
open("page.html", "wb").write(response.content)
print("saved", len(response.content), "bytes")

5. Mapping a website before crawling

A map operation discovers or organizes URLs. It is useful for audits, documentation indexing, migration planning, and selecting a bounded crawl set. A map can be cheaper than downloading every page because discovery and content extraction remain separate decisions.

DIY URL mapper in Python

from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
import requests
from bs4 import BeautifulSoup

start = "https://example.com/"
origin = urlparse(start).netloc
queue = deque([start])
seen = {start}

while queue:
    current = queue.popleft()
    try:
        response = requests.get(current, timeout=20, headers={"User-Agent": "site-mapper/1.0"})
        response.raise_for_status()
    except requests.RequestException as exc:
        print("ERROR", current, exc)
        continue

    print(current)
    content_type = response.headers.get("content-type", "")
    if "text/html" not in content_type:
        continue

    soup = BeautifulSoup(response.text, "html.parser")
    for anchor in soup.select("a[href]"):
        absolute, _ = urldefrag(urljoin(current, anchor["href"]))
        parsed = urlparse(absolute)
        if parsed.scheme in {"http", "https"} and parsed.netloc == origin and absolute not in seen:
            seen.add(absolute)
            queue.append(absolute)

This simple mapper does not implement robots.txt policy, canonical-link handling, sitemap parsing, authentication, JavaScript navigation, rate limiting, or URL budgets. Add those controls before using it against a production site.

6. Crawling an entire domain

A crawl follows links and usually combines URL discovery, fetching, rendering, extraction, retries, and storage. Define boundaries before starting:

  • Allowed hostnames and URL prefixes
  • Maximum pages and depth
  • Allowed content types and file-size limits
  • Concurrency and delay between requests
  • Query-parameter normalization and duplicate detection
  • Retry policy for timeouts, rate limits, and server errors
  • Checkpointing so a failed run can resume
  • Output schema and durable storage

Wayfern documents a 5,000-page crawl limit and concurrency capped at 5. Firecrawl documents separate Scrape, Crawl, Map, and Monitor operations: Scrape, Crawl, Map, and Monitor each cost 1 credit per page, while Search costs 2 credits per 10 results. Treat these figures as point-in-time documentation and verify current plans before purchase.

Reliable crawl sequence

  1. Run a map or parse the sitemap to estimate scope.
  2. Deduplicate and classify URLs.
  3. Fetch a small sample with the intended rendering and proxy settings.
  4. Validate extracted fields, not just HTTP status.
  5. Start the crawl with bounded concurrency and checkpoints.
  6. Retry transient failures with exponential backoff and a maximum attempt count.
  7. Write results incrementally and retain failure reasons.
  8. Run a completeness report: discovered, fetched, successful, rejected, retried, and permanently failed URLs.

7. Rendering, proxies, and anti-bot behavior

Requirement Configuration to look for Trade-off
Server-rendered HTML Plain HTTP fetch Lower latency and cost
JavaScript-generated content Browser rendering or JavaScript execution Higher latency and often higher cost
Country-specific output Geographic proxy or location parameter More consistent regional results, added cost
Login-protected pages Sessions, cookies, headers, or authentication More state to secure and maintain
Rate-limited targets Concurrency controls, delays, retries Lower pressure, longer completion time

Do not classify a response as successful solely because it is HTTP 200. Bot checks, consent pages, empty shells, login redirects, and error templates can all return successful transport status. Validate page identity and required fields.

8. cURL, Python, and Node.js request patterns

curl -G "$SCRAPE_ENDPOINT" \
  -H "Authorization: Bearer $SCRAPE_API_KEY" \
  --data-urlencode "url=https://example.com/article" \
  --data-urlencode "render_js=true" \
  -o page.html
import os
import requests

r = requests.get(
    os.environ["SCRAPE_ENDPOINT"],
    params={"url": "https://example.com/article", "render_js": "true"},
    headers={"Authorization": f"Bearer {os.environ['SCRAPE_API_KEY']}"},
    timeout=90,
)
r.raise_for_status()
open("page.html", "wb").write(r.content)
const endpoint = process.env.SCRAPE_ENDPOINT;
const key = process.env.SCRAPE_API_KEY;
const q = new URLSearchParams({
  url: 'https://example.com/article',
  render_js: 'true'
});
const res = await fetch(`${endpoint}?${q}`, {
  headers: { Authorization: `Bearer ${key}` }
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.html', body));

9. Or skip the browser setup

If your deliverable is a visual capture rather than extracted HTML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

10. Performance, reliability, and cost

  • Concurrency: increase workers only after measuring target-site throttling and provider limits.
  • Timeouts: use separate connection and overall deadlines; browser renders need longer limits than plain HTTP.
  • Retries: retry timeouts, connection resets, and rate-limit responses with backoff. Do not blindly retry permanent authorization or validation errors.
  • Caching: cache stable pages and search responses where allowed. Record the cache key and freshness policy.
  • Pagination: stop when the business requirement is met; result pages can dominate search costs.
  • Rendering: enable JavaScript only for targets that require it.
  • Storage: stream large responses and write checkpoints so a process restart does not discard completed work.
  • Cost model: calculate request, result, page, proxy, rendering, storage, and retry costs together.

11. Troubleshooting

Symptom Likely cause Fix
HTTP 200 but empty content JavaScript shell, consent wall, or bot page Enable rendering, inspect the returned body, and validate required fields
Many 403 or 429 responses Rate limiting, blocked IP, or missing headers Lower concurrency, add backoff, use supported proxy or session settings, and send an appropriate user agent
Results differ by run Location, language, device, personalization, or rotating proxies Pin those parameters and record them with every result
Crawl never finishes Unbounded query parameters, calendars, or duplicate URLs Normalize URLs, cap depth and pages, and define allowed parameters
High bill with little useful data Unnecessary rendering, retries, pagination, or duplicate fetches Sample first, deduplicate, limit fields, and enable rendering selectively
Missing internal pages Client-side navigation or links generated after load Use browser crawling or combine sitemap and map discovery
Authentication failures Expired key, wrong header, or environment variable missing Check secret loading, authorization format, and provider account status
Parser breaks after a redesign Selectors tied to unstable markup Prefer semantic attributes, version schemas, and keep fixture pages

12. Compliance and operational safeguards

Check each target’s robots directives, terms, contracts, privacy obligations, authentication requirements, and jurisdiction-specific rules. This dossier does not provide a universal legal conclusion. Keep credentials server-side, minimize stored personal data, honor deletion requirements, and document why each URL is collected.

13. FAQ

Do I need both a map API and a crawl API?

Not always. Use a map first when you need scope and URL discovery; use a crawl when you also need page content. A combined product can perform both operations, but separating them helps control cost.

Is a SERP API the same as a scraping API?

No. A SERP API queries a search index and returns ranked results. A scraping API fetches a specified page and returns its content or extracted fields.

When should I use a place-search API?

Use one when the primary records are local businesses or geographic entities. General web search is a poorer fit for structured place data.

Should every crawl use a headless browser?

No. Start with plain HTTP and enable browser rendering only for pages whose required content is created by JavaScript.

How should I compare providers?

Run the same representative corpus through each candidate and compare field accuracy, completeness, JavaScript coverage, latency, geographic consistency, retry behavior, and effective cost.