Migrating From Crawlbase to a Web Scraping API
A practical Crawlbase migration guide: map legacy APIs, preserve rendering and proxy behavior, compare providers, and cut over safely.
Short answer: inventory the Crawlbase surface you use, map it to the modern API or a replacement with the same rendering and proxy behavior, then prove parity with a small acceptance suite before changing production traffic. Crawlbase’s current surfaces are the Crawling API for most new integrations, Smart AI Proxy for a proxy-shaped interface, and Enterprise Crawler for very large asynchronous queues.
The safest migration is contract-first. Capture your current endpoint, parameters, response format, wait actions, proxy and country rules, session behavior, retry policy, and billing unit. Recreate those behaviors with the smallest provider-specific adapter possible, then run both systems against the same URLs.
1. Identify your Crawlbase integration
Crawlbase has legacy products whose modern destinations differ. Write down these fields before selecting a provider:
- Endpoint and token type; Crawlbase documents one token authenticating its APIs.
- Target URL encoding, HTTP method, and whether requests are synchronous or queued.
- JavaScript rendering, wait-for-selector, delay, scroll, click, and AJAX-idle behavior.
- Proxy type (residential or datacenter), country targeting, sticky sessions, and CAPTCHA or bot handling.
- Output contract: HTML, Markdown, JSON, screenshot, PDF, extracted fields, or callback payload.
- Timeouts, retry and backoff rules, concurrency limits, and how a billable request is counted.
| Legacy surface | Modern Crawlbase mapping | What to verify |
|---|---|---|
| Scraper API | Crawling API plus scraper= parameters |
Rendered output, extraction fields, and response metadata |
| Screenshots API | Crawling API screenshot parameters or an MCP screenshot tool | Viewport, full-page behavior, image format, and failure semantics |
| Proxy API | Smart AI Proxy | Proxy authentication, country, sticky session, and headers |
| Leads API | No direct replacement; Crawlbase describes its email-extractor scraper as the closest workflow | Validate extraction quality and legal requirements separately |
2. Build a parity checklist
Do not compare providers by endpoint names alone. Mark each item as required, optional, or removable:
| Area | Questions |
|---|---|
| Rendering | Does JavaScript run? Can you wait for a selector, delay, scroll, click, or network/AJAX idle? |
| Access | Are residential or datacenter exits available? Can you target a country and keep a sticky session? How are bot challenges handled? |
| Output | Do downstream jobs require raw HTML, Markdown (format=md), JSON, screenshots, PDF, or a callback? |
| Extraction | Are CSS/XPath selectors, structured extraction, or an AI extractor part of the contract? |
| Operations | What are timeout, retry, rate-limit, concurrency, storage, and webhook guarantees? |
| Billing | Is a request, browser render, successful response, proxy use, or data volume the billable unit? |
3. Choose a migration target
Use the workload, not a feature checklist, to narrow the field:
| Option | Best fit | Migration watch-outs |
|---|---|---|
| Crawlbase Crawling API | Stay with Crawlbase while leaving legacy endpoints. | Update endpoint and parameters; preserve token and rendering assumptions. |
| ScraperAPI | A broad URL, API, image, document, and PDF scraping API. | Verify response format, crawler behavior, and credit or concurrency limits. |
| ScrapingBee | Simple hosted calls to JavaScript-heavy pages. | Convert request parameters and account for credit multipliers for browser or AI features. |
| Zyte API | Difficult targets, automatic ban avoidance, extraction, and pay-as-you-go usage. | Change GET query calls to POST JSON and adapt RPM/concurrency assumptions. |
| Apify | Prebuilt Actors, scheduled jobs, and multi-step pipelines. | This is a workflow migration; validate data contracts and orchestration. |
These products expose different request shapes. ScrapingBee uses GET query parameters; Zyte documents POST requests with JSON bodies. Billing and ban-handling models also differ, so normalize the full cost of rendering, proxying, anti-bot handling, and extraction before choosing.
4. Create a provider-neutral request contract
Keep application code independent from vendor syntax. A small internal object is enough:
{
"url": "https://example.com/products",
"render_js": true,
"wait": {"selector": ".product-card", "timeout_ms": 15000},
"proxy": {"country": "us", "sticky": true},
"actions": [{"type": "scroll", "y": 1200}],
"output": "html",
"timeout_ms": 90000
}
Write one adapter per provider. Store the original request and normalized response metadata so a failed migration can be replayed. Never silently drop a field: reject unsupported options or record that the behavior changed.
5. Example migration with cURL
The following is a provider-neutral template. Replace the endpoint and parameter names with the target provider’s documentation, then keep the response handling unchanged in your application.
curl -G "https://api.example-provider.test/v1/fetch" \
--data-urlencode "url=https://example.com/products" \
--data "render_js=true" \
--data "wait_for=.product-card" \
--data "country=us" \
--data "format=html" \
-H "Authorization: Bearer $SCRAPER_API_KEY" \
--max-time 90 \
-o response.html
If the replacement requires POST JSON, use this shape:
curl https://api.example-provider.test/v1/fetch \
-H "Authorization: Bearer $SCRAPER_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"url": "https://example.com/products",
"browserHtml": true,
"actions": [{"action": "waitForSelector", "selector": ".product-card"}]
}'
6. Python adapter with retries
import os
import time
import requests
API_URL = "https://api.example-provider.test/v1/fetch"
KEY = os.environ["SCRAPER_API_KEY"]
def fetch_page(url: str) -> str:
params = {
"url": url,
"render_js": "true",
"wait_for": ".product-card",
"country": "us",
"format": "html",
}
for attempt in range(4):
try:
r = requests.get(API_URL, params=params, headers={"Authorization": f"Bearer {KEY}"}, timeout=(10, 90))
if r.status_code in (429, 500, 502, 503, 504):
if attempt == 3:
r.raise_for_status()
time.sleep(2 ** attempt)
continue
r.raise_for_status()
return r.text
except requests.RequestException:
if attempt == 3:
raise
time.sleep(2 ** attempt)
html = fetch_page("https://example.com/products")
open("response.html", "w", encoding="utf-8").write(html)
7. Node.js adapter
const key = process.env.SCRAPER_API_KEY;
const params = new URLSearchParams({
url: 'https://example.com/products',
render_js: 'true',
wait_for: '.product-card',
country: 'us',
format: 'html'
});
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 90_000);
try {
const res = await fetch('https://api.example-provider.test/v1/fetch?' + params, {
headers: { Authorization: 'Bearer ' + key },
signal: controller.signal
});
if (!res.ok) throw new Error('scraper status ' + res.status);
const html = await res.text();
require('node:fs').writeFileSync('response.html', html);
} finally {
clearTimeout(timer);
}
8. Validate before switching traffic
- Select representative URLs: static HTML, JavaScript-rendered content, a consent wall, a bot challenge, a slow page, and a 404.
- Replay each URL through Crawlbase and the candidate with identical country, session, wait, and timeout settings.
- Compare status, final URL, title, content markers, extracted fields, screenshot dimensions, and latency. Compare failure reasons separately from successful bodies.
- Run a shadow period where the replacement is called but its result is not published. Record provider status, retry count, billable units, and response size.
- Cut over gradually and keep a rollback switch. Remove the old integration only after queued jobs and webhooks have drained.
9. Edge cases that break migrations
- JavaScript timing: a fixed delay can finish before hydration. Prefer a selector or network-idle condition when supported.
- Infinite scroll: reproduce the exact scroll count or stop condition.
- Sticky sessions: preserve the same session identifier across pagination and login flows.
- Country-specific pages: set proxy country, locale headers, timezone, and cookies consistently.
- Consent and login walls: transfer cookies and custom headers deliberately; never log credentials.
- Large responses: stream to storage and enforce a size limit before parsing.
- Async queues: make webhook handlers idempotent and ignore duplicate deliveries.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| HTML is a shell with no data | JavaScript was disabled or the wait ended too soon. | Enable browser rendering and wait for a stable selector or AJAX idle. |
| Different country or language | Proxy location and request locale disagree. | Set country, Accept-Language, timezone, and locale cookies consistently. |
| 403, 429, or CAPTCHA | Exit reputation, request rate, or missing session state. | Use supported proxy and anti-bot options, lower concurrency, add backoff, and preserve sessions. |
| Timeouts after migration | Browser startup, proxy connection, or page idle exceeds the old timeout. | Separate connect and overall timeouts; retry transient failures only. |
| Costs rise unexpectedly | Browser, AI, proxy, or failed-request billing differs. | Normalize billable units per URL and set job budgets. |
| Parser receives the wrong format | Markdown, HTML, JSON, or compressed bytes changed. | Pin output format and content-type checks in the adapter. |
| Duplicate callback records | Webhook delivery was retried. | Use an idempotency key or provider job ID. |
11. Performance, reliability, and cost
- Performance: measure DNS/connect, proxy acquisition, browser startup, page load, wait time, transfer, and parsing separately.
- Reliability: classify errors as permanent, target-side, or transient. Retry transient classes with exponential backoff and jitter.
- Concurrency: start below the documented limit and increase while watching 429s, queue time, and target errors.
- Cost: record successful requests, browser renders, proxy use, extraction, bytes, and retries. Crawlbase notes that successful requests, normal versus JavaScript requests, and domain complexity affect billing.
- Caching: cache immutable pages and include URL, locale, session, and rendering options in the cache key.
12. Or skip the browser setup
If your migration only needs reliable screenshots or PDFs, ScreenshotNeo is a direct capture API. It accepts a URL and returns PNG, JPEG, WebP, or PDF; its parameter names are compatible with those used by other screenshot services. See the ScreenshotNeo docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing state. An MCP server lets Claude, Cursor, and other MCP clients take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
13. Migration checklist
- Inventory every Crawlbase endpoint and legacy feature.
- Freeze representative URLs and expected output fixtures.
- Choose a provider and document unsupported fields.
- Implement an adapter with explicit timeouts, retries, and idempotency.
- Shadow traffic, compare usable output and effective cost, then roll out gradually.
- Keep rollback credentials and drain asynchronous jobs before decommissioning.
FAQ
Do I have to rewrite every caller?
No. Put provider-specific syntax behind one adapter and keep the internal request contract stable.
Can I keep JavaScript rendering and proxy rotation?
Usually, but verify browser rendering, proxy type, country, sticky sessions, waits, and anti-bot behavior as separate acceptance criteria.
Is an API swap the same as moving to Apify?
No. Apify migration includes Actors, schedules, storage, and orchestration, so validate the workflow and data contracts.
When is ScreenshotNeo a better fit?
When the required artifact is a clean screenshot or PDF and you want consent and popup removal, explicit non-billing for failed captures, or MCP access for AI agents.


