Migrating From ScrapingBee to a Web Scraping API
A practical migration guide covering request mapping, response changes, feature gaps, testing, cost, reliability, and a safe production cutover.

Moving from ScrapingBee to another web scraping API is a transport migration and a behavior-parity exercise. You must change how requests are authenticated and serialized, then prove that rendering, waits, browser actions, proxy behavior, extraction, response parsing, retries, latency, throughput, and cost still meet your application’s needs.
The safest sequence is: inventory the ScrapingBee integration, map each option to the destination API, update request and response handling, compare representative pages, recalculate real costs, and cut over gradually with a rollback path.
What changes when you switch providers?
Do not assume that replacing a hostname and API key creates a drop-in replacement. ScrapingBee’s HTML API accepts a target URL and API key, recommends bearer-token authentication, and enables JavaScript rendering by default. Rendering and proxy choices affect credit usage. A destination provider may use a different HTTP method, body format, authentication scheme, response envelope, and billing model.
Zyte’s documented ScrapingBee migration is a useful concrete example. A ScrapingBee GET request with URL-encoded query parameters becomes a Zyte POST request with JSON in the body. Authentication changes to HTTP Basic authentication in Zyte’s example. The target response is returned inside a JSON object, and the body is base64 encoded. Those differences require changes to both the HTTP client and downstream parsing; they are not solved by changing a base URL. See the Zyte migration guide.
1. Inventory your current ScrapingBee behavior
Read the client code, configuration, and consumers of its output. Create a table of every option that is actually used and the assumption made by the rest of your system.
| Area | Questions to answer |
|---|---|
| Authentication | Is the key in an Authorization header or a query string? Are keys rotated from one secret store? |
| Navigation | Is JavaScript rendering enabled? Do you use wait, wait_for, or a navigation wait? |
| Actions | Does a js_scenario click, fill, scroll, or wait before extraction? |
| Access | Do requests require a country, premium or stealth proxy, custom proxy, headers, cookies, or a user agent? |
| Output | Does your code expect raw HTML, text, Markdown, a screenshot, structured extraction, AI extraction, status data, or target headers? |
| Operations | What are timeout values, retry rules, concurrency limits, cache behavior, and usage alerts? |
ScrapingBee documents these as separate controls, so an option being optional does not mean your integration does not depend on it. Record defaults as well as explicit parameters. For example, a client that never sends render_js may still rely on ScrapingBee’s documented JavaScript-rendering default. Start with the ScrapingBee API documentation.
2. Separate transport changes from scraping behavior
Make transport changes in one layer and page-behavior changes in another. This keeps a parsing bug from being confused with a browser configuration bug.
Transport checklist
- HTTP method: GET, POST, or another method.
- Authentication: bearer token, query parameter, Basic auth, or a provider-specific header.
- Parameter encoding: URL query parameters versus JSON or form data.
- Response envelope: direct page bytes versus JSON containing fields.
- Body encoding: plain text, compressed data, or base64.
- Content type and character encoding.
- Error format, status codes, request IDs, and rate-limit headers.
ScrapingBee’s current documentation recommends the Authorization bearer-token header and describes query-string api_key authentication as deprecated while still supported for backward compatibility. Do not carry an old query-string pattern into a new provider without checking its authentication model.
Behavior checklist
- Browser rendering and wait conditions.
- Clicks, form fills, scrolling, and other actions.
- Proxy type, country, geolocation, and escalation.
- Cookies, request headers, and user-agent behavior.
- Ad or resource blocking.
- Extraction rules, selectors, AI extraction, and screenshots.
3. Map features one at a time
Build a mapping document with four columns: ScrapingBee option, destination equivalent, code change, and validation case. Zyte maps common controls such as render_js to browser HTML, wait and wait_for to browser actions, premium proxy settings to a residential IP type, and country_code to geolocation. ScrapingBee actions such as click, fill, scroll, and wait also have Zyte action equivalents.
The same guide lists unsupported options, including ad or resource blocking, custom proxies, server-side extraction rules, selected screenshot targeting, and some request controls and headers. Treat each unsupported feature as a design decision: reproduce it in your own pipeline, change the workflow, or retain that capability elsewhere only after testing. Do not claim full feature parity between providers.
| ScrapingBee dependency | Migration decision |
|---|---|
| JavaScript rendering | Enable the destination’s browser mode and compare content after hydration. |
wait_for selector |
Use an equivalent selector wait; verify that the selector appears before timeout. |
js_scenario |
Translate each action and preserve ordering, pauses, and failure handling. |
| Country or premium proxy | Map to the closest location and IP type, then test geo-sensitive pages. |
| Custom headers or cookies | Confirm which headers are accepted and whether cookies must be sent in a structured field. |
| Extraction or screenshot option | Find an equivalent output or move extraction and image processing into your application. |
4. Update the client and response parser
Keep the provider call behind a small interface such as fetch_page(target, options). Return a normalized internal object containing body, status, headers, provider request ID, timing, and billing metadata where available. Decode provider-specific envelopes inside the adapter so the rest of the application continues to consume one shape.
Generic cURL migration pattern
# Old shape: ScrapingBee-style GET with query parameters
curl -G "https://app.scrapingbee.com/api/v1/" \
-H "Authorization: Bearer $SCRAPINGBEE_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode "render_js=true"
# Destination shape: example POST/JSON pattern used by Zyte's guide
curl -X POST "https://api.zyte.com/v1/extract" \
-u "$ZYTE_API_KEY:" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","browserHtml":true}'
The exact destination endpoint and field names must come from that provider’s documentation. In Zyte’s response, inspect the JSON object and decode the target body from its documented base64 field before parsing HTML.
Python adapter skeleton
import base64
import requests
def fetch_page(url: str, api_key: str) -> str:
response = requests.post(
"https://api.zyte.com/v1/extract",
auth=(api_key, ""),
json={"url": url, "browserHtml": True},
timeout=90,
)
response.raise_for_status()
payload = response.json()
encoded = payload["browserHtml"]
return base64.b64decode(encoded).decode("utf-8", errors="replace")
html = fetch_page("https://example.com", "YOUR_API_KEY")
print(len(html))
Node.js adapter skeleton
const response = await fetch('https://api.zyte.com/v1/extract', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Basic ${Buffer.from(`${process.env.ZYTE_API_KEY}:`).toString('base64')}`
},
body: JSON.stringify({ url: 'https://example.com', browserHtml: true })
});
if (!response.ok) throw new Error(`Provider error: ${response.status}`);
const payload = await response.json();
const html = Buffer.from(payload.browserHtml, 'base64').toString('utf8');
console.log(html.length);
These examples illustrate the transport and decoding changes documented by Zyte. Confirm the current endpoint, field names, and authentication details before deploying.
5. Build a representative comparison suite
Choose pages from your real workload, not only a simple static page. Include server-rendered HTML, JavaScript applications, pages requiring a selector wait, interaction flows, geo-sensitive content, cookie-dependent pages, blocked or challenged pages, large documents, and pages that frequently time out.
- Send equivalent requests to ScrapingBee and the candidate provider.
- Normalize encoding, whitespace, and dynamic timestamps before comparison.
- Compare required extracted fields, links, table rows, and content completeness.
- Record status, provider errors, timeout rate, latency percentiles, retries, and response size.
- Run the suite repeatedly at realistic concurrency.
- Test complex cases before moving production traffic, as recommended in Zyte’s migration guidance.
A successful HTTP status is not enough. A browser can return a page shell without the data your parser needs, or a challenge page with status 200. Add semantic assertions such as “product list contains at least one item” or “article body exceeds the minimum length.”
6. Recalculate cost and throughput
Compare the bill for your actual request mix. ScrapingBee documents credit costs that vary by configuration: its current documentation says JavaScript rendering is enabled by default and costs 5 credits for a standard request; premium proxy use is documented as 25 credits with JavaScript rendering and 10 without; stealth proxy use is documented as 75 credits per successful API call with JavaScript rendering; and AI extraction options add 5 credits. These are provider terms that can change, so verify them before purchasing.
ScrapingBee’s Auto-Mode can try configurations from cheaper to more expensive and charge for the configuration that succeeds, with an optional cap. Model that escalation when estimating spend. Zyte describes pay-as-you-go pricing, spending limits or commitments, and RPM-based limits, while ScrapingBee describes concurrency-based limits. A fair comparison therefore needs page mix, browser-rendering rate, proxy escalation, successful volume, concurrency, retries, and required throughput.
| Metric | Why it matters |
|---|---|
| Cost per successful page | Separates useful output from retries and failed attempts. |
| Cost per rendered page | Shows the impact of JavaScript defaults and browser mode. |
| p50/p95 latency | Reveals whether browser requests meet your job deadlines. |
| Concurrency or RPM headroom | Determines queue growth during traffic spikes. |
| Extraction success rate | Captures silent partial pages that status-code metrics miss. |
7. Cut over safely
- Ship the destination adapter behind a feature flag.
- Run shadow requests where permitted, or replay a sanitized sample.
- Compare normalized outputs and alert on missing fields.
- Start with a small percentage of traffic and keep ScrapingBee available for rollback.
- Watch errors, latency, retries, throughput, extraction quality, and spend.
- Expand only after the workload meets its acceptance thresholds.
A staged rollout and rollback plan are engineering recommendations inferred from the documented request, response, feature, and billing differences. No provider comparison can guarantee that a migration succeeds without workload-specific validation.
Or skip the browser setup
If your migration need is specifically clean screenshots or PDFs rather than HTML extraction, ScreenshotNeo provides a focused website capture API. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.


See the ScreenshotNeo API documentation for all options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, selector or network-idle waits, blocking controls, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account.
Troubleshooting common migration failures
401 or 403 authentication errors
Cause: the destination expects a different header or Basic auth format, or the key is scoped differently. Fix: reproduce the provider’s documented authentication example exactly, remove deprecated query-string keys, and verify the key in the correct environment.
400 invalid request
Cause: query parameters were sent where JSON was required, a boolean became a string, or a field name has no destination equivalent. Fix: log the serialized request with secrets redacted and validate it against the destination schema.
Parser receives JSON instead of HTML
Cause: the destination wraps the target body in an envelope or base64 encodes it, as in Zyte’s documented migration. Fix: parse the envelope, decode the body, then pass text to the existing HTML parser.
JavaScript content is missing
Cause: browser rendering is disabled, the wait condition is wrong, or an action was not translated. Fix: enable browser mode, wait for a stable selector or navigation event, and validate each action independently.
Pages are geo-inconsistent
Cause: country, proxy type, timezone, or cookies differ between providers. Fix: map all location inputs and compare from the same intended region.
Unexpected cost increase
Cause: rendering or premium proxy escalation is occurring more often, retries are counted, or extraction options add charges. Fix: tag requests by configuration, measure cost per successful output, set provider spending controls, and compare the same workload mix.
Rate-limit or queue failures
Cause: the destination uses RPM limits while the old integration was tuned for concurrency, or vice versa. Fix: add a bounded queue, exponential backoff for retryable responses, and provider-specific concurrency or RPM settings.
Performance and reliability checklist
- Reuse HTTP connections and set explicit connect and total timeouts.
- Cap concurrency below the provider limit and monitor queue depth.
- Retry only transient network, timeout, and rate-limit failures; do not blindly retry invalid requests.
- Use idempotency keys or deduplication for jobs that may be retried.
- Cache immutable pages where your terms and freshness requirements allow it.
- Store provider request IDs with every result for support and incident analysis.
- Track semantic extraction failures separately from transport failures.
- Keep the old adapter available until production traffic confirms parity.
FAQ
Is Zyte API a drop-in replacement for ScrapingBee?
No. Zyte documents a direct migration path, but the HTTP method, authentication, request encoding, response envelope, base64 body, and supported options differ. Map and test the features your application uses.
Should I migrate the parser or the request client first?
Build the destination adapter and normalized response object first. Then point existing extraction code at that normalized object. This isolates transport changes from parsing changes.
How many pages should be in a migration test set?
Use enough samples to cover every behavior your production traffic relies on: rendering, waits, actions, geo, cookies, large pages, failures, and extraction edge cases. A fixed page count is less useful than behavioral coverage.
Can I compare providers using headline prices?
No. Rendering, proxy escalation, extraction charges, retries, rate limits, and successful-output rates make the actual workload mix decisive.
When does ScreenshotNeo fit this migration?
Use it when the output you need is a clean screenshot or PDF. It is a capture API with consent and widget removal, verdict-based billing, and an MCP server for AI agents, rather than a general HTML extraction replacement.