Migrating From Scrape.do to a Web Scraping API
A provider-neutral migration plan for replacing Scrape.do: inventory controls, map behavior, validate results, rebaseline cost, and cut over safely.
Replacing Scrape.do is an API contract migration, not a search-and-replace exercise. First inventory how your integration authenticates, encodes target URLs, selects proxies, preserves sessions, forwards headers, renders JavaScript, waits for content, retries failures, reports cost, and handles asynchronous jobs. Then map each required behavior to a documented capability at the destination provider, run both providers against representative pages, and shift traffic gradually with rollback ready.
What do I need to change when switching scraping API providers?
Change the provider-specific contract while preserving the page behavior your application actually needs. The destination may use a different endpoint, authentication header, request body, URL encoding rule, proxy model, rendering switch, wait syntax, error format, concurrency limit, webhook contract, and billing unit.
| Contract area | Record in the Scrape.do integration | Map and verify at the destination |
|---|---|---|
| Target | Absolute URL, HTTP method, body, URL encoding | Accepted protocols, encoding, methods, redirects |
| Authentication | Account token location and secret loading | Query parameter, header, signed request, scopes |
| Network route | Datacenter, residential/mobile, geography | Equivalent proxy classes, countries, fallback behavior |
| Session | Sticky session or session identifier | Cookie jar, session lifetime, reuse semantics |
| Browser behavior | JavaScript rendering, waits, screenshots or HTML | Rendering engine, selector/network-idle waits, timeout limits |
| Request shaping | Custom headers, cookies, user agent | Forwarding rules, blocked headers, cookie format |
| Reliability | Retries, timeout handling, status interpretation | Retryable errors, idempotency, rate-limit headers |
| Scale | Concurrency, async jobs, polling, webhooks | Queue limits, callback signing, result retention |
| Cost | Credits and the Scrape.do-Request-Cost header |
Successful-result charging, failure charging, surcharges |
1. Inventory the existing Scrape.do integration
- Search source, deployment manifests, secrets, and runbooks for the Scrape.do host, token names, proxy settings, callbacks, and cost metrics.
- Capture one real request for each code path: ordinary API mode, rendered page, geographic route, session request, and asynchronous task if used.
- Separate required behavior from experiments. A parameter added during debugging may no longer be needed and can create unnecessary cost or incompatibility.
- Write down the response fields your parser consumes: status, headers, body, extracted metadata, and error payloads.
Scrape.do API mode requires an account token and target URL. Its documentation states that API-mode target URLs must be URL-encoded so query parameters are not misinterpreted. Preserve that rule in your inventory, then confirm the destination’s encoding requirements rather than assuming parameter names are portable.
2. Confirm which Scrape.do access pattern you use
API mode
API mode is an explicit request to the scraping endpoint. Record the target URL, token placement, rendering and wait controls, proxy or geography controls, headers, sessions, timeout, and retry logic.
Proxy mode
Proxy mode routes ordinary HTTP(S) traffic through proxy.scrape.do:8080, with the token and parameters embedded in proxy credentials. Scrape.do documents TLS certificate implications and says customHeaders=true by default. Proxy mode and API mode use the same subscription, but they are different integration contracts. A destination API that only accepts explicit requests cannot be substituted by changing a proxy hostname.
3. Build a behavior mapping before writing code
Create one row for every behavior your application depends on. Mark each destination capability as supported, different semantics, or unavailable. Do not map by spelling alone: two providers can both expose a parameter called render while waiting for different browser events.
Behavior Scrape.do setting Destination decision
Target URL encoded url parameter documented URL field
Geography proxy/country option country or region capability
Residential/mobile proxy class equivalent route or redesign
Sticky session session option cookie/session identifier
Custom headers headers/customHeaders allowed header list
JavaScript rendering/headless browser execution mode
Wait selector/time/network idle exact wait condition
Retry client or provider retries retryable status matrix
Response HTML/status/headers parser and error adapter
Keep an adapter boundary in your application. Your scraper should call a provider-neutral function such as fetch_page(target, options); only that adapter should know the destination endpoint and authentication format.
4. A provider-neutral adapter you can run
The following examples use environment variables so endpoint names and options are not invented. Set DESTINATION_ENDPOINT and DESTINATION_TOKEN to values from the provider you selected. Replace the option names with that provider’s documented contract.
cURL
curl --fail-with-body --get "$DESTINATION_ENDPOINT" \
--header "Authorization: Bearer $DESTINATION_TOKEN" \
--data-urlencode "url=https://example.com/products" \
--data "render=true" \
--data "timeout=60000" \
--output page.html
Python
import os
import requests
endpoint = os.environ["DESTINATION_ENDPOINT"]
token = os.environ["DESTINATION_TOKEN"]
params = {
"url": "https://example.com/products",
"render": "true",
"timeout": 60000,
}
response = requests.get(
endpoint,
params=params,
headers={"Authorization": f"Bearer {token}"},
timeout=90,
)
response.raise_for_status()
with open("page.html", "wb") as output:
output.write(response.content)
Node.js
const endpoint = process.env.DESTINATION_ENDPOINT;
const token = process.env.DESTINATION_TOKEN;
const query = new URLSearchParams({
url: 'https://example.com/products',
render: 'true',
timeout: '60000'
});
const response = await fetch(`${endpoint}?${query}`, {
headers: { Authorization: `Bearer ${token}` }
});
if (!response.ok) {
throw new Error(`scrape request failed: ${response.status} ${await response.text()}`);
}
const body = Buffer.from(await response.arrayBuffer());
await Bun.write('page.html', body);
For production, replace the simple error path with a typed result object containing provider, request ID, target URL hash, HTTP status, retry classification, and billing metadata. Never log tokens or full URLs when they may contain credentials.
5. Rebuild asynchronous scraping explicitly
If you use Scrape.do Async API, treat migration as a separate workflow. Scrape.do documents an async base URL of https://q.scrape.do, X-Token authentication, job and task identifiers, status polling, webhooks, cancellation, and expiring results.
- Create a job and persist the job/task identifier durably.
- Poll with exponential backoff only when webhooks are not available or as a reconciliation process.
- Validate status and error fields before downloading a result.
- Retrieve results before the provider’s expiration window.
- Make webhook handling idempotent: store event IDs or task states and safely repeat a delivery.
- Map destination queue and concurrency limits separately from synchronous limits.
Do not assume a synchronous endpoint can absorb an asynchronous workload. Recheck callback signing, result retention, cancellation, and maximum batch size in the destination documentation.
6. Rebaseline cost and capacity
Do not convert Scrape.do credits directly into destination request counts. Scrape.do documents base costs of 1 credit for a standard datacenter request, 5 with headless rendering, 10 for residential/mobile, and 25 for residential/mobile plus rendering. Costs can vary by target domain and options; the Scrape.do-Request-Cost response header is the authoritative cost for an actual call.
Measure a representative sample by target domain and mode. Compare effective cost per valid result, failed-request charging, retry charging, concurrency, geographic availability, rendering, asynchronous throughput, and rate limits. Record p50 and p95 latency and the percentage of responses that pass your content validation.
7. Validate providers in parallel
- Select static pages, JavaScript-heavy pages, redirects, consent walls, session-dependent pages, regional pages, and known difficult domains.
- Replay identical targets and required options against both providers.
- Compare status codes, extracted fields, completeness, title and canonical URL, latency, error category, retry outcome, and effective cost.
- Store raw responses temporarily for debugging, with secrets and sensitive data redacted.
- Define acceptance thresholds before looking at results. For example: required fields present, no increase in invalid pages, bounded latency, and a cost ceiling.
8. Cut over gradually
Deploy the destination behind a feature flag or routing percentage. Start with internal traffic or a small production slice, monitor validity as well as HTTP success, and increase traffic only when acceptance criteria hold. Keep the Scrape.do route available until queued jobs, webhooks, billing reconciliation, and rollback have been exercised.
Common migration errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Target URL loses its query string | URL was not encoded or was encoded twice | Use the destination SDK or one standards-compliant URL encoder; log the parsed target safely. |
| HTTP 401 or 403 | Token moved from a query parameter to a header, or wrong token type | Check authentication placement and environment selection. |
| Static HTML replaces rendered content | JavaScript mode is absent or wait condition is too short | Enable documented rendering and wait for a selector or network-idle condition. |
| Different regional content | Proxy geography is unsupported or silently falls back | Verify the response region and destination’s fallback semantics. |
| Logged-in pages fail | Cookies, authorization headers, or sticky sessions were not migrated | Recreate the session contract and confirm header/cookie forwarding rules. |
| Costs rise unexpectedly | Rendering, residential routing, retries, or domain surcharges differ | Capture billing headers/metadata and compare cost per valid result. |
| Async jobs disappear | Result retention or task expiration differs | Download promptly and persist task state; do not rely on indefinite polling. |
| Webhook duplicates data | Handler is not idempotent | Use a task ID or event ID as a deduplication key. |
Performance and reliability checklist
- Reuse HTTP connections and set explicit client timeouts.
- Bound concurrency below the provider and target-site limits.
- Use exponential backoff with jitter for 429 and transient 5xx responses.
- Do not retry authentication errors, invalid URLs, or deterministic parsing failures.
- Track valid-result rate, not only transport success.
- Cache immutable pages where permitted and avoid duplicate requests.
- Use async jobs for long-running browser work and bulk workloads.
- Keep provider credentials in a secret manager and rotate them during cutover.
Or skip the browser setup
If your goal is clean screenshots rather than raw HTML extraction, ScreenshotNeo provides a one-call website screenshot API and MCP server. See the ScreenshotNeo documentation for the full option list.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
How do I replace Scrape.do in my scraper?
Put the provider call behind an adapter, inventory required behaviors, map each behavior to documented destination controls, then validate parallel results before changing traffic.
Can I keep my existing Scrape.do parameters?
Keep the concepts your application needs, but translate parameter names and semantics. A successful HTTP response does not prove equivalent rendering, geography, session, or retry behavior.
Should I migrate API mode and Proxy Mode the same way?
No. API mode is an explicit scraping request; Proxy Mode is a network proxy contract. Inventory and test them separately.
When should I use asynchronous jobs?
Use them for long browser renders, large queues, or workloads that need webhooks and independent concurrency. Persist identifiers and retrieve results before expiration.
What is the safest cutover?
Run representative targets in parallel, define acceptance criteria, shift a small traffic percentage, monitor valid results and cost, and retain a tested rollback route.


