ScreenshotNeo

BlogEngineering

Migrating From Scrape.do to a Web Scraping API

A provider-neutral migration plan for replacing Scrape.do: inventory controls, map behavior, validate results, rebaseline cost, and cut over safely.

By the ScreenshotNeo team1 October 20268 min read

Replacing Scrape.do is an API contract migration, not a search-and-replace exercise. First inventory how your integration authenticates, encodes target URLs, selects proxies, preserves sessions, forwards headers, renders JavaScript, waits for content, retries failures, reports cost, and handles asynchronous jobs. Then map each required behavior to a documented capability at the destination provider, run both providers against representative pages, and shift traffic gradually with rollback ready.

What do I need to change when switching scraping API providers?

Change the provider-specific contract while preserving the page behavior your application actually needs. The destination may use a different endpoint, authentication header, request body, URL encoding rule, proxy model, rendering switch, wait syntax, error format, concurrency limit, webhook contract, and billing unit.

Contract area Record in the Scrape.do integration Map and verify at the destination
Target Absolute URL, HTTP method, body, URL encoding Accepted protocols, encoding, methods, redirects
Authentication Account token location and secret loading Query parameter, header, signed request, scopes
Network route Datacenter, residential/mobile, geography Equivalent proxy classes, countries, fallback behavior
Session Sticky session or session identifier Cookie jar, session lifetime, reuse semantics
Browser behavior JavaScript rendering, waits, screenshots or HTML Rendering engine, selector/network-idle waits, timeout limits
Request shaping Custom headers, cookies, user agent Forwarding rules, blocked headers, cookie format
Reliability Retries, timeout handling, status interpretation Retryable errors, idempotency, rate-limit headers
Scale Concurrency, async jobs, polling, webhooks Queue limits, callback signing, result retention
Cost Credits and the Scrape.do-Request-Cost header Successful-result charging, failure charging, surcharges

1. Inventory the existing Scrape.do integration

  1. Search source, deployment manifests, secrets, and runbooks for the Scrape.do host, token names, proxy settings, callbacks, and cost metrics.
  2. Capture one real request for each code path: ordinary API mode, rendered page, geographic route, session request, and asynchronous task if used.
  3. Separate required behavior from experiments. A parameter added during debugging may no longer be needed and can create unnecessary cost or incompatibility.
  4. Write down the response fields your parser consumes: status, headers, body, extracted metadata, and error payloads.

Scrape.do API mode requires an account token and target URL. Its documentation states that API-mode target URLs must be URL-encoded so query parameters are not misinterpreted. Preserve that rule in your inventory, then confirm the destination’s encoding requirements rather than assuming parameter names are portable.

2. Confirm which Scrape.do access pattern you use

API mode

API mode is an explicit request to the scraping endpoint. Record the target URL, token placement, rendering and wait controls, proxy or geography controls, headers, sessions, timeout, and retry logic.

Proxy mode

Proxy mode routes ordinary HTTP(S) traffic through proxy.scrape.do:8080, with the token and parameters embedded in proxy credentials. Scrape.do documents TLS certificate implications and says customHeaders=true by default. Proxy mode and API mode use the same subscription, but they are different integration contracts. A destination API that only accepts explicit requests cannot be substituted by changing a proxy hostname.

3. Build a behavior mapping before writing code

Create one row for every behavior your application depends on. Mark each destination capability as supported, different semantics, or unavailable. Do not map by spelling alone: two providers can both expose a parameter called render while waiting for different browser events.

Behavior                 Scrape.do setting             Destination decision
Target URL               encoded url parameter          documented URL field
Geography                proxy/country option           country or region capability
Residential/mobile       proxy class                    equivalent route or redesign
Sticky session           session option                 cookie/session identifier
Custom headers           headers/customHeaders          allowed header list
JavaScript               rendering/headless              browser execution mode
Wait                     selector/time/network idle     exact wait condition
Retry                    client or provider retries     retryable status matrix
Response                 HTML/status/headers             parser and error adapter

Keep an adapter boundary in your application. Your scraper should call a provider-neutral function such as fetch_page(target, options); only that adapter should know the destination endpoint and authentication format.

4. A provider-neutral adapter you can run

The following examples use environment variables so endpoint names and options are not invented. Set DESTINATION_ENDPOINT and DESTINATION_TOKEN to values from the provider you selected. Replace the option names with that provider’s documented contract.

cURL

curl --fail-with-body --get "$DESTINATION_ENDPOINT" \
  --header "Authorization: Bearer $DESTINATION_TOKEN" \
  --data-urlencode "url=https://example.com/products" \
  --data "render=true" \
  --data "timeout=60000" \
  --output page.html

Python

import os
import requests

endpoint = os.environ["DESTINATION_ENDPOINT"]
token = os.environ["DESTINATION_TOKEN"]
params = {
    "url": "https://example.com/products",
    "render": "true",
    "timeout": 60000,
}
response = requests.get(
    endpoint,
    params=params,
    headers={"Authorization": f"Bearer {token}"},
    timeout=90,
)
response.raise_for_status()
with open("page.html", "wb") as output:
    output.write(response.content)

Node.js

const endpoint = process.env.DESTINATION_ENDPOINT;
const token = process.env.DESTINATION_TOKEN;
const query = new URLSearchParams({
  url: 'https://example.com/products',
  render: 'true',
  timeout: '60000'
});
const response = await fetch(`${endpoint}?${query}`, {
  headers: { Authorization: `Bearer ${token}` }
});
if (!response.ok) {
  throw new Error(`scrape request failed: ${response.status} ${await response.text()}`);
}
const body = Buffer.from(await response.arrayBuffer());
await Bun.write('page.html', body);

For production, replace the simple error path with a typed result object containing provider, request ID, target URL hash, HTTP status, retry classification, and billing metadata. Never log tokens or full URLs when they may contain credentials.

5. Rebuild asynchronous scraping explicitly

If you use Scrape.do Async API, treat migration as a separate workflow. Scrape.do documents an async base URL of https://q.scrape.do, X-Token authentication, job and task identifiers, status polling, webhooks, cancellation, and expiring results.

  1. Create a job and persist the job/task identifier durably.
  2. Poll with exponential backoff only when webhooks are not available or as a reconciliation process.
  3. Validate status and error fields before downloading a result.
  4. Retrieve results before the provider’s expiration window.
  5. Make webhook handling idempotent: store event IDs or task states and safely repeat a delivery.
  6. Map destination queue and concurrency limits separately from synchronous limits.

Do not assume a synchronous endpoint can absorb an asynchronous workload. Recheck callback signing, result retention, cancellation, and maximum batch size in the destination documentation.

6. Rebaseline cost and capacity

Do not convert Scrape.do credits directly into destination request counts. Scrape.do documents base costs of 1 credit for a standard datacenter request, 5 with headless rendering, 10 for residential/mobile, and 25 for residential/mobile plus rendering. Costs can vary by target domain and options; the Scrape.do-Request-Cost response header is the authoritative cost for an actual call.

Measure a representative sample by target domain and mode. Compare effective cost per valid result, failed-request charging, retry charging, concurrency, geographic availability, rendering, asynchronous throughput, and rate limits. Record p50 and p95 latency and the percentage of responses that pass your content validation.

7. Validate providers in parallel

  1. Select static pages, JavaScript-heavy pages, redirects, consent walls, session-dependent pages, regional pages, and known difficult domains.
  2. Replay identical targets and required options against both providers.
  3. Compare status codes, extracted fields, completeness, title and canonical URL, latency, error category, retry outcome, and effective cost.
  4. Store raw responses temporarily for debugging, with secrets and sensitive data redacted.
  5. Define acceptance thresholds before looking at results. For example: required fields present, no increase in invalid pages, bounded latency, and a cost ceiling.

8. Cut over gradually

Deploy the destination behind a feature flag or routing percentage. Start with internal traffic or a small production slice, monitor validity as well as HTTP success, and increase traffic only when acceptance criteria hold. Keep the Scrape.do route available until queued jobs, webhooks, billing reconciliation, and rollback have been exercised.

Common migration errors and fixes

Symptom Likely cause Fix
Target URL loses its query string URL was not encoded or was encoded twice Use the destination SDK or one standards-compliant URL encoder; log the parsed target safely.
HTTP 401 or 403 Token moved from a query parameter to a header, or wrong token type Check authentication placement and environment selection.
Static HTML replaces rendered content JavaScript mode is absent or wait condition is too short Enable documented rendering and wait for a selector or network-idle condition.
Different regional content Proxy geography is unsupported or silently falls back Verify the response region and destination’s fallback semantics.
Logged-in pages fail Cookies, authorization headers, or sticky sessions were not migrated Recreate the session contract and confirm header/cookie forwarding rules.
Costs rise unexpectedly Rendering, residential routing, retries, or domain surcharges differ Capture billing headers/metadata and compare cost per valid result.
Async jobs disappear Result retention or task expiration differs Download promptly and persist task state; do not rely on indefinite polling.
Webhook duplicates data Handler is not idempotent Use a task ID or event ID as a deduplication key.

Performance and reliability checklist

  • Reuse HTTP connections and set explicit client timeouts.
  • Bound concurrency below the provider and target-site limits.
  • Use exponential backoff with jitter for 429 and transient 5xx responses.
  • Do not retry authentication errors, invalid URLs, or deterministic parsing failures.
  • Track valid-result rate, not only transport success.
  • Cache immutable pages where permitted and avoid duplicate requests.
  • Use async jobs for long-running browser work and bulk workloads.
  • Keep provider credentials in a secret manager and rotate them during cutover.

Or skip the browser setup

If your goal is clean screenshots rather than raw HTML extraction, ScreenshotNeo provides a one-call website screenshot API and MCP server. See the ScreenshotNeo documentation for the full option list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

How do I replace Scrape.do in my scraper?

Put the provider call behind an adapter, inventory required behaviors, map each behavior to documented destination controls, then validate parallel results before changing traffic.

Can I keep my existing Scrape.do parameters?

Keep the concepts your application needs, but translate parameter names and semantics. A successful HTTP response does not prove equivalent rendering, geography, session, or retry behavior.

Should I migrate API mode and Proxy Mode the same way?

No. API mode is an explicit scraping request; Proxy Mode is a network proxy contract. Inventory and test them separately.

When should I use asynchronous jobs?

Use them for long browser renders, large queues, or workloads that need webhooks and independent concurrency. Persist identifiers and retrieve results before expiration.

What is the safest cutover?

Run representative targets in parallel, define acceptance criteria, shift a small traffic percentage, monitor valid results and cost, and retain a tested rollback route.