ScreenshotNeo

BlogEngineering

Migrating From Decodo to a Web Scraping API

A practical migration plan for moving from Decodo: inventory the contract, build an adapter, compare providers, and cut over safely.

By the ScreenshotNeo team1 October 20268 min read

Migrating from Decodo to another web scraping API is an interface-compatibility project. Preserve your application’s normalized data contract, put provider-specific requests behind an adapter, run both providers against the same corpus, and switch traffic gradually only after output quality and effective cost are understood.

Decodo documents a Web Scraping API with target templates, JavaScript rendering, geo-targeted proxy pools and outputs such as HTML, JSON, CSV, XHR, PNG and Markdown. Its documented task example uses POST https://scraper-api.decodo.com/v1/tasks with authorization plus fields including target, url, proxy_pool, headless and locale. Treat those capabilities and names as the contract you must inventory, not as assumptions you can silently carry into a replacement.

1. Freeze the Decodo contract before changing code

Export a table for every production job. Record the exact request and the fields your parser expects.

Area Record Why it matters
Endpoint and auth URL, HTTP method, authorization header, key rotation Prevents hidden authentication differences
Target Template name, generic URL mode, required parameters Template coverage is rarely identical between vendors
Rendering headless, JavaScript execution, wait condition, device profile Server HTML and browser HTML can differ materially
Network Proxy pool, country, language, session or sticky-IP behavior Geography and block rates depend on these controls
Output HTML, JSON, CSV, XHR, PNG, Markdown, encoding and envelope shape Downstream parsers should keep their current schema
Reliability Timeout, retries, backoff, idempotency, pagination checkpoints A successful HTTP response is not necessarily a complete record
Billing Standard versus premium proxy, JavaScript mode, request units A blended average can hide the expensive paths

2. Define a provider-neutral schema and adapter

Keep business code independent of vendor parameter names. The adapter should accept one internal request type and return one internal result type.

// provider.js
export function normalizeRequest(input) {
  return {
    url: input.url,
    target: input.target ?? "universal",
    country: input.country ?? null,
    locale: input.locale ?? "en-US",
    javascript: Boolean(input.javascript),
    proxyTier: input.proxyTier ?? "standard",
    timeoutMs: input.timeoutMs ?? 60000,
    page: input.page ?? null
  };
}

export function validateResult(result) {
  if (!result || result.status !== "ok") {
    throw new Error(`provider status: ${result?.status ?? "missing"}`);
  }
  if (!result.url || result.body == null) {
    throw new Error("provider returned an incomplete record");
  }
  return result;
}

Do not expose Decodo’s response envelope to the rest of the application. Map it once:

export function fromDecodo(taskResponse) {
  return validateResult({
    status: taskResponse.status === "success" ? "ok" : "error",
    url: taskResponse.url,
    body: taskResponse.results ?? taskResponse.html ?? taskResponse.data,
    provider: "decodo",
    raw: taskResponse
  });
}

export function fromReplacement(response, requestedUrl) {
  // Adjust these mappings to the replacement's documented response.
  return validateResult({
    status: response.ok ? "ok" : "error",
    url: response.url ?? requestedUrl,
    body: response.data ?? response.html ?? response.body,
    provider: "replacement",
    raw: response
  });
}

3. Translate Decodo controls explicitly

Build a mapping document before implementing the replacement adapter. Every row needs a test case and an owner.

Decodo concept Replacement question Acceptance test
target Is there an equivalent template, or must you parse a generic URL? Required fields are present for representative pages
headless / JavaScript Is browser rendering enabled, and what wait controls exist? Client-rendered content appears in the response
proxy_pool Which pool and premium tier provide comparable access? Block and challenge rates stay within your threshold
locale and country Are language, IP country and timezone separate settings? Localized price, currency and content match expectations
Session behavior Can cookies, sticky sessions or browser state persist? Login or multi-page flows retain required state
Pagination Does the API paginate tasks or return one page per request? Checkpointing resumes without duplicates

4. Implement a replacement adapter

The replacement endpoint and field names must come from its current documentation. Keep them isolated in one module. This example shows the shape without pretending that an unspecified provider has a fixed URL or schema.

# replacement_adapter.py
import time
import requests

class ReplacementClient:
    def __init__(self, endpoint, api_key):
        self.endpoint = endpoint
        self.session = requests.Session()
        self.session.headers.update({"Authorization": f"Bearer {api_key}"})

    def fetch(self, req):
        payload = {
            "url": req["url"],
            "target": req.get("target", "universal"),
            "javascript": req.get("javascript", False),
            "country": req.get("country"),
            "locale": req.get("locale", "en-US"),
            "proxy_tier": req.get("proxyTier", "standard")
        }
        response = self.session.post(
            self.endpoint,
            json=payload,
            timeout=req.get("timeoutMs", 60000) / 1000
        )
        response.raise_for_status()
        data = response.json()
        if not data.get("data") and not data.get("html") and not data.get("body"):
            raise ValueError("incomplete replacement response")
        return data

def fetch_with_backoff(client, req, attempts=3):
    for attempt in range(attempts):
        try:
            return client.fetch(req)
        except (requests.Timeout, requests.ConnectionError) as exc:
            if attempt == attempts - 1:
                raise
            time.sleep(2 ** attempt)

5. Keep a Decodo reference client during the migration

Use the documented Decodo task shape as a reference implementation. Do not change parsers until both providers produce normalized records.

import requests

def decodo_task(api_key, url, target="universal", proxy_pool="standard", headless=True, locale="en-US"):
    response = requests.post(
        "https://scraper-api.decodo.com/v1/tasks",
        headers={"Authorization": f"Bearer {api_key}"},
        json={
            "target": target,
            "url": url,
            "proxy_pool": proxy_pool,
            "headless": headless,
            "locale": locale
        },
        timeout=90
    )
    response.raise_for_status()
    return response.json()

6. Run shadow traffic and compare useful outcomes

  1. Select URLs covering every target, country, device mode, JavaScript path and pagination state.
  2. Send identical logical requests to Decodo and the replacement.
  3. Normalize both responses before comparing them.
  4. Store status, block or challenge outcome, required-field completeness, encoding, response size, latency percentiles and cost.
  5. Inspect a sample of raw HTML or JSON manually; field counts alone can miss incorrect pages.
Metric Comparison rule
Completeness Validate required fields and types, not only HTTP status
Block rate Classify CAPTCHA, bot checks, access denied and empty responses separately
Latency Track p50, p95 and timeout rate by target and rendering mode
Cost Calculate cost per successful record for standard, premium and JavaScript paths
Stability Repeat the corpus across time windows and geographies

7. Cut over gradually and preserve rollback

  1. Release the adapter behind a feature flag.
  2. Route a small percentage of traffic to the replacement.
  3. Alert on completeness, block rate, timeout rate, latency and effective cost.
  4. Increase traffic only after each important target passes its acceptance thresholds.
  5. Keep Decodo credentials, parsers and rollback configuration until historical jobs and pagination checkpoints have completed.

8. cURL, Python and Node.js migration smoke tests

Replace the placeholder endpoint and fields with the provider’s documented values. These tests verify authentication, response shape and timeout behavior before you integrate the adapter.

curl -X POST "$REPLACEMENT_ENDPOINT" \
  -H "Authorization: Bearer $REPLACEMENT_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{"url":"https://example.com","javascript":true,"locale":"en-US"}'
import os, requests
r = requests.post(
    os.environ["REPLACEMENT_ENDPOINT"],
    headers={"Authorization": f"Bearer {os.environ['REPLACEMENT_API_KEY']}"},
    json={"url": "https://example.com", "javascript": True, "locale": "en-US"},
    timeout=90,
)
r.raise_for_status()
print(r.json())
const res = await fetch(process.env.REPLACEMENT_ENDPOINT, {
  method: 'POST',
  headers: {
    Authorization: `Bearer ${process.env.REPLACEMENT_API_KEY}`,
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({ url: 'https://example.com', javascript: true, locale: 'en-US' })
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
console.log(await res.json());

9. Or skip the browser setup

If your requirement is a clean visual capture rather than structured fields, ScreenshotNeo is the alternative to try first. It is a website screenshot API and MCP server. Cookie and consent banners are accepted and 60+ known consent platforms, newsletter popups and chat widgets are removed before capture; bot checks, blank pages and failed loads are not billed. An MCP server lets Claude, Cursor and other MCP clients take screenshots, inspect pages and capture PDFs. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots.

See the ScreenshotNeo API documentation for the complete option list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Sign up free for 1,000 screenshots a month with no card.

10. Troubleshooting

Symptom Likely cause Fix
401 or 403 Header format, expired key or wrong project Check the replacement’s auth scheme and rotate credentials safely
200 response with empty fields Challenge page, wrong target mapping or rendering disabled Classify body content, enable JavaScript where required and validate fields
Localized data differs IP country, locale, timezone or cookies were not mapped Set each control explicitly and compare raw responses
Timeouts increased Browser rendering, proxy tier or wait condition differs Set a bounded wait, raise timeout selectively and measure by mode
Duplicate pages after retry Non-idempotent task creation Use idempotency keys or persist task state before retrying
Parser crashes Envelope, encoding or nullability changed Normalize at the adapter boundary and add schema validation
Costs exceed estimate Premium proxy or JavaScript requests are priced differently Report cost per successful record by request class

11. Performance, reliability and cost notes

  • Reuse HTTP connections and cap concurrency below the provider’s documented limit.
  • Separate browser-rendered jobs from simple HTTP jobs so slow pages do not consume every worker.
  • Use exponential backoff for transient network failures, with a maximum attempt count and a dead-letter queue.
  • Persist pagination checkpoints and normalized records before acknowledging a job.
  • Cache only when the source’s freshness requirements permit it.
  • Model standard and premium proxy requests separately. Decodo’s displayed pricing examples include $19, $49 and $99 monthly plans, with request prices varying by proxy tier and JavaScript use; those figures are time-sensitive procurement data.
  • Do not treat vendor claims such as a stated 99.99% success rate or 125M+ IP network as an independent benchmark. Recheck current documentation before purchase.

12. Migration checklist

  • Every endpoint, target, output and parser is inventoried.
  • A provider-neutral request and response schema is enforced.
  • JavaScript, geography, locale, proxy tier and session behavior have explicit mappings.
  • Retries, idempotency, pagination and timeout budgets are implemented.
  • Shadow results are compared for completeness, blocks, latency and effective cost.
  • Feature-flagged rollout and rollback are ready.
  • Pricing, rate limits, acceptable-use rules and data-handling requirements were rechecked with the replacement provider.

FAQ

Can I keep my existing Decodo parser?

Keep the parser’s normalized output contract, but add an adapter that converts the replacement response into that contract. Directly parsing a new vendor envelope spreads migration work through the application.

Is a generic URL scraper equivalent to a Decodo target template?

No. A template may provide provider-generated fields or specialized extraction. Prove equivalence with required-field and sample-record tests before removing a target.

Should every HTTP 200 be counted as success?

No. Validate required fields, content type, challenge indicators and pagination state. A bot page can be returned with HTTP 200.

When does ScreenshotNeo fit this migration?

Use ScreenshotNeo when the output you need is a clean PNG, JPEG, WebP or PDF, or when an AI agent needs screenshot and page-inspection tools. It is not a drop-in replacement for structured Decodo target extraction.