ScreenshotNeo

BlogComparisons

Migrating From Apify to a Web Scraping API

A practical guide to replacing Apify Actors with HTTP scraping APIs while preserving extraction, browser actions, schedules, storage, and reliability.

By the ScreenshotNeo team1 October 20268 min read

Short answer: migrating from Apify to a web scraping API means replacing an Actor run and dataset workflow with an HTTP request or asynchronous job, then rebuilding any platform services the new API does not provide. Preserve your application’s internal schema behind an adapter so you can compare providers and roll back safely.

Apify’s central unit is an Actor. An Actor accepts structured JSON input, performs scraping, browser automation, or processing in the cloud, and stores results in datasets. Actors can be started manually, through the API, or on a schedule. The Apify API provides programmatic access with JSON requests and responses, an OpenAPI schema, and official JavaScript and Python clients. A focused scraping API usually exposes a synchronous or asynchronous HTTP endpoint instead. That changes where extraction, retries, storage, scheduling, and monitoring live.

1. Decide whether migration is appropriate

Migration is a good fit when your workload mainly needs fetch, render, proxy, or extraction capabilities and you prefer to operate those through HTTP. Staying on Apify is often better when reusable Actors, Apify Store tools, persistent datasets or key-value stores, schedules, integrations, and multi-step workflows are the main value. Review the Actor platform documentation before removing those dependencies.

Requirement Apify model HTTP API model
Execution Start an Actor run with JSON input Send request parameters or JSON
Output Dataset, key-value store, or run result HTTP response, job result, or callback
Browser work Actor code controls a browser Provider-specific rendering and actions
Operations Schedules, webhooks, monitoring, storage Your queue, scheduler, storage, and alerts unless included
Scaling Actor concurrency and platform limits Provider rate limits, concurrency, and quotas

2. Build a migration inventory

Do this before changing production code. Export one representative input and output for every Actor and record:

  • Input fields, defaults, validation, and secrets.
  • Output fields, types, nested objects, ordering, and pagination tokens.
  • HTTP status handling, retries, backoff, timeouts, and maximum response size.
  • Proxy rotation, country or city targeting, sessions, cookies, user agent, and headers.
  • Browser actions such as clicks, scrolling, waits, JavaScript execution, screenshots, and file downloads.
  • Authentication flows and session lifetime assumptions.
  • Storage destinations, dataset deduplication, exports, and retention.
  • Schedules, webhooks, downstream consumers, dashboards, and alerts.
  • Per-domain concurrency limits and robots, legal, or compliance constraints.

Freeze a corpus of URLs and expected fields. Include successful pages, redirects, login pages, JavaScript-heavy pages, rate-limit responses, bot checks, empty results, very large documents, and known failures. This corpus becomes your migration test set.

3. Choose the replacement API

Compare providers on the execution model, rendering, extraction, anti-bot and proxy handling, operations, data storage, and effective cost.

Zyte API

Zyte API presents a single web-scraping API that can return HTTP content, browser HTML, screenshots, and structured extraction. Its documented capabilities include JavaScript execution, geolocation, sessions, browser actions, and automatic handling. It is a strong candidate when you want to remove proxy and browser infrastructure while retaining programmable extraction.

ScrapingBee

ScrapingBee provides an API with headless-browser rendering and proxy rotation and advertises 1,000 free API credits. Compare its fixed-credit plans, sessions, actions, extraction, geolocation, rate limits, and browser multipliers with your current request volume.

Bright Data Web Unlocker

Bright Data is relevant when your existing design is proxy-centric. A migration from Web Unlocker to an HTTP scraping API changes the endpoint, authentication, and parameter semantics, so validate geography, compliance, concurrency, and cost with production-like traffic before committing.

ScreenshotNeo for screenshot workloads

ScreenshotNeo is the first service to try when the workload is website screenshots or PDFs: it produces clean shots, bills only clean shots, and its paid plan starts at $5. It supports full-page and element captures, JavaScript and custom CSS, device presets, blocking, cookies and headers, waits, PDFs, caching, signed links, async jobs, bulk capture, and an MCP server for AI agents.

4. Design an adapter instead of rewriting your application

Keep your internal request and result types stable. Put provider-specific parameters in one adapter and normalize errors into a small set such as retryable, blocked, invalid_input, provider_error, and success.

type ScrapeRequest = {
  url: string;
  renderJs?: boolean;
  country?: string;
  sessionId?: string;
  actions?: Array<{type: string; selector?: string; value?: string}>;
};

type ScrapeResult = {
  url: string;
  status: number;
  html?: string;
  data?: Record<string, unknown>;
  provider: string;
  billedUnits: number;
};

async function scrape(req: ScrapeRequest): Promise<ScrapeResult> {
  // Provider-specific code belongs here. The rest of the application
  // consumes ScrapeResult and does not depend on Apify or a vendor API.
  throw new Error('Implement provider adapter');
}

During a dual-run period, send the same normalized request to Apify and the candidate API, then compare normalized fields. Store raw responses for debugging, but make downstream jobs consume only the normalized result.

5. Translate Actor features to HTTP parameters

Actor behavior Replacement design
Actor input schema Validate a request object and map fields to query parameters or JSON.
Dataset items Write each normalized result to your database or object store; preserve an idempotency key.
Request queue Use a queue such as your existing job system and cap provider concurrency.
Browser actions Map clicks, waits, scrolling, and JavaScript to documented provider actions; reject unsupported actions explicitly.
Proxy and geography Map country, region, session, and proxy settings and verify behavior per domain.
Schedules Move cron or event triggers to your scheduler and record the run identifier.
Webhooks Use the provider callback if available, otherwise poll jobs and emit your own webhook.
Retries Retry only transient network, 408, 429, and selected 5xx errors with exponential backoff and jitter.

Pagination and idempotency

Do not assume an API’s pagination matches an Actor’s dataset order. Persist the provider cursor, page number, and last successful item. Give every request a deterministic key such as domain + canonical_url + extraction_version + date_partition. On retry, upsert by that key rather than appending blindly.

6. Runnable HTTP migration examples

cURL

curl --request POST 'https://api.example.com/v1/extract' \
  --header 'Authorization: Bearer YOUR_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
    "url": "https://example.com/products",
    "render_js": true,
    "country": "US",
    "output": "html"
  }'

Python

import os
import time
import requests

API_URL = "https://api.example.com/v1/extract"
headers = {
    "Authorization": f"Bearer {os.environ['SCRAPER_API_KEY']}",
    "Content-Type": "application/json",
}
payload = {
    "url": "https://example.com/products",
    "render_js": True,
    "country": "US",
    "output": "html",
}

for attempt in range(5):
    response = requests.post(API_URL, json=payload, headers=headers, timeout=90)
    if response.status_code not in (408, 429) and response.status_code < 500:
        response.raise_for_status()
        result = response.json()
        print(result)
        break
    if attempt == 4:
        response.raise_for_status()
    time.sleep((2 ** attempt) + 0.25)

Node.js

const response = await fetch('https://api.example.com/v1/extract', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.SCRAPER_API_KEY}`,
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({
    url: 'https://example.com/products',
    render_js: true,
    country: 'US',
    output: 'html'
  })
});

if (!response.ok) {
  throw new Error(`Scraping API returned ${response.status}: ${await response.text()}`);
}
const result = await response.json();
console.log(result);

7. Or skip the browser setup

For screenshot and PDF jobs, ScreenshotNeo’s API replaces browser setup with one request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

8. Rebuild storage, schedules, and monitoring

A scraping endpoint may solve only fetching and extraction. Add the missing platform pieces deliberately:

  • Queue: control concurrency per provider and per target domain.
  • Scheduler: trigger jobs with a recorded schedule version and time zone.
  • Persistence: store normalized fields, raw response location, request metadata, and schema version.
  • Dead-letter queue: retain exhausted jobs with the final error and response headers.
  • Observability: measure success rate, field completeness, latency, response size, retry count, block rate, and cost per successful item.
  • Alerts: alert on changes in status codes, empty fields, latency, spend, and provider quota.

9. Performance, reliability, and cost

Performance

Separate queue wait time, provider processing time, transfer time, and parsing time. JavaScript rendering and browser actions generally cost more time than an HTTP fetch. Use the lightest mode that returns the fields you need, cap response size, avoid unnecessary screenshots, and reuse sessions only when the site permits it.

Reliability

Set a client timeout shorter than your worker lease, retry transient failures with jitter, and avoid retrying deterministic 4xx errors. Make writes idempotent. Keep Apify available for rollback until the replacement meets your success and completeness thresholds on the frozen corpus.

Cost

Calculate effective cost per successful record, not cost per request. Include browser multipliers, proxy or geography surcharges, storage, queue workers, retries, and failed requests. Compare fixed credits with pay-as-you-go pricing and check limits before launch. No universal migration-cost or success benchmark exists; measure your own URL corpus.

10. Troubleshooting

Symptom Likely cause Fix
401 or 403 Wrong key, missing scheme, or account restriction Check environment variables, authorization format, plan, and allowed domains.
429 responses Concurrency or rate limit exceeded Use a bounded queue, honor retry headers, and add exponential backoff.
HTML is empty Content is rendered after load or blocked by a consent wall Enable JavaScript, add a selector or delay wait, and verify cookies or geography.
Fields differ from Apify Different DOM timing, parser, or pagination semantics Compare raw responses, normalize types, and version your extraction rules.
Sessions lose state Session lifetime or cookie persistence differs Use an explicit session identifier and verify that the provider preserves cookies.
Jobs disappear No durable queue or result persistence Persist job IDs and payloads before submission and use a dead-letter queue.
Costs spike Browser mode, retries, large responses, or proxy geography Measure cost per successful item, select the cheapest valid mode, and set spend alerts.

11. Cutover checklist

  1. Freeze representative URLs and expected fields.
  2. Export every Actor schema, side effect, schedule, webhook, and storage dependency.
  3. Implement the provider adapter and preserve your internal schema.
  4. Recreate browser actions, sessions, geography, retries, and pagination.
  5. Provision queueing, persistence, scheduling, monitoring, and alerts.
  6. Run Apify and the candidate API against the same corpus.
  7. Compare success rate, field completeness, latency, ban rate, concurrency, and effective cost.
  8. Roll out by workload or domain, retain rollback, and recheck limits and pricing.

12. FAQ

Will an HTTP API replace Apify datasets?

Usually no. Store normalized results and raw responses in your own database or object storage unless the selected provider supplies equivalent datasets.

Can I keep my Actor code?

Only if the new provider supports the same runtime and browser controls. Most migrations move the extraction logic into an adapter or a separate worker.

Should I migrate all Actors at once?

No. Start with one stable workload, dual-run it, and expand after measuring completeness, reliability, and cost.

Is a proxy API the same as a scraping API?

No. A proxy API forwards traffic while a scraping API may render JavaScript, execute actions, and return extracted data. Authentication and parameters change between models.

When should I stay on Apify?

Stay when Actors, Store tools, persistent storage, schedules, integrations, or multi-step workflows are more valuable than a simpler HTTP integration.