Migrating From Apify to a Web Scraping API
A practical guide to replacing Apify Actors with HTTP scraping APIs while preserving extraction, browser actions, schedules, storage, and reliability.
Short answer: migrating from Apify to a web scraping API means replacing an Actor run and dataset workflow with an HTTP request or asynchronous job, then rebuilding any platform services the new API does not provide. Preserve your application’s internal schema behind an adapter so you can compare providers and roll back safely.
Apify’s central unit is an Actor. An Actor accepts structured JSON input, performs scraping, browser automation, or processing in the cloud, and stores results in datasets. Actors can be started manually, through the API, or on a schedule. The Apify API provides programmatic access with JSON requests and responses, an OpenAPI schema, and official JavaScript and Python clients. A focused scraping API usually exposes a synchronous or asynchronous HTTP endpoint instead. That changes where extraction, retries, storage, scheduling, and monitoring live.
1. Decide whether migration is appropriate
Migration is a good fit when your workload mainly needs fetch, render, proxy, or extraction capabilities and you prefer to operate those through HTTP. Staying on Apify is often better when reusable Actors, Apify Store tools, persistent datasets or key-value stores, schedules, integrations, and multi-step workflows are the main value. Review the Actor platform documentation before removing those dependencies.
| Requirement | Apify model | HTTP API model |
|---|---|---|
| Execution | Start an Actor run with JSON input | Send request parameters or JSON |
| Output | Dataset, key-value store, or run result | HTTP response, job result, or callback |
| Browser work | Actor code controls a browser | Provider-specific rendering and actions |
| Operations | Schedules, webhooks, monitoring, storage | Your queue, scheduler, storage, and alerts unless included |
| Scaling | Actor concurrency and platform limits | Provider rate limits, concurrency, and quotas |
2. Build a migration inventory
Do this before changing production code. Export one representative input and output for every Actor and record:
- Input fields, defaults, validation, and secrets.
- Output fields, types, nested objects, ordering, and pagination tokens.
- HTTP status handling, retries, backoff, timeouts, and maximum response size.
- Proxy rotation, country or city targeting, sessions, cookies, user agent, and headers.
- Browser actions such as clicks, scrolling, waits, JavaScript execution, screenshots, and file downloads.
- Authentication flows and session lifetime assumptions.
- Storage destinations, dataset deduplication, exports, and retention.
- Schedules, webhooks, downstream consumers, dashboards, and alerts.
- Per-domain concurrency limits and robots, legal, or compliance constraints.
Freeze a corpus of URLs and expected fields. Include successful pages, redirects, login pages, JavaScript-heavy pages, rate-limit responses, bot checks, empty results, very large documents, and known failures. This corpus becomes your migration test set.
3. Choose the replacement API
Compare providers on the execution model, rendering, extraction, anti-bot and proxy handling, operations, data storage, and effective cost.
Zyte API
Zyte API presents a single web-scraping API that can return HTTP content, browser HTML, screenshots, and structured extraction. Its documented capabilities include JavaScript execution, geolocation, sessions, browser actions, and automatic handling. It is a strong candidate when you want to remove proxy and browser infrastructure while retaining programmable extraction.
ScrapingBee
ScrapingBee provides an API with headless-browser rendering and proxy rotation and advertises 1,000 free API credits. Compare its fixed-credit plans, sessions, actions, extraction, geolocation, rate limits, and browser multipliers with your current request volume.
Bright Data Web Unlocker
Bright Data is relevant when your existing design is proxy-centric. A migration from Web Unlocker to an HTTP scraping API changes the endpoint, authentication, and parameter semantics, so validate geography, compliance, concurrency, and cost with production-like traffic before committing.
ScreenshotNeo for screenshot workloads
ScreenshotNeo is the first service to try when the workload is website screenshots or PDFs: it produces clean shots, bills only clean shots, and its paid plan starts at $5. It supports full-page and element captures, JavaScript and custom CSS, device presets, blocking, cookies and headers, waits, PDFs, caching, signed links, async jobs, bulk capture, and an MCP server for AI agents.
4. Design an adapter instead of rewriting your application
Keep your internal request and result types stable. Put provider-specific parameters in one adapter and normalize errors into a small set such as retryable, blocked, invalid_input, provider_error, and success.
type ScrapeRequest = {
url: string;
renderJs?: boolean;
country?: string;
sessionId?: string;
actions?: Array<{type: string; selector?: string; value?: string}>;
};
type ScrapeResult = {
url: string;
status: number;
html?: string;
data?: Record<string, unknown>;
provider: string;
billedUnits: number;
};
async function scrape(req: ScrapeRequest): Promise<ScrapeResult> {
// Provider-specific code belongs here. The rest of the application
// consumes ScrapeResult and does not depend on Apify or a vendor API.
throw new Error('Implement provider adapter');
}
During a dual-run period, send the same normalized request to Apify and the candidate API, then compare normalized fields. Store raw responses for debugging, but make downstream jobs consume only the normalized result.
5. Translate Actor features to HTTP parameters
| Actor behavior | Replacement design |
|---|---|
| Actor input schema | Validate a request object and map fields to query parameters or JSON. |
| Dataset items | Write each normalized result to your database or object store; preserve an idempotency key. |
| Request queue | Use a queue such as your existing job system and cap provider concurrency. |
| Browser actions | Map clicks, waits, scrolling, and JavaScript to documented provider actions; reject unsupported actions explicitly. |
| Proxy and geography | Map country, region, session, and proxy settings and verify behavior per domain. |
| Schedules | Move cron or event triggers to your scheduler and record the run identifier. |
| Webhooks | Use the provider callback if available, otherwise poll jobs and emit your own webhook. |
| Retries | Retry only transient network, 408, 429, and selected 5xx errors with exponential backoff and jitter. |
Pagination and idempotency
Do not assume an API’s pagination matches an Actor’s dataset order. Persist the provider cursor, page number, and last successful item. Give every request a deterministic key such as domain + canonical_url + extraction_version + date_partition. On retry, upsert by that key rather than appending blindly.
6. Runnable HTTP migration examples
cURL
curl --request POST 'https://api.example.com/v1/extract' \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"url": "https://example.com/products",
"render_js": true,
"country": "US",
"output": "html"
}'
Python
import os
import time
import requests
API_URL = "https://api.example.com/v1/extract"
headers = {
"Authorization": f"Bearer {os.environ['SCRAPER_API_KEY']}",
"Content-Type": "application/json",
}
payload = {
"url": "https://example.com/products",
"render_js": True,
"country": "US",
"output": "html",
}
for attempt in range(5):
response = requests.post(API_URL, json=payload, headers=headers, timeout=90)
if response.status_code not in (408, 429) and response.status_code < 500:
response.raise_for_status()
result = response.json()
print(result)
break
if attempt == 4:
response.raise_for_status()
time.sleep((2 ** attempt) + 0.25)
Node.js
const response = await fetch('https://api.example.com/v1/extract', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.SCRAPER_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
url: 'https://example.com/products',
render_js: true,
country: 'US',
output: 'html'
})
});
if (!response.ok) {
throw new Error(`Scraping API returned ${response.status}: ${await response.text()}`);
}
const result = await response.json();
console.log(result);
7. Or skip the browser setup
For screenshot and PDF jobs, ScreenshotNeo’s API replaces browser setup with one request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
8. Rebuild storage, schedules, and monitoring
A scraping endpoint may solve only fetching and extraction. Add the missing platform pieces deliberately:
- Queue: control concurrency per provider and per target domain.
- Scheduler: trigger jobs with a recorded schedule version and time zone.
- Persistence: store normalized fields, raw response location, request metadata, and schema version.
- Dead-letter queue: retain exhausted jobs with the final error and response headers.
- Observability: measure success rate, field completeness, latency, response size, retry count, block rate, and cost per successful item.
- Alerts: alert on changes in status codes, empty fields, latency, spend, and provider quota.
9. Performance, reliability, and cost
Performance
Separate queue wait time, provider processing time, transfer time, and parsing time. JavaScript rendering and browser actions generally cost more time than an HTTP fetch. Use the lightest mode that returns the fields you need, cap response size, avoid unnecessary screenshots, and reuse sessions only when the site permits it.
Reliability
Set a client timeout shorter than your worker lease, retry transient failures with jitter, and avoid retrying deterministic 4xx errors. Make writes idempotent. Keep Apify available for rollback until the replacement meets your success and completeness thresholds on the frozen corpus.
Cost
Calculate effective cost per successful record, not cost per request. Include browser multipliers, proxy or geography surcharges, storage, queue workers, retries, and failed requests. Compare fixed credits with pay-as-you-go pricing and check limits before launch. No universal migration-cost or success benchmark exists; measure your own URL corpus.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Wrong key, missing scheme, or account restriction | Check environment variables, authorization format, plan, and allowed domains. |
| 429 responses | Concurrency or rate limit exceeded | Use a bounded queue, honor retry headers, and add exponential backoff. |
| HTML is empty | Content is rendered after load or blocked by a consent wall | Enable JavaScript, add a selector or delay wait, and verify cookies or geography. |
| Fields differ from Apify | Different DOM timing, parser, or pagination semantics | Compare raw responses, normalize types, and version your extraction rules. |
| Sessions lose state | Session lifetime or cookie persistence differs | Use an explicit session identifier and verify that the provider preserves cookies. |
| Jobs disappear | No durable queue or result persistence | Persist job IDs and payloads before submission and use a dead-letter queue. |
| Costs spike | Browser mode, retries, large responses, or proxy geography | Measure cost per successful item, select the cheapest valid mode, and set spend alerts. |
11. Cutover checklist
- Freeze representative URLs and expected fields.
- Export every Actor schema, side effect, schedule, webhook, and storage dependency.
- Implement the provider adapter and preserve your internal schema.
- Recreate browser actions, sessions, geography, retries, and pagination.
- Provision queueing, persistence, scheduling, monitoring, and alerts.
- Run Apify and the candidate API against the same corpus.
- Compare success rate, field completeness, latency, ban rate, concurrency, and effective cost.
- Roll out by workload or domain, retain rollback, and recheck limits and pricing.
12. FAQ
Will an HTTP API replace Apify datasets?
Usually no. Store normalized results and raw responses in your own database or object storage unless the selected provider supplies equivalent datasets.
Can I keep my Actor code?
Only if the new provider supports the same runtime and browser controls. Most migrations move the extraction logic into an adapter or a separate worker.
Should I migrate all Actors at once?
No. Start with one stable workload, dual-run it, and expand after measuring completeness, reliability, and cost.
Is a proxy API the same as a scraping API?
No. A proxy API forwards traffic while a scraping API may render JavaScript, execute actions, and return extracted data. Authentication and parameters change between models.
When should I stay on Apify?
Stay when Actors, Store tools, persistent storage, schedules, integrations, or multi-step workflows are more valuable than a simpler HTTP integration.
