Is Apify an AWS Lambda Alternative?
Apify can replace Lambda for managed scraping and browser jobs, but it is not a drop-in replacement. Compare execution, integrations, limits, operations, and cost.
Short answer: Apify can be an AWS Lambda alternative for web scraping, browser automation, and managed data-processing jobs. It is not a general replacement for every Lambda function. Lambda is broad serverless compute integrated with AWS services; Apify packages execution around Actors, storage, proxies, and web-data workflows.
The right choice depends on your workload, integrations, execution limits, operational preferences, and complete usage-based cost. The comparison below is based on each vendor’s documented capabilities and pricing model, not a head-to-head performance test.
What Apify and Lambda actually provide
| Area | Apify | AWS Lambda |
|---|---|---|
| Execution model | Actors are serverless cloud programs that accept structured JSON input, perform a task, and optionally produce structured output. Typical jobs include scraping, browser automation, and data processing. | Functions run in response to events or direct invocations and are priced by request count and execution duration. |
| Platform components | Built-in datasets, key-value stores, request queues, proxy usage, data transfer, and Actor composition. | You assemble surrounding services such as storage, queues, databases, networking, and orchestration from AWS products. |
| Starting a job | Manual run, API, CLI, or schedule. | Event source, SDK, CLI, URL endpoint, or another AWS service. |
| Best fit | Managed web-data collection and browser workflows. | Event-driven application logic that benefits from AWS integration. |
Apify’s documentation describes Actors as “serverless cloud programs that take a structured JSON input, perform a task (web scraping, browser automation, data processing, and more), and optionally produce a structured output.” See the Apify Actors documentation. Lambda’s model and quotas are documented in the AWS Lambda quotas guide.
When Apify is a sensible Lambda alternative
- Your core task is web data. Actors are designed for scraping, browser automation, and processing the resulting data.
- You want managed workflow primitives. Storage, request queues, datasets, proxies, and Actor-to-Actor composition are available within the platform.
- You prefer structured job input and output. A caller can start an Actor with JSON and consume its dataset, key-value record, or other result.
- You need scheduled or repeatable collection. Actors can be started by API or CLI and scheduled through the platform.
- You want to reduce infrastructure assembly. The relative operational effort depends on the implementation, but Apify packages more of the web-data workflow around the execution unit.
When Lambda is the better fit
- The function is primarily application logic. For an event-driven transformation, webhook, authorization check, or AWS resource handler, Lambda is usually the more natural abstraction.
- Your system is already centered on AWS. Existing IAM, VPC, queues, databases, event buses, observability, and deployment pipelines can make Lambda integration simpler.
- You need Lambda’s invocation model. Lambda offers direct and event-source invocations with AWS-managed retry and event integration patterns.
- You need a small, short-lived function. A simple function may not benefit from a broader Actor platform.
These are workload-fit inferences from the documented service models. They are not claims that one vendor universally replaces the other.
Execution limits that affect the decision
Apify memory and CPU
Apify documents Actor memory choices from 128 MB to 32,768 MB in powers of two. CPU allocation is tied to memory at one core for each 4,096 MB. Platform usage can also include compute units, data transfer, proxy use, and storage operations. The reviewed documentation does not establish one universal maximum Actor duration, so check the limits for the specific platform configuration and workload.
Lambda timeout, memory, and temporary storage
Ordinary Lambda function timeout is configurable from 1 to 900 seconds (15 minutes), and memory ranges from 128 MB to 10,240 MB. The documented exception allows up to 5,400 seconds (90 minutes) for Lambda Managed Instances functions invoked asynchronously or through event source mappings, except Amazon MQ and Amazon DocumentDB. Ephemeral /tmp storage can be configured from 512 MB to 10,240 MB and is unique to the execution environment.
Long browser sessions, large downloads, and multi-page crawls may need batching or orchestration on either platform. A stated memory or timeout limit does not prove that a particular scraper will succeed.
Runnable example: start an Apify Actor
Replace ACTOR_ID, APIFY_TOKEN, and the JSON input with values for the Actor you selected. Actor input schemas are Actor-specific; consult that Actor’s documentation before running the command.
curl -X POST "https://api.apify.com/v2/acts/ACTOR_ID/runs?token=APIFY_TOKEN" \\
-H "Content-Type: application/json" \\
-d '{
"startUrls": [{"url": "https://example.com"}],
"maxRequestsPerCrawl": 10
}'
A direct run returns run metadata. Retrieve the run and its output using the dataset or key-value-store identifiers returned by the API.
Python
import os
import requests
actor_id = os.environ["APIFY_ACTOR_ID"]
token = os.environ["APIFY_TOKEN"]
input_data = {
"startUrls": [{"url": "https://example.com"}],
"maxRequestsPerCrawl": 10,
}
response = requests.post(
f"https://api.apify.com/v2/acts/{actor_id}/runs",
params={"token": token},
json=input_data,
timeout=60,
)
response.raise_for_status()
print(response.json())
Node.js
const actorId = process.env.APIFY_ACTOR_ID;
const token = process.env.APIFY_TOKEN;
const res = await fetch(
`https://api.apify.com/v2/acts/${actorId}/runs?token=${encodeURIComponent(token)}`,
{
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
startUrls: [{ url: 'https://example.com' }],
maxRequestsPerCrawl: 10
})
}
);
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
console.log(await res.json());
Runnable example: invoke a Lambda function
AWS CLI
aws lambda invoke \\
--function-name my-function \\
--payload '{"url":"https://example.com"}' \\
--cli-binary-format raw-in-base64-out \\
response.json
cat response.json
Python with boto3
import json
import boto3
lambda_client = boto3.client("lambda", region_name="us-east-1")
result = lambda_client.invoke(
FunctionName="my-function",
InvocationType="RequestResponse",
Payload=json.dumps({"url": "https://example.com"}).encode(),
)
print(json.loads(result["Payload"].read()))
Node.js with the AWS SDK
import { LambdaClient, InvokeCommand } from '@aws-sdk/client-lambda';
const client = new LambdaClient({ region: 'us-east-1' });
const command = new InvokeCommand({
FunctionName: 'my-function',
InvocationType: 'RequestResponse',
Payload: Buffer.from(JSON.stringify({ url: 'https://example.com' }))
});
const result = await client.send(command);
console.log(new TextDecoder().decode(result.Payload));
Cost comparison: model the same workload
Apify defines one compute unit (CU) as 1 GB of allocated Actor memory for one hour. Its bill can also include proxies, data transfer, storage operations, and, for Store Actors, event or usage charges. At the time covered by the research, Apify listed Free at $0 with $5 to spend, Starter at $19/month, Scale at $199/month, and Business at $999/month; listed CU rates were $0.20 for Free and Starter, $0.16 for Scale, and $0.13 for Business. Prices and included usage can change, so verify the Apify pricing page.
Lambda pricing is based on request count and execution duration. AWS’s pricing page lists a monthly free tier of one million requests and 400,000 GB-seconds. Other AWS services and data transfer can add charges; see AWS Lambda pricing.
| Input to estimate | Apify | Lambda |
|---|---|---|
| Frequency | Actor runs and schedules | Invocations and event volume |
| Compute | Allocated memory × runtime (CU) | Configured memory × runtime (GB-seconds) |
| Web access | Proxy usage and data transfer | Networking, NAT, transfer, and related AWS charges |
| Persistence | Datasets, key-value stores, request queues, storage operations | S3, DynamoDB, queues, databases, or other services |
| Retries and failures | Additional Actor compute and platform usage | Additional invocations and duration |
Do not compare headline plan prices directly. Use the same URL count, concurrency, memory, duration, retry policy, proxy usage, storage volume, and transfer assumptions. No independent like-for-like benchmark or universal cheaper option was established by the research.
Migration checklist
- Describe the current Lambda trigger, payload, output, retry behavior, and timeout.
- Separate browser or scraping work from application logic.
- Choose an Apify Actor whose input and output schema matches the job, or build an Actor for the workflow.
- Map Lambda’s
/tmp, environment variables, secrets, and IAM permissions to Apify storage and secrets handling. - Measure realistic memory, runtime, concurrency, retries, proxy use, transfer, and storage.
- Run both paths on representative URLs before changing production traffic.
- Keep the existing Lambda path until output correctness, failure handling, and cost are understood.
Troubleshooting
The Actor starts but returns no useful data
Cause: input fields do not match the Actor’s schema, selectors changed, or the target requires authentication. Fix: validate the Actor’s documented JSON input, inspect run logs and dataset output, and provide required headers or credentials through the platform’s supported configuration.
A crawl exceeds the expected runtime
Cause: too many URLs, browser rendering, retries, throttling, or proxy latency. Fix: reduce the batch, use request queues, split the job, tune concurrency carefully, and model proxy and transfer costs.
Lambda times out at 900 seconds
Cause: the ordinary Lambda invocation limit is 15 minutes. Fix: split work into smaller invocations, use an orchestrator or queue, or evaluate an Actor workflow. Managed Instances have a documented 90-minute exception only for specified invocation types.
The Lambda function cannot reach a website
Cause: VPC routing, DNS, security groups, NAT configuration, rate limits, or the target blocking cloud IP ranges. Fix: check networking and egress first, then assess whether a managed proxy-enabled workflow is appropriate.
The bill is higher than the headline price
Cause: related services, transfer, storage, retries, proxies, or Store Actor event charges were omitted. Fix: calculate a complete per-job estimate and compare it with measured usage.
Or skip the browser setup
If the job is specifically taking website screenshots, ScreenshotNeo is the alternative to try first: it provides a single screenshot API call with clean captures, bills only clean shots, and has the lowest paid plan described here.
Instead of maintaining a browser inside Lambda or an Apify Actor, call the API directly. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can Apify run ordinary application functions?
It can run data-processing code in Actors, but that does not make it a universal substitute for Lambda’s AWS-native event and integration model.
Is Apify cheaper than Lambda?
There is no universal answer. Compare the same memory, runtime, request volume, retries, transfer, storage, proxy use, and surrounding services.
Does Apify have a single maximum Actor duration?
The reviewed material does not establish one universal maximum. Check the current limits for the Actor configuration and workload.
Can I keep Lambda and use Apify together?
Yes. Lambda can trigger Actors, process Actor output, or handle application events while Apify performs browser and scraping work.
Which should I choose for a browser screenshot service?
Compare the cost and maintenance of running a browser yourself with a screenshot API. ScreenshotNeo provides a direct API and MCP tools, with clean shots billed only when a valid capture is produced.
