How to Run Web Scraping Actors Locally from Your Terminal
Run an Apify Actor from your terminal, provide local input, inspect storage, fix common errors, and deploy it when the scraper is ready.
To run an Apify Actor locally, install the Apify CLI, create or open an Actor project, place its input in storage/key_value_stores/default/INPUT.json, and run apify run from the project directory. The run writes datasets, key-value records, and request queues into the project’s storage directory.
What you need before you start
- A terminal and a working JavaScript/TypeScript or Python development environment matching your Actor template.
- The Apify CLI installed using Apify’s current installation instructions.
- An Actor project created with the CLI or an existing project containing its Actor metadata and source.
- A target website that you are allowed to access and scrape. Respect its terms, robots rules, privacy requirements, and rate limits.
An Actor is a program that accepts structured JSON input, performs a task such as crawling or browser automation, and can produce structured output. On Apify’s platform, Actors run in Docker containers; the project Dockerfile defines the image used for that runtime.
1. Install the Apify CLI and create a project
Install the CLI by following Apify’s current installation documentation for your operating system. Then create a project with the CLI or initialize an existing project:
# Create a new Actor project
apify create
# Or change into an existing Actor project
cd path/to/your-actor
The generated project normally contains an .actor directory, an actor.json file, input and output schemas, source code, a storage directory, a Dockerfile, and project metadata. The CLI provides JavaScript/TypeScript and Python templates in its quick-start workflow.
2. Inspect the project structure
your-actor/
├── .actor/
│ ├── actor.json
│ ├── input_schema.json
│ └── output_schema.json
├── src/ # or the template's source directory
├── storage/
│ ├── datasets/default/
│ ├── key_value_stores/default/
│ └── request_queues/default/
├── Dockerfile
└── package or Python project files
Names can vary by template, so treat the generated project as authoritative. The important local locations are:
| Location | Purpose |
|---|---|
storage/key_value_stores/default/INPUT.json |
The default input object for a local run. |
storage/datasets/default/ |
Dataset output, normally one JSON file per scraped item. |
storage/key_value_stores/default/ |
Named key-value records produced by the Actor. |
storage/request_queues/default/ |
Enqueued requests used by crawlers and queues. |
3. Provide Actor input locally
Open storage/key_value_stores/default/INPUT.json. Put in the JSON object expected by the Actor’s input schema. For example, an Actor whose schema defines startUrls and maxPages might use:
{
"startUrls": [
{"url": "https://example.com/"}
],
"maxPages": 10
}
Use the exact property names and value types declared by the input schema. If you add a field to the schema, update INPUT.json at the same time. A valid JSON document is required; comments and trailing commas will cause parsing errors.
4. Run the Actor from your terminal
cd path/to/your-actor
apify run
apify run starts the local development workflow. Watch the terminal output for startup messages, request progress, warnings, and the final result. Keep this terminal open until the process exits.
When the run finishes, inspect the output files:
find storage/datasets/default -type f -maxdepth 1 -print
find storage/key_value_stores/default -maxdepth 1 -type f -print
find storage/request_queues/default -maxdepth 1 -type f -print
Open the JSON files in the dataset directory to verify that each item has the expected fields. If the Actor writes a named key-value record, inspect that record in the key-value store directory.
5. Reset local state between runs
Local storage persists in the project directory. That is useful for continuing development, but stale datasets, queues, or input can make a new run look incorrect. Clear the default local storages before a clean run:
apify run --purge
Use this when you need to remove previous default datasets, key-value records, and request queues. Preserve any output you still need before purging it.
6. A practical local development loop
- Change the Actor source or its schema.
- Update
storage/key_value_stores/default/INPUT.jsonto match the schema. - Run
apify run. - Inspect terminal logs and files under
storage/. - Fix one failure at a time and rerun.
- Use
apify run --purgewhen old state could affect the result.
For repeatable debugging, keep a small input that exercises one page or one request. Increase the scope only after the smallest case produces correct output.
7. Local versus hosted execution
| Concern | Local run | Hosted run |
|---|---|---|
| Control | You control the terminal, files, runtime, and local environment. | Apify runs the Actor in its managed infrastructure and Docker-based runtime. |
| Persistence | Data is written to the project’s storage directory. |
Results are managed through the Apify platform. |
| Authentication | Not required merely to execute a local project. | Authenticate before deploying and managing the hosted Actor. |
| Deployment | No deployment; source runs on your machine. | Use apify push or a repository-based deployment workflow. |
| Scheduling and monitoring | You start and observe the process from the terminal. | Platform features can handle scheduling, monitoring, and managed execution. |
| Infrastructure | You maintain the local runtime, browser dependencies, network, and disk. | Apify supplies the execution environment defined by the project image. |
8. Deploy the Actor after local testing
First make sure a local run completes and its output is correct. Then authenticate the CLI:
apify login
From the Actor project directory, push the source to Apify:
apify push
For projects hosted in a source repository, use the repository deployment workflow described in Apify’s documentation. After deployment, review the hosted input and run configuration rather than assuming that local files automatically become hosted input. Keep secrets out of source code and local input files that you commit.
9. Capturing pages produced by an Actor
A scraper often needs HTML or structured data. If you also need a visual record of a page, you can run a browser locally, but browser setup adds dependencies, timing problems, consent banners, and cleanup work. A screenshot service can handle that separate concern after your Actor identifies the URL.
10. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API accepts one GET request and returns PNG, JPEG, WebP, or PDF output. Before capture, it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for the full option list and schemas. The basic call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper size and margins, HTML/CSS to image, custom CSS and JavaScript, clicks, selector or network-idle waits, ad and tracker blocking, custom headers and cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. Create a free ScreenshotNeo account to try it.
11. Troubleshooting local runs
| Symptom | Likely cause | Fix |
|---|---|---|
apify: command not found |
The CLI is not installed or is not on PATH. |
Install the current Apify CLI and restart the terminal so its executable path is loaded. |
| Input parsing fails before the Actor starts | INPUT.json is invalid JSON or uses the wrong field type. |
Validate the JSON, remove comments and trailing commas, and compare every field with the input schema. |
| The Actor starts with unexpected URLs | Old input or queue state remains in local storage. | Rewrite INPUT.json, then run apify run --purge. |
| No dataset files appear | The Actor did not push items, exited early, or wrote another storage record. | Read the terminal error, confirm the success path pushes dataset items, and inspect all three storage locations. |
| Requests fail locally but work elsewhere | Local DNS, proxy, credentials, browser dependencies, or site access differs from the hosted environment. | Check the local network and environment variables, verify credentials, and compare the project Dockerfile with the runtime you intend to deploy. |
| Browser pages are blank or incomplete | The page needs more time, JavaScript, or resources that are blocked locally. | Inspect browser logs, wait for the page’s real readiness condition, and verify that required resources are available. |
apify push is rejected |
You are not authenticated or the project metadata is incomplete. | Run apify login, check .actor/actor.json, and retry from the project directory. |
| Hosted output differs from local output | Environment, image, network, timing, or input differs between runs. | Pin configuration in the project, compare inputs, review the Dockerfile, and log the effective URL and key settings without exposing secrets. |
12. Performance, reliability, and cost considerations
Performance
- Start with a small input and a narrow request scope; scale page counts only after correctness is established.
- Reuse the project’s request queue and dataset mechanisms instead of retaining a large in-memory result set.
- Choose waits based on a page readiness condition rather than an unnecessarily long fixed delay.
- Keep local storage on a disk with enough capacity for datasets, logs, browser caches, and queued requests.
Reliability
- Make the input schema and
INPUT.jsonevolve together. - Make reruns safe: deduplicate output where appropriate and record enough source information to diagnose a bad item.
- Use
--purgefor deliberately clean experiments, but archive useful output first. - Expect websites to change. Selectors, consent flows, rate limits, and login requirements can invalidate an otherwise correct Actor.
Cost
A local run uses your own machine and infrastructure. Hosted execution adds the platform’s resource and operational model, so estimate page volume, browser time, storage, retries, and scheduling needs before moving a frequent job to hosted runs. For screenshots, ScreenshotNeo bills only clean shots; failed loads, bot checks, blank pages, timeouts, and cache hits are not billed. Its free tier includes 1,000 shots each month without a card.
13. Local-run checklist
- CLI installation works in a new terminal.
- The Actor project contains its
.actormetadata and source. INPUT.jsonis valid and matches the input schema.apify runcompletes for a small test input.- Dataset, key-value, and request-queue outputs are where you expect them.
- A clean run with
apify run --purgeproduces the same intended result. - Secrets are supplied securely and are not committed to the repository.
apify loginsucceeds before deployment.apify pushcompletes and the hosted input is reviewed.
FAQ
What command runs an Actor locally?
Run apify run from the Actor project directory.
Where does local Actor input go?
The default input record is storage/key_value_stores/default/INPUT.json.
Where are local results stored?
Dataset items are under storage/datasets/default/; key-value records and request queues are under their corresponding directories in storage/.
Does a local run require an Apify account?
The local command runs the project on your machine. An authenticated account is required when you deploy or manage the hosted Actor.
How do I clear data from a previous run?
Use apify run --purge, after saving any output you need.
How can an Actor produce screenshots without maintaining a browser?
Pass the resulting URL to ScreenshotNeo’s API or MCP server. It handles consent cleanup and reports whether a response was clean and billable.


