Cloud Scrapers: How to Scrape Websites at Scale
Build a cloud crawler that scales with bounded workers, respectful per-host pacing, reliable parsing, and browser rendering only where it is needed.
A cloud scraper scales by coordinating a controlled pipeline—not by sending as many requests as possible. Define the pages you are authorized to collect, fetch them with per-host concurrency and rate limits, render only pages that need JavaScript, validate extracted fields, and persist results with enough logs to diagnose failures. Increase worker count only when the target sites, error rates, and extraction checks support it.
This guide covers a practical architecture, runnable Python examples for HTTP and browser-based collection, cURL and Node.js request patterns, rate control, reliability, troubleshooting, and cost tradeoffs. Website access rules and applicable laws vary; the operating practices below are engineering guidance, not legal advice.
1. Design the crawler as a bounded pipeline
Separate discovery, fetching, rendering, parsing, validation, and storage. A queue or batch coordinator can hand work to a limited number of workers. Store raw responses or rendered HTML when appropriate, alongside normalized records, so you can investigate extraction changes without rerunning every request.
- Define scope. Start with an explicit URL list, sitemap, or controlled discovery process. Set crawl depth and breadth limits and exclude paths outside the intended dataset.
- Check crawl instructions. Read and respect the target site’s
robots.txt, identify your crawler in its user agent, and review the site’s terms and privacy information. Be prepared to stop if asked. - Schedule bounded work. Use a queue or batch coordinator, limit active work per host, and persist job state so workers can restart safely.
- Choose a fetch path. Use ordinary HTTP when the response contains the needed data. For JavaScript-dependent content, first consider whether the underlying data request can be called directly; use a browser when rendering or browser interaction is necessary.
- Parse and validate. Extract into a defined schema. Check required fields, types, and completeness before marking a record successful.
- Persist and monitor. Store records and useful diagnostics. Track fetch, rendering, parsing, and validation outcomes separately.
AWS documents one provider-specific architecture using AWS Batch to manage jobs, ECS containers to run crawlers, and S3 for collected files. The services are examples, not requirements: use equivalent queue, worker, and storage components that fit your workload and operating ownership. AWS also advises splitting large crawls into smaller batches; short-lived jobs may fit serverless functions, while long-running jobs may suit EC2 or ECS. AWS crawler architecture guidance.
Keep job state restartable
Represent each discovered URL as a durable work item with a stable identifier and state such as queued, fetched, parsed, validated, or failed. Make writes idempotent: retrying a worker should update the same record or output key rather than silently duplicating results. Record attempt count, timestamps, response status, final URL, and failure category. Keep sensitive headers and cookies out of ordinary logs.
2. Choose HTTP fetching or browser rendering
Use the least complex path that returns the data you need. An HTTP client and parser generally use fewer resources than a browser when the server response already includes the content. If JavaScript supplies the needed data, compare calling the underlying request with rendering the page. Reverse-engineering the request can take more development effort but fewer resources after it is built; browser automation can save that development work while consuming more resources and creating scaling constraints. Zyte’s scraping guide.
| Approach | Use when | Tradeoffs to plan for |
|---|---|---|
| HTTP client plus parser | The response HTML or endpoint contains the required content. | Low execution overhead; requires handling parsing and site-specific response behavior. |
| Direct data request | The page obtains data from a request you can reproduce for an approved collection purpose. | Can reduce browser work; requires investigation and can depend on session, headers, or request formats. |
| Browser automation or browser API | The required content appears only after scripts run, or the workflow needs browser actions. | Uses more resources, needs timeouts and wait conditions, and can be harder to scale. Confirm the service supports the required actions and output. |
Proxy rotation alone is not a complete strategy. Target-specific behavior may involve sessions, cookies, browser JavaScript, and HTTP protocol details. Technical capability does not establish permission to access a site or bypass its controls.
3. Build an HTTP crawler with bounded concurrency
This Python example fetches a fixed list of pages with a small global worker limit, a per-host delay, bounded retries for transient server errors, and basic extraction. It deliberately does not discover links or evade access controls. Adapt the parser and pacing to the target’s instructions and observed responses.
import asyncio
import json
import time
from collections import defaultdict
from urllib.parse import urlparse
import httpx
from bs4 import BeautifulSoup
URLS = [
"https://example.com/",
"https://example.com/about",
]
USER_AGENT = "ExampleResearchBot/1.0 (contact: crawler@example.org)"
GLOBAL_CONCURRENCY = 4
PER_HOST_DELAY_SECONDS = 10.0
MAX_ATTEMPTS = 3
host_locks = defaultdict(asyncio.Lock)
last_request_at = defaultdict(float)
def retry_delay(attempt):
return min(2 ** attempt, 30)
async def pace_for_host(host):
async with host_locks[host]:
elapsed = time.monotonic() - last_request_at[host]
if elapsed < PER_HOST_DELAY_SECONDS:
await asyncio.sleep(PER_HOST_DELAY_SECONDS - elapsed)
last_request_at[host] = time.monotonic()
async def fetch_one(client, semaphore, url):
host = urlparse(url).netloc
async with semaphore:
for attempt in range(MAX_ATTEMPTS):
await pace_for_host(host)
try:
response = await client.get(url)
if response.status_code == 429:
# Stop this host's work and investigate the site's limit.
return {"url": url, "status": 429, "error": "rate_limited"}
if response.status_code == 403:
# Repeated forbidden responses should trigger a stop/review.
return {"url": url, "status": 403, "error": "forbidden"}
if response.status_code in (500, 502, 503, 504):
if attempt + 1 < MAX_ATTEMPTS:
await asyncio.sleep(retry_delay(attempt))
continue
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else None
text = soup.get_text(" ", strip=True)
return {
"url": str(response.url),
"status": response.status_code,
"title": title,
"text_excerpt": text[:500],
}
except (httpx.TimeoutException, httpx.NetworkError) as exc:
if attempt + 1 == MAX_ATTEMPTS:
return {"url": url, "error": type(exc).__name__}
await asyncio.sleep(retry_delay(attempt))
except httpx.HTTPStatusError as exc:
return {"url": url, "status": exc.response.status_code,
"error": "http_status"}
async def main():
timeout = httpx.Timeout(20.0, connect=10.0)
limits = httpx.Limits(max_connections=GLOBAL_CONCURRENCY)
semaphore = asyncio.Semaphore(GLOBAL_CONCURRENCY)
async with httpx.AsyncClient(
timeout=timeout,
limits=limits,
headers={"User-Agent": USER_AGENT},
follow_redirects=True,
) as client:
results = await asyncio.gather(
*(fetch_one(client, semaphore, url) for url in URLS)
)
with open("results.jsonl", "w", encoding="utf-8") as output:
for row in results:
output.write(json.dumps(row, ensure_ascii=False) + "\n")
if __name__ == "__main__":
asyncio.run(main())
Install dependencies with python -m pip install httpx beautifulsoup4. The example’s delay is intentionally conservative and illustrative, not a universal threshold. Its per-host lock serializes requests to each host; for larger workloads, implement a shared per-host scheduler so independent worker processes honor the same limit.
cURL for a single diagnostic fetch
Use cURL to inspect a response manually before adding a target to a crawl. Replace the URL and user agent with values appropriate to your permitted use.
curl --fail-with-body --location --max-time 30 \
--user-agent "ExampleResearchBot/1.0 (contact: crawler@example.org)" \
--dump-header response-headers.txt \
--output response.html \
"https://example.com/"
This makes one request path easy to inspect; it is not a crawler coordinator. Check status, redirects, response headers, and whether the content you need is present in the returned HTML.
Node.js HTTP example
Node’s built-in fetch works for a small bounded list. This example is sequential, which is a useful starting point for a low-volume job; add a queue and per-host scheduler before expanding concurrency.
const urls = ["https://example.com/", "https://example.com/about"];
const userAgent = "ExampleResearchBot/1.0 (contact: crawler@example.org)";
for (const url of urls) {
try {
const response = await fetch(url, {
headers: { "User-Agent": userAgent },
signal: AbortSignal.timeout(20000),
redirect: "follow",
});
if (response.status === 429) {
console.error("Rate limited; pause this host:", url);
break;
}
if (response.status === 403) {
console.error("Forbidden; review and stop repeated requests:", url);
break;
}
if (!response.ok) {
console.error("HTTP error", response.status, url);
continue;
}
const html = await response.text();
const title = html.match(/<title[^>]*>([\s\S]*?)<\/title>/i)?.[1] ?? null;
console.log({ url: response.url, status: response.status, title });
} catch (error) {
console.error("Fetch failed", url, error.name);
}
}
The title extraction expression is intentionally minimal and is not a general HTML parser. Use a proper parser and schema validation for production extraction.
4. Render JavaScript pages only when the content requires it
When a page needs JavaScript execution, a browser worker should have explicit navigation and action timeouts, a wait condition tied to the content, and a resource budget. Avoid waiting indefinitely for all network activity: analytics or long-lived connections can prevent a page from becoming idle. Verify that the expected element exists before extracting.
For example, with Playwright for Python, install the package and browser once in the worker image using the Playwright installation instructions, then run a small, bounded browser job:
import asyncio
from playwright.async_api import async_playwright
async def capture_page(url):
async with async_playwright() as playwright:
browser = await playwright.chromium.launch(headless=True)
page = await browser.new_page()
page.set_default_navigation_timeout(30000)
page.set_default_timeout(10000)
await page.goto(url, wait_until="domcontentloaded")
await page.locator("main").wait_for(state="visible")
title = await page.title()
main_text = await page.locator("main").inner_text()
await browser.close()
return {"url": url, "title": title, "text": main_text[:2000]}
async def main():
result = await capture_page("https://example.com/")
print(result)
asyncio.run(main())
Replace main with a selector that represents the content you need, and handle missing elements as an extraction failure. Browser rendering can provide rendered DOM and browser actions, but services differ in supported interactions and outputs; verify those details in the relevant documentation. Zyte API reference.
Screenshot checks for extraction quality
For pages with visual layouts, a screenshot can help diagnose whether a selector or extracted field still corresponds to what a visitor sees. Treat it as a quality inspection artifact, not a substitute for schema validation. ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo; its screenshot and PDF options can be used as part of a visual inspection workflow.
5. Control rate, retries, and stop conditions
Set limits per host rather than relying only on a global worker count. The AWS ethical-crawler guidance gives examples of one request every 10–15 seconds for small or medium-sized sites and 1–2 requests per second for larger sites or sites with explicit crawl permission. Those are examples from AWS guidance, not universal thresholds or permission to crawl. AWS says to pause after HTTP 429 responses and consider stopping if HTTP 403 responses continue. AWS best practices for ethical web crawlers.
- Use a clear user agent that identifies the crawler and provides a contact route.
- Keep per-host concurrency and delay configurable, and honor crawl instructions.
- On 429, pause the affected host and review its stated or observed limits before resuming.
- On repeated 403 responses, stop and review access conditions rather than cycling identities or proxies.
- Retry only errors that may be transient; use a bounded attempt count and exponential backoff with jitter.
- Do not retry parsing or schema failures as though they were network errors.
- Use small batches to contain timeouts and make restart behavior understandable.
Scrapy’s AutoThrottle is one implementation of adaptive pacing: its documentation describes adjusting delays from response latency and target concurrency, and warns that a fixed short delay can inadvertently increase request rate when errors return faster. The cited material is for Scrapy 2.5.1, so confirm settings against the version deployed. Scrapy AutoThrottle documentation.
6. Make parsing and operations observable
Websites change, so successful HTTP responses do not guarantee successful extraction. Track field-level completeness and schema shape, and alert when those checks change. Separate fetch failure, render failure, parser failure, validation failure, and storage failure in logs and metrics.
Keep enough context to diagnose a failed record: requested and final URL, status, redirect chain where available, content type, duration, attempt number, parser version, and a short error category. Retain response bodies only when appropriate for the data and retention rules. Avoid logging secrets, authorization headers, or unnecessary personal data.
For JavaScript-heavy sites, event-driven navigation can hide links from crawlers that do not simulate the interaction. AWS’s Bedrock crawler troubleshooting suggests explicit seed URLs or a sitemap in such cases. AWS crawler troubleshooting.
Useful operational measures
- Queue health: queued, active, completed, and abandoned work items.
- Fetch health: status-code distribution, timeouts, redirects, and latency by host.
- Extraction health: records passing required-field checks and changes in field completeness.
- Resource use: worker memory, CPU, browser count, and storage volume.
- Retry behavior: attempts per URL and the share of work completed only after retries.
7. Scale the workload and estimate cost
Increase concurrency in measured steps and observe target responses, completion rate, extraction quality, and worker resources. More workers do not make an unresponsive or rate-limited site faster. If adding workers increases 429s, timeouts, or incomplete extraction, reduce pressure and investigate instead of scaling further.
Batching makes failures cheaper to reason about: a crashed worker affects a bounded portion of the crawl, and completed batches can be checked before the next begins. Size batches to fit job time limits, memory, restart behavior, and the delay policy. Use long-running containers or machines when job duration and control needs justify them; use short-lived or serverless execution for suitable smaller tasks.
Compare total workload cost, not just the request price or compute bill. Include browser compute, retries, storage, data transfer, orchestration, monitoring, maintenance time, and the work needed to repair parsers. For managed services, include the service’s pricing and limits for the features you actually use. The reviewed sources do not establish a neutral provider price or performance winner, so measure on representative targets you are permitted to crawl.
8. Troubleshooting common failures
| Symptom | Likely cause | Response |
|---|---|---|
| HTTP 429 | The target is limiting request frequency or concurrency. | Pause work for that host, reduce its rate, and review crawl instructions. Resume cautiously only when appropriate. |
| Repeated HTTP 403 | The target denies the request or access conditions are not met. | Stop repeated requests and review the site’s rules or contact route. Do not treat identity rotation as a fix. |
| HTTP 200 but empty fields | The parser no longer matches the page, content is loaded by JavaScript, or the response is an interstitial. | Inspect the response and extraction checks. Try a browser only if rendering is required and permitted; update the parser and add a regression fixture. |
| Navigation or worker timeout | Slow target, overloaded worker, unbounded wait condition, or page that never reaches network idle. | Set separate connection, navigation, and action timeouts; wait for a content-specific selector; reduce concurrency and inspect worker resources. |
| Many duplicate records | Redirects, URL variants, query parameters, or retries are being stored as distinct pages. | Define canonicalization rules for the dataset, retain the requested and final URL, and make writes idempotent. |
| Sudden extraction completeness drop | Site markup or response behavior changed, or a different page variant is being returned. | Alert on required fields, inspect representative responses, compare screenshots where useful, and version parser changes. |
| Browser workers run out of memory | Too many simultaneous browser contexts, oversized pages, or browser processes not closed after failures. | Limit contexts per worker, close browsers in a finally block, recycle workers, and measure memory by page class. |
| Retries make the crawl slower and noisier | Retrying permanent statuses or parser errors, or retry schedule is too aggressive. | Retry only transient failures, cap attempts, back off, and classify errors before deciding whether to retry. |
9. Or skip the browser setup
If the job is to capture visual evidence of pages rather than extract a large structured dataset, a screenshot API can remove browser setup from that part of the workflow. ScreenshotNeo accepts one GET request and returns a PNG, JPEG, WebP, or PDF. It can load lazy images for full-page captures, capture a CSS-selected element, use device presets or custom viewports, and wait for a selector, delay, or network idle. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server exposes screenshot, page-info, and PDF tools to AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
10. FAQ
Should I use a browser for every page?
No. Use HTTP when it provides the required content. Render only when JavaScript or browser interaction is necessary, because browser workers add resource and scaling costs.
Does a sitemap guarantee a complete crawl?
No. It is a useful source of focused URLs, but scope and completeness depend on the site and the collection goal. Validate discovered pages and extraction results.
Can I keep increasing concurrency if the crawler falls behind?
Only if target responses and extraction quality remain healthy. A growing queue can reflect target limits, slow pages, retries, or parser failures; diagnose the bottleneck before adding workers.
Is rotating proxies enough to make a crawler reliable?
No. Reliability also depends on pacing, sessions where appropriate, rendering needs, request behavior, parsing checks, and respecting access conditions.


