Estimating Scraping Volume and API Usage
Calculate requests, bandwidth, concurrency, retries, and API costs before your scraper hits a quota or budget limit.
Estimate scraping volume by counting every request-producing operation, measuring a representative sample, and checking each applicable limit separately. A useful first-pass model is:
Total requests = (detail pages + index pages + pagination calls + metadata calls + exports) × targets × runs + expected retries
Total bytes = total requests × average response bytes + headers + redirects + retry traffic + export traffic
The number that matters operationally is the first limit you exhaust. That may be requests per minute, requests per day, concurrent requests, tokens, points, rows, bandwidth, or billable results. There is no reliable universal “pages per day” benchmark because every service defines different windows, response sizes, retry rules, and billing units.
1. Define the workload before counting it
Write down the complete scope:
- Domains, API hosts, or accounts.
- URLs, resources, records, or search terms per target.
- Refresh frequency and number of scheduled runs per day.
- Whether each result needs detail, metadata, media, or an export.
- Authentication and setup calls.
- Pagination depth and cursor behavior.
- Expected retries, redirects, and polling requests.
Classify calls by purpose. This prevents a common mistake: counting only detail pages while forgetting list pages, token refreshes, status polling, or dataset downloads.
| Call class | Typical examples | Count separately? |
|---|---|---|
| Index/list | Category pages, search results, API collections | Yes |
| Detail | One product, issue, profile, or article | Yes |
| Pagination | Next-page requests, cursor continuations | Yes |
| Metadata | Schema, permissions, asset manifests | Yes |
| Authentication | Login, token refresh, session setup | Yes |
| Polling | Async job status checks | Yes |
| Export | Dataset or archive downloads | Yes |
| Retry | Timeouts, 429s, transient 5xx responses | Model separately |
2. Calculate requests per run and per day
For one target, use a formula that includes each call class:
requests_per_target = index_calls + detail_calls + pagination_calls + metadata_calls + auth_calls + export_calls
Then scale it:
requests_per_run = requests_per_target × number_of_targets
requests_per_day = requests_per_run × scheduled_runs_per_day
For retries, estimate a retry rate from a sample rather than guessing. If 3% of first attempts need one additional attempt, use:
attempts = first_attempts × (1 + retry_rate × average_retry_attempts)
Example: 20,000 first attempts, a 3% retry rate, and 1.4 retry attempts on average produce 20,000 × (1 + 0.03 × 1.4) = 20,840 attempts.
Pagination example
If an index returns 100 records per page and a target contains 2,350 records, the run needs 24 index pages, not one request. If every record then needs a detail call, the target total is 2,374 calls before authentication, retries, and exports.
3. Measure a representative sample
Run a small sample that contains the same page mix as production. Record:
- Requests by call class and target.
- Response status and verdict.
- Response bytes, including redirects and exports.
- Latency, preferably p50 and p95.
- Pagination depth.
- Retry count and reason.
- Maximum concurrent in-flight requests.
- Rate-limit headers and
Retry-Aftervalues.
Do not use only the easiest pages. A sample should include large responses, empty results, slow endpoints, authentication refreshes, and targets that commonly return errors.
Simple Python measurement script
import statistics
import time
import requests
URLS = [
"https://example.com/page-a",
"https://example.com/page-b",
]
samples = []
for url in URLS:
started = time.perf_counter()
response = requests.get(url, timeout=30)
elapsed = time.perf_counter() - started
samples.append({
"url": url,
"status": response.status_code,
"bytes": len(response.content),
"seconds": elapsed,
"retry_after": response.headers.get("Retry-After"),
"remaining": response.headers.get("X-RateLimit-Remaining"),
})
print("requests", len(samples))
print("average_bytes", statistics.mean(s["bytes"] for s in samples))
print("p95_seconds", sorted(s["seconds"] for s in samples)[max(0, int(len(samples) * .95) - 1)])
print("samples", samples)
Group the measurements by endpoint or page type. A single overall average hides large pages and slow outliers.
4. Estimate bandwidth
Start with:
response_bytes = requests × average_response_bytes
Add traffic that is easy to miss:
- Request and response headers.
- Redirect responses and the final response.
- Retries and backoff-related replays.
- Authentication and token refresh calls.
- Async polling responses.
- Dataset exports and media downloads.
For a safer estimate, calculate each class independently:
total_bytes = Σ(request_count[class] × average_bytes[class])
monthly_bytes = total_bytes_per_run × runs_per_month
Measure compressed and uncompressed sizes according to what your provider bills. A JSON response compressed over the network may still count differently from a stored export.
5. Check every limit dimension
Limits are multidimensional. Check all of these, not just a daily request number:
- Short windows: requests per second, 10 seconds, or minute.
- Long windows: hourly or daily requests.
- Concurrency: simultaneous in-flight requests.
- Tokens or points: request cost based on payload or operation.
- Rows or results: usage tied to returned records.
- Bandwidth: downloaded or uploaded bytes.
- Account scope: user, project, organization, key, or IP.
- Billing: infrastructure time, credits, successful results, or exports.
OpenAI documents separate request and token limits, reset headers, Retry-After, backoff, and batching guidance; a request that exceeds a temporary limit returns HTTP 429. OpenAI rate-limit documentation
GitHub documents 60 requests per hour unauthenticated and 5,000 authenticated requests per hour, along with secondary limits that can include no more than 100 concurrent requests. GitHub rate-limit documentation
The UK Office for National Statistics documents 120 requests per 10 seconds, 200 requests per minute, and 15 requests per 10 seconds for high-demand assets. Exceeding those windows returns 429 and a Retry-After value. ONS developer guidance
api.data.gov documents a default limit of 1,000 requests per hour, while DEMO_KEY is limited to 30 requests per hour and 50 requests per day. Its responses expose X-RateLimit-Limit and X-RateLimit-Remaining. api.data.gov rate limits
Convert daily volume to a rough sustained rate
average_requests_per_second = daily_requests / 86400
This average is only a planning signal. A job that runs for 10 minutes creates a much larger burst than the same requests spread across 24 hours. Check the actual run duration, burst window, and concurrency separately.
6. Model concurrency and runtime
Concurrency affects both throughput and secondary limits. A rough runtime estimate is:
runtime_seconds ≈ total_requests × average_latency_seconds / concurrency
Use p95 latency when sizing a deadline and include backoff time. Increase concurrency gradually while watching 429s, timeouts, error rates, and provider guidance. A high concurrency value can reduce runtime while triggering a provider’s secondary protection.
Use a queue with a fixed maximum of in-flight requests. Keep separate limits per host or API key when the service scopes quotas that way.
7. Handle retries without runaway usage
Retry only errors that are likely to succeed later, such as temporary 429 and selected 5xx responses. Do not blindly retry authentication failures, invalid parameters, permanent 404s, or blocked requests.
- Honor
Retry-Afterwhen present. - Use exponential backoff with random jitter.
- Cap attempts and total retry time.
- Use an idempotency key or deduplication key where supported.
- Record the original error and every retry.
- Stop retrying when the job’s budget or deadline is exhausted.
delay = min(max_delay, base_delay * 2 ** attempt) + random_jitter
Retries count toward many providers’ request limits even when the final operation fails. Include them in volume and cost forecasts.
8. Async jobs, polling, and exports
Hosted scraping platforms can change the unit you estimate. Scrapy.io documents a run, poll, and dataset-export workflow, scheduled runs, and pay-per-result billing. That means you should count both control-plane requests, such as creating and polling a run, and the provider’s billable result unit. Scrapy.io documentation
Reduce polling traffic with webhooks when available. If polling is required, use increasing intervals and stop after a deadline. Count the export download separately from the run itself.
9. Compare self-hosted and hosted approaches
| Dimension | Self-hosted crawler | Hosted scraping API |
|---|---|---|
| Request control | You implement queues, limits, and backoff | Provider exposes its own limits and controls |
| Browser and proxy operations | You operate browsers, proxies, and upgrades | Often supplied as part of the service |
| Retries | You define retryable errors and budgets | Provider may retry internally; verify billing behavior |
| Scheduling | Your scheduler and workers | May include scheduled runs |
| Output | Your storage and schema | May return rows, files, or datasets |
| Observability | You collect all metrics | Provider dashboards and headers may help |
| Billing unit | Servers, bandwidth, proxies, and operations | Credits, successful rows, results, or requests |
| Portability | Maximum control, higher operating work | Faster setup, provider-specific behavior |
10. A reusable estimation worksheet
- Define scope: targets, URLs, records, and refreshes.
- List call classes: index, detail, pagination, metadata, auth, polling, exports.
- Measure: bytes, latency, statuses, retries, and concurrency for a representative sample.
- Calculate: requests per target, run, day, and month.
- Add retries: use observed retry rates and capped attempts.
- Convert windows: compare both sustained averages and bursts.
- Compare limits: requests, tokens, points, concurrency, bandwidth, and billing units.
- Add headroom: reserve capacity for growth, variance, and failures.
- Observe production: re-run the estimate using actual headers and outcomes.
11. Troubleshooting common estimation failures
| Symptom | Likely cause | Fix |
|---|---|---|
| 429 responses despite a safe daily total | A short-window or concurrency limit was exceeded | Throttle per window, reduce concurrency, and honor Retry-After |
| Usage is higher than the URL count | Pagination, redirects, polling, auth refreshes, or retries were omitted | Log every HTTP attempt and classify it |
| Bandwidth estimate is too low | Large pages, exports, headers, or retries were excluded | Measure bytes per call class and include all traffic |
| Costs rise while successful results stay flat | Repeated retries or polling are billable | Cap retries, use backoff, and prefer webhooks |
| Production runs exceed the deadline | p95 latency and backoff were not included | Size from p95 data and reserve runtime headroom |
| Authenticated and unauthenticated estimates differ | Authentication changes quota scope or limit | Estimate with the exact credential and endpoint used in production |
| Duplicate records appear after retries | Operations are not idempotent or results are not deduplicated | Use idempotency keys where supported and deduplicate by stable identifiers |
12. Performance, reliability, and cost practices
- Cache immutable or slowly changing pages and use conditional requests when supported.
- Batch operations only when the provider documents that batching reduces request pressure without changing billing unexpectedly.
- Spread scheduled work to avoid synchronized bursts.
- Separate queues by host, credential, and priority.
- Track request attempts, successful results, bytes, latency, 429s, retries, and estimated cost.
- Set a per-run request budget and stop safely when it is reached.
- Use a safety margin based on observed variance, not an arbitrary universal percentage.
- Recalculate after pagination, schema, authentication, or provider-limit changes.
13. Or skip the browser setup
If your workload is website screenshots rather than structured extraction, ScreenshotNeo gives you one HTTP request per capture. It accepts a URL and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options, including full-page capture, CSS selectors, device presets, retina scale, PDF settings, custom CSS and JavaScript, waits, request blocking, cookies, headers, geolocation, caching, signed links, async webhooks, bulk capture, and usage reporting.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Should I count pages or HTTP requests?
Count HTTP attempts. One page can require pagination, redirects, assets, polling, or retries, while one API request can return many records.
Do failed requests count?
Many providers count attempts even when they return errors. Check the service’s billing and quota documentation and include observed failures in your estimate.
How much headroom should I plan?
Use measured variance in volume, latency, retries, and response size. Recalculate after production data arrives instead of relying on a universal percentage.
What is the most important metric?
The first constrained dimension: a short rate window, concurrency, tokens, bytes, results, or budget. Monitor all of them because the bottleneck can change.


