ScreenshotNeo

BlogEngineering

Estimating Scraping Volume and API Usage

Calculate requests, bandwidth, concurrency, retries, and API costs before your scraper hits a quota or budget limit.

By the ScreenshotNeo team1 October 20268 min read

Estimate scraping volume by counting every request-producing operation, measuring a representative sample, and checking each applicable limit separately. A useful first-pass model is:

Total requests = (detail pages + index pages + pagination calls + metadata calls + exports) × targets × runs + expected retries
Total bytes = total requests × average response bytes + headers + redirects + retry traffic + export traffic

The number that matters operationally is the first limit you exhaust. That may be requests per minute, requests per day, concurrent requests, tokens, points, rows, bandwidth, or billable results. There is no reliable universal “pages per day” benchmark because every service defines different windows, response sizes, retry rules, and billing units.

1. Define the workload before counting it

Write down the complete scope:

  • Domains, API hosts, or accounts.
  • URLs, resources, records, or search terms per target.
  • Refresh frequency and number of scheduled runs per day.
  • Whether each result needs detail, metadata, media, or an export.
  • Authentication and setup calls.
  • Pagination depth and cursor behavior.
  • Expected retries, redirects, and polling requests.

Classify calls by purpose. This prevents a common mistake: counting only detail pages while forgetting list pages, token refreshes, status polling, or dataset downloads.

Call class Typical examples Count separately?
Index/list Category pages, search results, API collections Yes
Detail One product, issue, profile, or article Yes
Pagination Next-page requests, cursor continuations Yes
Metadata Schema, permissions, asset manifests Yes
Authentication Login, token refresh, session setup Yes
Polling Async job status checks Yes
Export Dataset or archive downloads Yes
Retry Timeouts, 429s, transient 5xx responses Model separately

2. Calculate requests per run and per day

For one target, use a formula that includes each call class:

requests_per_target = index_calls + detail_calls + pagination_calls + metadata_calls + auth_calls + export_calls

Then scale it:

requests_per_run = requests_per_target × number_of_targets
requests_per_day = requests_per_run × scheduled_runs_per_day

For retries, estimate a retry rate from a sample rather than guessing. If 3% of first attempts need one additional attempt, use:

attempts = first_attempts × (1 + retry_rate × average_retry_attempts)

Example: 20,000 first attempts, a 3% retry rate, and 1.4 retry attempts on average produce 20,000 × (1 + 0.03 × 1.4) = 20,840 attempts.

Pagination example

If an index returns 100 records per page and a target contains 2,350 records, the run needs 24 index pages, not one request. If every record then needs a detail call, the target total is 2,374 calls before authentication, retries, and exports.

3. Measure a representative sample

Run a small sample that contains the same page mix as production. Record:

  • Requests by call class and target.
  • Response status and verdict.
  • Response bytes, including redirects and exports.
  • Latency, preferably p50 and p95.
  • Pagination depth.
  • Retry count and reason.
  • Maximum concurrent in-flight requests.
  • Rate-limit headers and Retry-After values.

Do not use only the easiest pages. A sample should include large responses, empty results, slow endpoints, authentication refreshes, and targets that commonly return errors.

Simple Python measurement script

import statistics
import time
import requests

URLS = [
    "https://example.com/page-a",
    "https://example.com/page-b",
]

samples = []
for url in URLS:
    started = time.perf_counter()
    response = requests.get(url, timeout=30)
    elapsed = time.perf_counter() - started
    samples.append({
        "url": url,
        "status": response.status_code,
        "bytes": len(response.content),
        "seconds": elapsed,
        "retry_after": response.headers.get("Retry-After"),
        "remaining": response.headers.get("X-RateLimit-Remaining"),
    })

print("requests", len(samples))
print("average_bytes", statistics.mean(s["bytes"] for s in samples))
print("p95_seconds", sorted(s["seconds"] for s in samples)[max(0, int(len(samples) * .95) - 1)])
print("samples", samples)

Group the measurements by endpoint or page type. A single overall average hides large pages and slow outliers.

4. Estimate bandwidth

Start with:

response_bytes = requests × average_response_bytes

Add traffic that is easy to miss:

  • Request and response headers.
  • Redirect responses and the final response.
  • Retries and backoff-related replays.
  • Authentication and token refresh calls.
  • Async polling responses.
  • Dataset exports and media downloads.

For a safer estimate, calculate each class independently:

total_bytes = Σ(request_count[class] × average_bytes[class])
monthly_bytes = total_bytes_per_run × runs_per_month

Measure compressed and uncompressed sizes according to what your provider bills. A JSON response compressed over the network may still count differently from a stored export.

5. Check every limit dimension

Limits are multidimensional. Check all of these, not just a daily request number:

  • Short windows: requests per second, 10 seconds, or minute.
  • Long windows: hourly or daily requests.
  • Concurrency: simultaneous in-flight requests.
  • Tokens or points: request cost based on payload or operation.
  • Rows or results: usage tied to returned records.
  • Bandwidth: downloaded or uploaded bytes.
  • Account scope: user, project, organization, key, or IP.
  • Billing: infrastructure time, credits, successful results, or exports.

OpenAI documents separate request and token limits, reset headers, Retry-After, backoff, and batching guidance; a request that exceeds a temporary limit returns HTTP 429. OpenAI rate-limit documentation

GitHub documents 60 requests per hour unauthenticated and 5,000 authenticated requests per hour, along with secondary limits that can include no more than 100 concurrent requests. GitHub rate-limit documentation

The UK Office for National Statistics documents 120 requests per 10 seconds, 200 requests per minute, and 15 requests per 10 seconds for high-demand assets. Exceeding those windows returns 429 and a Retry-After value. ONS developer guidance

api.data.gov documents a default limit of 1,000 requests per hour, while DEMO_KEY is limited to 30 requests per hour and 50 requests per day. Its responses expose X-RateLimit-Limit and X-RateLimit-Remaining. api.data.gov rate limits

Convert daily volume to a rough sustained rate

average_requests_per_second = daily_requests / 86400

This average is only a planning signal. A job that runs for 10 minutes creates a much larger burst than the same requests spread across 24 hours. Check the actual run duration, burst window, and concurrency separately.

6. Model concurrency and runtime

Concurrency affects both throughput and secondary limits. A rough runtime estimate is:

runtime_seconds ≈ total_requests × average_latency_seconds / concurrency

Use p95 latency when sizing a deadline and include backoff time. Increase concurrency gradually while watching 429s, timeouts, error rates, and provider guidance. A high concurrency value can reduce runtime while triggering a provider’s secondary protection.

Use a queue with a fixed maximum of in-flight requests. Keep separate limits per host or API key when the service scopes quotas that way.

7. Handle retries without runaway usage

Retry only errors that are likely to succeed later, such as temporary 429 and selected 5xx responses. Do not blindly retry authentication failures, invalid parameters, permanent 404s, or blocked requests.

  1. Honor Retry-After when present.
  2. Use exponential backoff with random jitter.
  3. Cap attempts and total retry time.
  4. Use an idempotency key or deduplication key where supported.
  5. Record the original error and every retry.
  6. Stop retrying when the job’s budget or deadline is exhausted.
delay = min(max_delay, base_delay * 2 ** attempt) + random_jitter

Retries count toward many providers’ request limits even when the final operation fails. Include them in volume and cost forecasts.

8. Async jobs, polling, and exports

Hosted scraping platforms can change the unit you estimate. Scrapy.io documents a run, poll, and dataset-export workflow, scheduled runs, and pay-per-result billing. That means you should count both control-plane requests, such as creating and polling a run, and the provider’s billable result unit. Scrapy.io documentation

Reduce polling traffic with webhooks when available. If polling is required, use increasing intervals and stop after a deadline. Count the export download separately from the run itself.

9. Compare self-hosted and hosted approaches

Dimension Self-hosted crawler Hosted scraping API
Request control You implement queues, limits, and backoff Provider exposes its own limits and controls
Browser and proxy operations You operate browsers, proxies, and upgrades Often supplied as part of the service
Retries You define retryable errors and budgets Provider may retry internally; verify billing behavior
Scheduling Your scheduler and workers May include scheduled runs
Output Your storage and schema May return rows, files, or datasets
Observability You collect all metrics Provider dashboards and headers may help
Billing unit Servers, bandwidth, proxies, and operations Credits, successful rows, results, or requests
Portability Maximum control, higher operating work Faster setup, provider-specific behavior

10. A reusable estimation worksheet

  1. Define scope: targets, URLs, records, and refreshes.
  2. List call classes: index, detail, pagination, metadata, auth, polling, exports.
  3. Measure: bytes, latency, statuses, retries, and concurrency for a representative sample.
  4. Calculate: requests per target, run, day, and month.
  5. Add retries: use observed retry rates and capped attempts.
  6. Convert windows: compare both sustained averages and bursts.
  7. Compare limits: requests, tokens, points, concurrency, bandwidth, and billing units.
  8. Add headroom: reserve capacity for growth, variance, and failures.
  9. Observe production: re-run the estimate using actual headers and outcomes.

11. Troubleshooting common estimation failures

Symptom Likely cause Fix
429 responses despite a safe daily total A short-window or concurrency limit was exceeded Throttle per window, reduce concurrency, and honor Retry-After
Usage is higher than the URL count Pagination, redirects, polling, auth refreshes, or retries were omitted Log every HTTP attempt and classify it
Bandwidth estimate is too low Large pages, exports, headers, or retries were excluded Measure bytes per call class and include all traffic
Costs rise while successful results stay flat Repeated retries or polling are billable Cap retries, use backoff, and prefer webhooks
Production runs exceed the deadline p95 latency and backoff were not included Size from p95 data and reserve runtime headroom
Authenticated and unauthenticated estimates differ Authentication changes quota scope or limit Estimate with the exact credential and endpoint used in production
Duplicate records appear after retries Operations are not idempotent or results are not deduplicated Use idempotency keys where supported and deduplicate by stable identifiers

12. Performance, reliability, and cost practices

  • Cache immutable or slowly changing pages and use conditional requests when supported.
  • Batch operations only when the provider documents that batching reduces request pressure without changing billing unexpectedly.
  • Spread scheduled work to avoid synchronized bursts.
  • Separate queues by host, credential, and priority.
  • Track request attempts, successful results, bytes, latency, 429s, retries, and estimated cost.
  • Set a per-run request budget and stop safely when it is reached.
  • Use a safety margin based on observed variance, not an arbitrary universal percentage.
  • Recalculate after pagination, schema, authentication, or provider-limit changes.

13. Or skip the browser setup

If your workload is website screenshots rather than structured extraction, ScreenshotNeo gives you one HTTP request per capture. It accepts a URL and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options, including full-page capture, CSS selectors, device presets, retina scale, PDF settings, custom CSS and JavaScript, waits, request blocking, cookies, headers, geolocation, caching, signed links, async webhooks, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Should I count pages or HTTP requests?

Count HTTP attempts. One page can require pagination, redirects, assets, polling, or retries, while one API request can return many records.

Do failed requests count?

Many providers count attempts even when they return errors. Check the service’s billing and quota documentation and include observed failures in your estimate.

How much headroom should I plan?

Use measured variance in volume, latency, retries, and response size. Recalculate after production data arrives instead of relying on a universal percentage.

What is the most important metric?

The first constrained dimension: a short rate window, concurrency, tokens, bytes, results, or budget. Monitor all of them because the bottleneck can change.