ScreenshotNeo

BlogHow-to

How to handle screenshot API rate limits in a bulk capture job

Build a resumable screenshot capture job that respects rate limits, retries temporary throttling safely, and stops when an account quota is exhausted.

By the ScreenshotNeo team4 October 202610 min read

Handle screenshot API rate limits with a shared queue, a provider-aware request-rate limiter, and a separate cap on concurrent captures. On a temporary 429, honor a valid Retry-After value, add jitter, and retry only within bounded attempt and time limits. If the response indicates monthly quota exhaustion, pause the job and check usage or reset details; waiting a few seconds will not restore quota.

Make the job resumable: persist each target’s state and result, and reconcile requests that timed out before replaying them. Exact rate windows, concurrency caps, batch semantics, headers, and quota rules vary by provider and plan, so read the chosen API’s current documentation before configuring the limits.

How do I handle screenshot API rate limits in a bulk capture job?

  1. Read the provider contract. Record the request limit and window, concurrent-render cap, batch-size maximum, monthly quota and reset, error codes, relevant headers, and retry behavior. Check the account dashboard or usage endpoint if available.
  2. Queue targets durably. Store one item per URL, or store provider batch and job IDs when using a documented batch interface. Persist completion state and results so a restarted worker can skip finished captures.
  3. Control dispatch rate and concurrency separately. A request-per-minute limit and a concurrent-render limit are different constraints. Use a shared limiter across workers that use the same account or key; independent per-worker limits can exceed the account-wide allowance.
  4. Retry temporary throttling carefully. For a recognized temporary rate-limit response, parse Retry-After as specified by the provider, wait at least that long, then add random jitter. If it is absent or invalid, use capped exponential backoff with jitter. Reduce dispatch rate or concurrency after repeated throttling.
  5. Stop on quota and permanent errors. Inspect the error body and headers: some APIs use HTTP 429 for both temporary throttling and monthly quota exhaustion. Do not retry invalid requests, authentication failures, or exhausted quotas indefinitely.
  6. Reconcile uncertain outcomes. A client timeout does not prove the server failed to capture the page. Check the provider’s job/status lookup or documented idempotency mechanism before replaying an uncertain request.
  7. Monitor the run. Track queue depth, in-flight work, completed and failed items, retries, request IDs, error types, and provider-supplied rate or quota remaining/reset values. Never log API secrets.

Rate limit, concurrency, and monthly quota are different

Control What it limits How the job should respond
Request rate Requests over a provider-defined time window, such as per second or per minute. Pace dispatch through a shared limiter. Treat a temporary 429 as a signal to slow down.
Concurrency Captures or renders running at the same time. Keep a separate semaphore or in-flight cap if the provider documents one.
Monthly quota Captures allowed in an account period, with a provider-defined reset. Pause and inspect usage, reset time, or plan capacity. Immediate retries do not replenish it.
Batch limit Items accepted in one batch request, and possibly separate per-item accounting. Stay within the documented batch maximum and confirm how each item counts toward rate and quota limits.

Do not assume a batch endpoint bypasses rate limits or quota accounting. Providers publish different limit models and plan values, and these can change. For example, documented services distinguish request windows from monthly successful captures or renders; their example plan numbers are provider-specific and should not be copied into another API’s configuration. Check your selected provider’s current documentation and account settings.

Build a bounded retry policy

Retry only failures that may clear with time. Parse machine-readable error details as well as the HTTP status; a status alone may not distinguish a short burst limit from account quota exhaustion. If the provider documents Retry-After, support its documented format and treat the indicated wait as a minimum. Add jitter so workers do not all resume together.

When there is no valid retry delay, use exponential backoff with a cap. Bound both the number of attempts and total elapsed retry time. Account for SDK retries too: nested SDK and application retries can multiply requests. The OpenAI API guide specifically warns that unsuccessful requests count toward its per-minute limit, so continuously resending will not work; use that as a general caution and verify the screenshot provider’s own policy. OpenAI rate-limit guidance.

delay = min(max_delay, base_delay * 2 ** retry_number)
wait = max(valid_retry_after, delay) + random_jitter

This is policy pseudocode, not a universal set of constants. Choose base delay, cap, jitter range, maximum attempts, and overall deadline for your workload and provider contract. Never let one retry loop run forever.

Runnable Python worker pattern

The example below shows a minimal paced worker with bounded retries for an API that returns image bytes and reports temporary throttling with HTTP 429. Replace the endpoint, authentication, error-body parsing, response handling, and limits with the selected provider’s documented contract. It intentionally pauses on any 429 because a generic example cannot safely distinguish temporary throttling from quota exhaustion; add provider-specific parsing before enabling automated 429 retries in production.

import random
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from urllib.parse import urlparse

import requests

API_URL = "https://api.example.com/screenshot"
API_KEY = "YOUR_API_KEY"
TARGETS = ["https://example.com", "https://www.python.org"]
OUTPUT_DIR = Path("screenshots")
OUTPUT_DIR.mkdir(exist_ok=True)

# These are example job settings, not provider limits.
MAX_WORKERS = 3
MIN_SECONDS_BETWEEN_STARTS = 0.5
MAX_ATTEMPTS = 5
MAX_RETRY_SECONDS = 120
BASE_BACKOFF_SECONDS = 1.0
MAX_BACKOFF_SECONDS = 30.0

next_start = 0.0

def pace_dispatch():
    """Shared pacing for threads in this process."""
    global next_start
    now = time.monotonic()
    start_at = max(now, next_start)
    next_start = start_at + MIN_SECONDS_BETWEEN_STARTS
    delay = start_at - now
    if delay > 0:
        time.sleep(delay)

def output_name(url):
    host = urlparse(url).netloc.replace(":", "_") or "page"
    return OUTPUT_DIR / f"{host}.png"

def capture(url):
    started = time.monotonic()
    for attempt in range(MAX_ATTEMPTS):
        pace_dispatch()
        try:
            response = requests.get(
                API_URL,
                params={"key": API_KEY, "url": url},
                timeout=(10, 90),
            )
        except requests.Timeout as exc:
            # The provider may have completed the capture. Reconcile via its
            # status/idempotency mechanism before replaying in production.
            return {"url": url, "state": "uncertain", "error": str(exc)}
        except requests.RequestException as exc:
            return {"url": url, "state": "failed", "error": str(exc)}

        if response.status_code == 429:
            # Inspect provider-specific body/headers here. Quota exhaustion
            # must stop the job rather than enter this transient retry path.
            return {
                "url": url,
                "state": "throttled_or_quota",
                "retry_after": response.headers.get("Retry-After"),
                "error": response.text[:500],
            }

        if response.status_code in (401, 403, 400, 404):
            return {"url": url, "state": "failed", "status": response.status_code,
                    "error": response.text[:500]}
        if not response.ok:
            return {"url": url, "state": "failed", "status": response.status_code,
                    "error": response.text[:500]}

        path = output_name(url)
        path.write_bytes(response.content)
        return {"url": url, "state": "complete", "path": str(path)}

    elapsed = time.monotonic() - started
    return {"url": url, "state": "retry_limit", "elapsed_seconds": elapsed}

def main():
    results = []
    with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
        futures = [pool.submit(capture, url) for url in TARGETS]
        for future in as_completed(futures):
            result = future.result()
            results.append(result)
            print(result)
            if result["state"] == "throttled_or_quota":
                # Production orchestration should stop dispatch globally and
                # classify the provider error before deciding to resume.
                print("Pause the job and inspect provider error details/usage.")
    return results

if __name__ == "__main__":
    main()

For production, persist each result as it completes rather than keeping status only in memory. Use an atomic write or temporary file plus rename so a process crash cannot leave a partial image marked complete. The simple pacing variable above is shared only within this process; distributed workers need a shared queue, datastore, or rate-control service.

cURL, Python, and Node.js request shapes

These are single-request examples for exercising a provider’s capture endpoint. They do not implement a bulk scheduler or define a rate limit. Replace URL, endpoint, key parameter, output format, and timeout behavior according to the provider documentation.

cURL

curl --fail-with-body --get "https://api.example.com/screenshot" \
  --data-urlencode "key=YOUR_API_KEY" \
  --data-urlencode "url=https://example.com" \
  --output screenshot.png

Python

import requests

response = requests.get(
    "https://api.example.com/screenshot",
    params={"key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=(10, 90),
)
response.raise_for_status()
with open("screenshot.png", "wb") as image_file:
    image_file.write(response.content)

Node.js

const endpoint = new URL('https://api.example.com/screenshot');
endpoint.search = new URLSearchParams({
  key: 'YOUR_API_KEY',
  url: 'https://example.com',
});

const response = await fetch(endpoint, { signal: AbortSignal.timeout(90000) });
if (!response.ok) {
  throw new Error(`Screenshot request failed: HTTP ${response.status}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('screenshot.png', bytes));

Batch jobs and resumability

If your provider offers a batch or asynchronous job endpoint, it may simplify submission and progress tracking. Confirm whether the batch submission, each captured URL, or both count against request-rate and monthly limits. Record the returned batch/job identifier and use the documented status endpoint or event mechanism to follow progress. Preserve item-level status so a restart can resume unfinished work without resubmitting completed captures.

For synchronous calls, store a durable record keyed by a stable job item identifier. Useful fields include target URL, requested options, state, attempt count, last HTTP status and error, provider request or job ID, output location, and timestamps. Avoid using the URL alone if the same URL can be requested with different capture options.

Performance, reliability, and cost

  • Raise throughput gradually. Begin below the provider’s documented rate and concurrency ceilings. Increase in controlled steps while watching throttling and queue time. A higher worker count does not help when a shared account rate limit is already saturated.
  • Use backpressure. Stop or slow queue consumption when the provider throttles, the output store is slow, or too many jobs are in flight. This prevents unbounded memory use and retry buildup.
  • Keep timeouts realistic. Browser rendering can take longer than a normal HTTP response. Choose connect and response timeouts based on the provider’s documented behavior, but treat a response timeout as an uncertain outcome until reconciled.
  • Account for quota semantics. Providers may count successful captures, accepted jobs, or other units. Confirm whether failures, cache hits, and batch items consume allowance before estimating run cost.
  • Budget retries and polling. Status checks may themselves be rate-limited. Use the documented polling interval or event mechanism, and include retries and status calls in request-rate planning.
  • Keep secrets out of logs. Do not log authorization headers, API keys, or signed capture URLs. Retain provider request IDs and sanitized error details for support and diagnosis.

Common errors and fixes

Symptom Likely cause Fix
Repeated 429 responses Dispatch rate exceeds the account’s window, workers synchronize retries, or failed requests still consume request capacity. Use a shared limiter, honor Retry-After, add jitter, lower concurrency, and bound retries.
429 continues after waiting The response may indicate exhausted monthly quota rather than temporary throttling. Inspect the machine-readable body and quota headers or usage page. Pause until capacity or the documented reset is available.
Workers overwhelm the limit despite local pacing Each worker or machine is enforcing its own allowance instead of sharing the account’s total limit. Coordinate dispatch through one shared controller or divide the documented account allowance across workers.
Same items appear more than once after restart Completion state was held only in memory, or a timed-out result was replayed without reconciliation. Persist item state and output references; check provider job/status or idempotency support before replay.
Retries multiply unexpectedly Both the SDK and application have retry loops. Inspect SDK defaults and versions. Use one clearly bounded retry layer or account for combined attempts.
Batch submission is rejected or still throttled The batch exceeds the maximum, or batch items remain subject to per-item limits and quota. Reduce batch size to the documented maximum and verify batch accounting rules.
401 or 403 Invalid key, missing scope, or account permission issue. Correct authentication or permissions; do not retry unchanged credentials.
Capture times out but quota decreases The server may have completed the render after the client stopped waiting. Reconcile through a status lookup or logs before resubmitting.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its one-call API is useful when you want the capture service to handle browser setup; it also offers bulk capture for up to 100 URLs per call. See the ScreenshotNeo API documentation for request options and integration details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before the shot; known consent platforms, newsletter popups, and chat widgets can also be removed.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; responses report page verdict and billing status in headers.
  • An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.

FAQ

Should every HTTP 429 be retried?

No. First determine whether it means temporary throttling or exhausted account quota. Retry only the temporary case under a bounded policy.

Can I use one limiter for multiple API keys?

Only if the provider documents that those keys share a limit. Scope the limiter to the provider’s actual rate-limit unit, such as key, account, or organization.

How do I choose worker concurrency?

Start below the documented cap, then adjust based on observed throughput and throttling. Keep concurrency control separate from request pacing.

Does a timeout mean the screenshot was not created?

No. The server may have finished after the client timed out. Check the provider’s status or idempotency support before submitting the same work again.