How to Handle Screenshot API Rate Limit Errors
Learn how to classify screenshot API 429 errors, honor Retry-After, build bounded retries, and prevent throttling, quota, and billing failures.
Handle a screenshot API 429 by classifying it before retrying. A 429 may mean temporary request throttling, an exhausted monthly screenshot quota, or another usage or billing limit. Read the response body and headers first. For temporary throttling, wait at least the duration in Retry-After. If that header is missing or invalid, use capped exponential backoff with random jitter, a retry limit, and a total deadline. Do not automatically retry monthly quota, billing, authentication, or invalid-input errors.
What a 429 means
HTTP 429 is a status code, not a complete diagnosis. Screenshot providers commonly enforce more than one limit:
| Condition | Typical signal | Action |
|---|---|---|
| Temporary request throttling | 429 with a rate-limit message, remaining/reset headers, or Retry-After |
Queue the job, reduce concurrency, and retry after the server’s delay. |
| Monthly screenshot quota | 429 or a provider-specific error such as quota_exceeded |
Stop automatic retries. Check usage, upgrade, or wait for the billing-period reset. |
| Billing or organization cap | Usage, payment, or account-limit error | Fix the account or billing limit before sending more requests. |
| Invalid request or authentication | 4xx error identifying parameters or credentials | Correct the request or key. Retrying does not fix it. |
| Transient renderer or service failure | 500, 502, or 503, sometimes with retry guidance | Retry a small bounded number of times, then defer or surface the failure. |
A provider can apply both a burst or requests-per-minute limit and a monthly successful-render allowance. A request-rate limit controls how quickly you send work; a monthly quota controls how much successful capture volume your plan includes. They require different recovery actions.
Step 1: Capture diagnostics safely
Log enough information to classify the failure and investigate it later, but never log API keys or cookies.
- HTTP status and content type
- Machine-readable error code and response body
Retry-After,RateLimit-Remaining,RateLimit-Reset, and provider-specific quota headers- Provider request or correlation ID
- Endpoint, target URL (redacted if it contains secrets), attempt number, and UTC timestamp
- Client timeout and total elapsed time
Successful captures are often binary PNG, JPEG, WebP, or PDF responses, while errors are JSON or text. Branch on status and content type before trying to decode an image.
Step 2: Read Retry-After correctly
Retry-After is the first pacing signal. It can be a number of seconds or an HTTP date. Treat it as a minimum delay. OpenAI’s rate-limit guidance defines it as the minimum number of seconds before retrying a temporary rate-limit error when present (OpenAI rate limits). Apple gives the same order: wait for Retry-After, fall back to RateLimit-Reset, then use a default delay (Apple rate-limit guidance).
If the value is absent, malformed, negative, or impossibly large, use a bounded fallback. A common schedule is:
delay = min(max_delay, base_delay * 2 ** attempt) + random_jitter
Use a maximum delay and a total retry deadline. If Retry-After exceeds your deadline, persist the job for later instead of retrying early.
Step 3: Implement bounded retries
Provider-neutral pseudocode
for attempt in 0..max_retries:
response = capture()
if response.ok:
return response
error = parse_error(response)
if response.status == 429 and error.code indicates monthly_quota:
stop_and_surface_quota_action()
if response.status == 429 or response.status == 503:
delay = valid_retry_after(response)
or exponential_delay(attempt) + random_jitter()
if deadline_exceeded(delay):
defer_job()
sleep(delay)
continue
return classify_non_retryable_error(response)
Retry only errors that are plausibly temporary. Every retry consumes time, and unsuccessful requests can still contribute to request-rate limits. An uncontrolled loop can therefore prolong throttling.
Runnable Python example
This example treats a JSON error code containing quota as non-retryable, honors both seconds and HTTP-date forms of Retry-After, and caps attempts and total time.
import json
import random
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import requests
API_URL = "https://example.invalid/v1/shot" # Replace with your provider endpoint
MAX_RETRIES = 4
MAX_TOTAL_SECONDS = 90
BASE_DELAY = 1.0
MAX_DELAY = 30.0
def retry_after_seconds(value):
if not value:
return None
try:
seconds = float(value)
return max(0.0, seconds)
except ValueError:
try:
target = parsedate_to_datetime(value)
if target.tzinfo is None:
target = target.replace(tzinfo=timezone.utc)
return max(0.0, (target - datetime.now(timezone.utc)).total_seconds())
except (TypeError, ValueError, OverflowError):
return None
def parse_error(response):
try:
data = response.json()
if isinstance(data, dict):
return str(data.get("code", "")), str(data.get("message", ""))
except (ValueError, json.JSONDecodeError):
pass
return "", response.text[:500]
def capture():
return requests.get(
API_URL,
params={"url": "https://example.com"},
timeout=(10, 60),
)
def run():
started = time.monotonic()
for attempt in range(MAX_RETRIES + 1):
response = capture()
if response.ok:
content_type = response.headers.get("content-type", "")
if "image" in content_type or "pdf" in content_type:
with open("shot.bin", "wb") as output:
output.write(response.content)
return
raise RuntimeError(f"Unexpected success content type: {content_type}")
code, message = parse_error(response)
lowered = f"{code} {message}".lower()
print({
"status": response.status_code,
"code": code,
"retry_after": response.headers.get("Retry-After"),
"request_id": response.headers.get("X-Request-Id"),
"attempt": attempt,
})
if response.status_code == 429 and any(word in lowered for word in ("quota", "billing", "monthly")):
raise RuntimeError(f"Non-retryable usage limit: {message}")
retryable = response.status_code == 429 or response.status_code == 503
if not retryable or attempt == MAX_RETRIES:
raise RuntimeError(f"Capture failed ({response.status_code}): {message}")
delay = retry_after_seconds(response.headers.get("Retry-After"))
if delay is None:
delay = min(MAX_DELAY, BASE_DELAY * (2 ** attempt)) + random.uniform(0, 0.5)
if time.monotonic() - started + delay > MAX_TOTAL_SECONDS:
raise TimeoutError("Retry deadline exceeded; defer this job")
time.sleep(delay)
if __name__ == "__main__":
run()
Runnable Node.js example
const API_URL = 'https://example.invalid/v1/shot'; // Replace with your provider endpoint
const maxRetries = 4;
const maxTotalMs = 90_000;
const baseDelayMs = 1_000;
const maxDelayMs = 30_000;
const sleep = (ms) => new Promise(resolve => setTimeout(resolve, ms));
function retryAfterMs(value) {
if (!value) return null;
const seconds = Number(value);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1000);
const timestamp = Date.parse(value);
if (!Number.isNaN(timestamp)) return Math.max(0, timestamp - Date.now());
return null;
}
async function capture() {
const url = new URL(API_URL);
url.searchParams.set('url', 'https://example.com');
return fetch(url, { signal: AbortSignal.timeout(70_000) });
}
const started = Date.now();
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const response = await capture();
if (response.ok) {
const type = response.headers.get('content-type') || '';
if (!type.includes('image') && !type.includes('pdf')) {
throw new Error(`Unexpected success content type: ${type}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await require('node:fs').promises.writeFile('shot.bin', bytes);
break;
}
const type = response.headers.get('content-type') || '';
const body = type.includes('json') ? await response.json() : await response.text();
const code = typeof body === 'object' && body ? String(body.code || '') : '';
const message = typeof body === 'object' && body ? String(body.message || '') : String(body);
const lower = `${code} ${message}`.toLowerCase();
console.error({ status: response.status, code, retryAfter: response.headers.get('retry-after'), attempt });
if (response.status === 429 && /(quota|billing|monthly)/.test(lower)) {
throw new Error(`Non-retryable usage limit: ${message}`);
}
if (![429, 503].includes(response.status) || attempt === maxRetries) {
throw new Error(`Capture failed (${response.status}): ${message}`);
}
let delay = retryAfterMs(response.headers.get('retry-after'));
if (delay === null) {
delay = Math.min(maxDelayMs, baseDelayMs * (2 ** attempt)) + Math.random() * 500;
}
if (Date.now() - started + delay > maxTotalMs) {
throw new Error('Retry deadline exceeded; defer this job');
}
await sleep(delay);
}
cURL inspection and retry
curl --dump-header response.headers \
--fail-with-body \
--retry 0 \
-G "https://provider.example/v1/shot" \
--data-urlencode "url=https://example.com" \
-o shot.bin
Use --dump-header to inspect Retry-After and reset headers. Do not use an unbounded shell retry loop. Parse the error body, classify quota versus throttling, and apply the same attempt and deadline limits as in application code.
Step 4: Prevent rate-limit errors
- Use a bounded worker pool. Set a per-provider concurrency limit instead of starting one request per URL.
- Queue and pace dispatch. A token bucket or leaky bucket smooths bursts. Lower concurrency when remaining capacity falls or a reset approaches.
- Deduplicate and cache. Avoid capturing the same URL and option set repeatedly when freshness allows. Provider caching may also reduce work, but check each provider’s cache semantics.
- Batch where supported. Batch endpoints reduce request overhead; still account for each screenshot in the provider’s quota.
- Ramp traffic gradually. Releasing a backlog immediately after deployment can trigger throttling even when the average per-minute rate looks acceptable.
- Coordinate retries. Add jitter so many workers do not wake and retry together.
- Check SDK defaults. An SDK may already retry 429 and 503 responses. Disable nested retries or reduce your application retry budget so total attempts remain bounded.
Rate limits versus monthly quotas
Rate limits are usually short windows such as requests per second or minute. They recover after a delay or reset timestamp. Monthly quotas count successful screenshot renders over a billing period. A provider may return 429 for either condition, so branch on the documented machine-readable error code and current usage headers.
For example, ScreenshotEngine documents separate temporary 429 responses and monthly “Quota Exceeded” responses, with plan-specific screenshots-per-month and requests-per-minute values. Its guidance says to honor Retry-After, reduce concurrency, and avoid automatically retrying monthly quota errors (ScreenshotEngine documentation). Screenshot API documents rate_limited and quota_exceeded codes plus X-RateLimit-* and X-Quota-* headers (Screenshot API documentation). Verify current provider terms because plans and limits change.
Timeouts, duplicate captures, and idempotency
A client timeout does not prove that the provider failed. The renderer may have completed just before the connection dropped. Retrying blindly can create a duplicate successful capture and consume quota. Keep a client job ID, provider request ID, target URL, option hash, and approximate timestamp. If the provider supports idempotency keys or job lookup, use them. Otherwise, make duplicate handling explicit in your queue and accept that an uncertain timeout may require reconciliation.
Troubleshooting common errors
| Symptom | Likely cause | Fix |
|---|---|---|
| 429 immediately after a burst | Concurrency or requests-per-minute limit | Honor Retry-After, reduce workers, and pace the queue. |
| 429 continues for hours | Monthly quota or billing cap | Inspect the error code and usage dashboard; stop retries and resolve the account limit. |
| Retries happen twice per request | SDK and application both retry | Disable one layer or divide the total attempt budget. |
| Every retry gets 429 | No jitter, too much concurrency, or retrying before reset | Use server-provided delay, random jitter, and a bounded queue. |
| 429 body cannot be parsed as an image | Error response is JSON or text | Check status and content type before writing or decoding binary output. |
| 429 after valid credentials and parameters | Rate or usage policy, not authentication | Inspect machine-readable error fields and headers; do not assume all 4xx responses mean bad credentials. |
| 503 after a successful-looking request | Transient renderer or service failure | Retry a small number of times with backoff, then defer and retain request IDs. |
| Traffic works manually but fails in production | Production burst, shared account, or organization-wide limit | Measure aggregate traffic across workers and services, then enforce a single shared limiter. |
Performance, reliability, and cost considerations
- Throughput: Raising concurrency can lower queue time until the provider’s burst limit is reached; beyond that point it increases 429s and reduces useful throughput.
- Latency: Include queue wait, server retry delay, renderer time, and download time in your deadline. A short HTTP timeout with a long retry policy can create false failures.
- Reliability: Persist deferred jobs, classify permanent failures separately, and alert on quota exhaustion instead of repeatedly retrying it.
- Cost: Check whether the provider bills requests, successful renders, cache hits, or failed renders. A retry policy must match those rules. A client timeout can still correspond to a billable successful capture.
- Operations: Track status counts, error codes, retry attempts, delay time, queue age, and estimated quota consumed. Redact keys, authorization headers, cookies, and signed URLs from logs.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. It removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
Read the ScreenshotNeo API documentation for all options and response details. A one-call capture looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page and element capture, device presets, retina scale, PDF output, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture, and a usage API. Every feature is on every plan. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
FAQ
Should I retry every 429?
No. Retry temporary throttling when the provider indicates it. Stop for monthly quota, billing, authentication, and invalid-input errors.
What if Retry-After is missing?
Use capped exponential backoff with random jitter, a maximum attempt count, and a total deadline. Consider a reset header when its semantics are documented.
Can retries make throttling worse?
Yes. Unsuccessful attempts can consume request-rate capacity, and synchronized workers can create another burst. Queue work and add jitter.
How many retries should a screenshot job get?
There is no universal number. Set a small bounded count and deadline based on your latency budget, then defer jobs whose server-provided delay exceeds that budget.
Are 500 and 503 handled like 429?
They can be transient renderer or service failures. Retry only with provider guidance and the same bounded, jittered policy; do not retry indefinitely.


