ScreenshotNeo

BlogHow-to

Website screenshot API returns a 429 error: rate limit fixes

A 429 can mean a burst limit, concurrency cap, or exhausted quota. Inspect the response, honor Retry-After, and pace retries according to your provider’s rules.

By the ScreenshotNeo team4 October 202611 min read

HTTP 429 means the screenshot API rejected a request under a rate limit or another provider-defined limit. First inspect the status, response body, and headers; check your account usage; wait for Retry-After if present; then reduce request rate or concurrent renders. Do not assume every provider uses 429 for the same condition: one may use it for a short burst limit and a different status for exhausted monthly quota, while another may use 429 for either.

The HTTP standard does not prescribe a universal rate window, counter, or identity scope. A limit might apply per key, account, IP address, endpoint, or another provider-defined bucket. The response and that provider’s documentation determine the right fix. RFC 6585 defines 429 as “Too Many Requests” and says a response may include Retry-After.

Quick diagnostic sequence

  1. Confirm the response came from the API. Record the HTTP status, content type, response body, and relevant rate-limit headers. Screenshot APIs often return image bytes on success and JSON or text on errors, so check before saving a response as an image.
  2. Read the provider’s error code and message. Look for distinctions such as rate limited, concurrency limit, or quota reached.
  3. Check usage and quota. Use the provider’s dashboard or usage endpoint. Note the quota reset schedule if the account has exhausted a monthly allowance.
  4. Honor Retry-After. It can be a delay in seconds or an HTTP date. Wait at least that long before retrying.
  5. Reduce traffic. Lower request rate and in-flight concurrency; coordinate workers that share a key or account limit.
  6. Retry only temporary failures. Use bounded exponential backoff with jitter when appropriate. Do not retry invalid input, bad credentials, or confirmed quota exhaustion.
  7. For bulk work, use a queue or async mode if available. This controls bursts, but does not bypass a provider’s usage limits.

Why am I getting a 429 error from my screenshot API?

At the protocol level, the server is refusing a request because it considers too many requests to have arrived within its chosen rules. The server determines how requests are counted and how a client is identified; the RFC does not mandate a global requests-per-second limit. See RFC 6585, section 4.

For screenshot services, the triggering condition may be a short request burst, too many renders running at once, or a provider’s quota policy. For example, ScreenshotEngine documents both request-window limiting and monthly allowance exhaustion as possible 429 causes. Screenshot API documents 429 for a burst limit and a separate 402 response for exhausted monthly quota. These are examples of differing provider behavior, not universal conventions. Check the docs for the API you call: ScreenshotEngine errors and limits and Screenshot API documentation.

How do I inspect a 429 response?

Capture the response before your client library treats it as a file or throws away useful headers. Preserve the body as text for diagnosis; never log API keys, authorization headers, or cookies.

cURL: inspect status, headers, and body

curl -sS -D response-headers.txt \
  -o response-body.txt \
  -w 'HTTP %{http_code}\nContent type: %{content_type}\n' \
  -G 'https://YOUR_PROVIDER_ENDPOINT' \
  --data-urlencode 'url=https://example.com' \
  -H 'Authorization: Bearer YOUR_API_KEY'

cat response-headers.txt
cat response-body.txt

Replace the endpoint and authentication format with your provider’s actual values. If the endpoint returns an image on success, use a separate output filename for successful captures and inspect the status before treating the body as an image.

Python: inspect status before saving image bytes

import requests

response = requests.get(
    "https://YOUR_PROVIDER_ENDPOINT",
    params={"url": "https://example.com"},
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    timeout=90,
)

print("status:", response.status_code)
print("content-type:", response.headers.get("Content-Type"))
print("retry-after:", response.headers.get("Retry-After"))

if response.status_code == 429:
    print("error body:", response.text)
else:
    response.raise_for_status()
    with open("shot.png", "wb") as image_file:
        image_file.write(response.content)

Node.js: inspect status and headers

const response = await fetch('https://YOUR_PROVIDER_ENDPOINT?url=https%3A%2F%2Fexample.com', {
  headers: { Authorization: 'Bearer YOUR_API_KEY' },
});

console.log('status:', response.status);
console.log('content-type:', response.headers.get('content-type'));
console.log('retry-after:', response.headers.get('retry-after'));

if (response.status === 429) {
  console.error('error body:', await response.text());
} else if (!response.ok) {
  throw new Error(`Screenshot API returned HTTP ${response.status}: ${await response.text()}`);
} else {
  const image = Buffer.from(await response.arrayBuffer());
  await import('node:fs/promises').then(fs => fs.writeFile('shot.png', image));
}

These examples use a placeholder endpoint because the title does not identify a provider. Replace its endpoint, parameters, and authentication with the documented values for your service.

What does Retry-After mean?

Retry-After tells a client how long it ought to wait before making a follow-up request. Under RFC 9110, section 10.2.3, its value can be either a number of seconds or an HTTP date. Wait at least the indicated interval. If the value is a date, calculate the delay from the current time; account for clock skew conservatively.

Do not assume every API sends this header or uses the same format. If it is absent, follow the provider’s retry guidance and use a capped delay with jitter rather than immediately repeating the call.

Parse either standard form in Python

from datetime import datetime, timezone
from email.utils import parsedate_to_datetime


def retry_after_seconds(value: str | None) -> float | None:
    if not value:
        return None
    value = value.strip()
    try:
        return max(0.0, float(value))
    except ValueError:
        pass
    try:
        retry_at = parsedate_to_datetime(value)
        if retry_at.tzinfo is None:
            retry_at = retry_at.replace(tzinfo=timezone.utc)
        return max(0.0, (retry_at - datetime.now(timezone.utc)).total_seconds())
    except (TypeError, ValueError, OverflowError):
        return None

seconds = retry_after_seconds(response.headers.get("Retry-After"))
if seconds is not None:
    print(f"Wait at least {seconds:.1f} seconds before retrying")

How do I fix a 429 rate limit?

Lower request rate and concurrency

Rate and concurrency are different controls. Rate limits restrict how many requests arrive within a time window; concurrency limits restrict how many renders are in progress simultaneously. Lower the worker count and add a queue so bursts do not overwhelm either limit. If multiple services or processes use the same key, coordinate them with a shared limiter; per-process throttling alone can still exceed a shared account limit.

Make changes gradually and observe the status, body, and headers. Do not copy another provider’s published request limit: limits depend on the provider, account, plan, endpoint, and current configuration.

Use bounded retries with exponential backoff and jitter

Retry only if the provider indicates the failure is temporary, or its documentation says a 429 is retryable. Respect Retry-After when supplied. Without it, increase the delay after each attempt, add random jitter to spread requests from multiple workers, and cap both delay and attempts. Stop after the cap and surface the failure to the job queue or caller.

import random
import time


def retry_delay(attempt: int, retry_after: float | None,
                base: float = 1.0, cap: float = 30.0) -> float:
    """attempt starts at 0; returns a delay in seconds."""
    if retry_after is not None:
        # Add only positive jitter so the client never retries before the hint.
        return retry_after + random.uniform(0, min(1.0, retry_after * 0.1))
    ceiling = min(cap, base * (2 ** attempt))
    return random.uniform(0, ceiling)  # full jitter


# Use only after confirming this 429 is temporary and retryable.
max_attempts = 4
for attempt in range(max_attempts):
    result = make_screenshot_request()
    if result.status_code != 429:
        break
    if is_quota_exhausted(result):
        raise RuntimeError("Quota is exhausted; check usage and reset or plan details")
    if attempt == max_attempts - 1:
        raise RuntimeError("Rate limit persisted after bounded retries")
    delay = retry_delay(attempt, parse_retry_after(result.headers.get("Retry-After")))
    time.sleep(delay)

make_screenshot_request, is_quota_exhausted, and parse_retry_after are provider-specific integration points; implement them from the response schema and documentation. The code deliberately does not retry every 429 blindly.

Check whether quota is exhausted

If usage shows that the monthly allowance is spent, waiting a few seconds and retrying will not restore it. Check the provider’s reset time, account usage endpoint, and plan details. A provider may use 429, 402, or another documented response for quota exhaustion. Avoid repeatedly retrying a known exhausted quota.

Queue or batch large jobs

For a backlog of URLs, put work behind a bounded queue and keep only a controlled number of renders in flight. If the provider supports asynchronous jobs or bulk requests, consider those interfaces and their documented limits. Batching can reduce request overhead, but it does not make included renders free of quota accounting unless the provider explicitly says so.

Provider limits differ: what should I compare?

When reading the provider’s documentation or asking support, identify the exact limit dimension rather than asking only for “the 429 limit.” Compare these fields:

What to check Why it matters
Rate bucket It may be keyed by account, API key, IP, endpoint, or another identity.
Window A per-second burst rule behaves differently from a per-minute allowance.
Concurrency cap Slow renders can leave many jobs in flight even at a modest request rate.
Quota status and reset Monthly quota exhaustion needs a usage or plan remedy, not rapid retries.
Error code and response body These may distinguish a burst limit from concurrency or quota exhaustion.
Retry-After format Clients must handle seconds and HTTP dates correctly.
Usage visibility A dashboard or usage endpoint helps confirm which counter was reached.
Async or bulk support Queues can smooth traffic for large jobs, subject to provider rules.

Examples show why you must check the service you actually use: ScreenshotEngine documents 429 for request-window and monthly allowance conditions, while Screenshot API documents 429 for rate limiting and 402 for quota reached. ScreenshotNeo documents rate_limited and concurrency_limit 429 codes and says its Retry-After header gives retry timing; see the ScreenshotNeo documentation. These are provider-specific examples, not standard plan limits.

Common 429 troubleshooting cases

Symptom Likely cause Fix
429 appears only during bursts A short-window rate or burst limiter is firing. Throttle calls, smooth the queue, and honor Retry-After.
429 appears while request volume seems low Several workers may share a key, or a slow render may hit a concurrency cap. Check shared account activity and reduce simultaneous renders.
429 continues after waiting The response may mean quota exhaustion, the wait may be too short, or another worker keeps consuming the shared limit. Inspect body/code, usage, reset schedule, and other clients using the key.
Saved “image” will not open The client saved a JSON or text error response as image bytes. Check HTTP status and content type before writing the file. See ScreenshotEngine’s error guidance.
Retries make the problem worse Immediate or unbounded retries add traffic to the limited bucket. Use bounded exponential backoff, jitter, and the server’s retry hint.
Monthly usage appears depleted The provider may count successful captures or requests according to its own policy. Check its usage view and quota rules; stop retrying until the allowance or plan issue is resolved.
Duplicate capture or charge after a timeout The server may have completed the capture even though the client did not receive the response. Check for a job/result lookup or idempotency mechanism before resubmitting. ScreenshotEngine notes that a timeout after a successful capture can cause a duplicate billed request; verify your provider’s behavior.
Errors started after adding more workers Combined concurrency or request rate crossed a shared limit. Set a global worker cap or shared queue instead of independent per-worker pacing.

Performance, reliability, and cost considerations

  • Throughput: A queue and concurrency cap make throughput predictable, though a lower cap can increase completion time. Tune against observed provider responses rather than guessed universal limits.
  • Reliability: Bounded retries prevent a temporary limiter response from failing a whole batch immediately. Dead-letter or report jobs that remain unsuccessful after the retry budget.
  • Duplicate work: A client timeout does not prove the provider failed to complete the render. Before retrying a timed-out request, check whether the provider offers job IDs, result retrieval, or idempotency support.
  • Cost: Billing and quota counters vary. Verify whether failed renders, retries, cache hits, and asynchronous jobs count with your provider. Avoid assuming that a 429 is free or billed without documentation.
  • Observability: Log timestamp, endpoint, status, provider error code, content type, retry hint, and a request/job identifier if available. Redact API keys, cookies, and authorization values.

Escalate with useful diagnostics

If the error persists after the indicated wait and usage checks, provide support with the endpoint and method, approximate request time, HTTP status, response code and body, relevant non-secret rate-limit headers, and non-secret capture options. Remove API keys, bearer tokens, cookies, and private target URLs where needed. ScreenshotEngine’s troubleshooting guidance similarly asks for request details and warns against sharing keys.

Or skip the browser setup

If you need screenshot captures without managing browser infrastructure, ScreenshotNeo is a website screenshot API and MCP server for developers. A GET request returns an image or PDF. Its docs describe 429 rate_limited and concurrency_limit errors and Retry-After guidance, so still pace requests according to the account’s limits.

Install the Python dependency with python -m pip install requests, set your API key, then run this example:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

For options and response details, see the ScreenshotNeo API documentation. The equivalent cURL request is:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}: ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie and consent banners are accepted before capture; 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot. Each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and whether the request was billed.
  • An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
  • 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is on every plan.

Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.

FAQ

Is every HTTP 429 a rate limit?

It is the HTTP status for too many requests, but providers can attach different counting rules and may also use it for conditions such as exhausted quota. Read the provider’s response details.

Will waiting fix an exhausted monthly quota?

Only if the allowance resets during the wait. Check the account usage and reset policy; repeated requests do not replenish spent quota.

Should I retry a screenshot request after a timeout?

First check whether the provider completed the job or offers an idempotency or result lookup mechanism. A timeout can happen after server-side work finishes, so blind resubmission may create duplicate work.

Can I use one fixed request-per-second value for every screenshot API?

No. Rate buckets, windows, concurrency rules, and quota semantics are provider-specific and can vary by account. Use that provider’s current documentation and response headers.