ScreenshotNeo

BlogGuides

Screenshot API Request Limits: Requests per Minute and Concurrency Explained

Learn how screenshot API rate limits, concurrency, and monthly quotas differ, and how to pace jobs, handle 429 responses, and avoid wasteful retries.

By the ScreenshotNeo team4 October 202610 min read

A screenshot API limit can refer to three different things: how many requests you may start in a short time window, how many renders may run simultaneously, or how many successful screenshots your plan allows in a billing period. There is no universal limit or shared definition of “concurrency.” Check the API provider’s documentation and your account’s live usage data before setting a pace.

For a 429 response, inspect its body, headers, and any Retry-After value. A temporary rate cap calls for waiting before starting more work. An exhausted monthly quota calls for pausing or changing the plan; retrying immediately will not restore quota.

1. The three limits to track

Limit What it controls What to check
Request rate How many request starts are accepted during a time window, often per second or per minute. Limit, remaining capacity, reset time, and any burst rules.
Concurrency Potentially the number of active renders, but provider usage varies. Some APIs use this label for a request-start bucket instead. The provider’s definition of the field and whether a distinct active-render cap exists.
Monthly quota The plan-period allowance of screenshots or successful renders. What counts toward quota, remaining allowance, and reset date.

These constraints are independent. A client can stay below a per-minute cap and still use up its monthly allowance. It can also have monthly quota left and receive a 429 because it started requests too quickly.

2. Request rate is not the same as concurrency

A requests-per-minute limit normally governs request starts in some time window. The provider determines whether the window is fixed, rolling, bucketed, or combined with a burst limit. Do not assume that a cap of 60 requests per minute means 60 browser renders can run at once.

“Concurrency” also needs provider-specific interpretation. ScreenshotOne’s usage documentation says its concurrency values are request-bucket values, not active render counts. In that API, concurrency.limit is the number of requests that can be started during the current one-minute bucket, remaining is the number still startable before reset, and reset is a nanosecond Unix timestamp. Those meanings apply to ScreenshotOne’s fields; do not assume other APIs use the same model. ScreenshotOne: Get Usage.

ScreenshotNeo’s documentation describes a 429 rate_limited / concurrency_limit condition and directs clients to Retry-After. That is a reason to inspect the provider’s own explanation rather than infer active-render capacity from a field name. ScreenshotNeo API documentation.

3. Published examples, checked October 3, 2026

The figures below are examples shown in the named providers’ documentation when checked on 2026-10-03. They are configuration details, not a cross-provider benchmark or universal standard. Plans and limits can change, so confirm the live docs and your account before relying on them.

Provider and plan Short-window limit Period quota Documented details
Screenshot API (screenshot-api.org), Free 60 requests/minute 500 screenshots/month Documents rate-limit and quota remaining/reset headers.
ScreenshotEngine, Free 5 requests/minute 50 screenshots/month Distinguishes rate-limit 429 from monthly-quota 429; failed calls remain rate-limited even when they do not count toward successful captures.
ScreenshotEngine, Starter 40 requests/minute 3,000 screenshots/month Values listed in its limits documentation.
ScreenshotEngine, Professional 100 requests/minute 15,000 screenshots/month Values listed in its limits documentation.
ScreenshotEngine, Engine 250 requests/minute 60,000 screenshots/month Values listed in its limits documentation.
Screenshot API (screenshot-api.net), Free 1 request/second 100 renders/month This is a separate service from screenshot-api.org.
Screenshot API (screenshot-api.net), Starter 5 requests/second 2,000 renders/month Separate request-rate and render quota.
Screenshot API (screenshot-api.net), Pro 10 requests/second 10,000 renders/month Values shown in its documentation.
Screenshot API (screenshot-api.net), Team 25 requests/second 25,000 renders/month Values shown in its documentation.
Screenshot API (screenshot-api.net), Business 50 requests/second 100,000 renders/month Values shown in its documentation.

Sources: Screenshot API (.org) REST API reference, ScreenshotEngine errors and limits, and Screenshot API (.net) documentation. Each provider defines its own enforcement and what counts toward usage; these rows should not be treated as an apples-to-apples ranking.

4. Inspect usage and diagnose 429 responses

  1. Read the response body. Look for a provider-specific code or message that distinguishes a short-window rate cap, active-render cap, or depleted quota.
  2. Read Retry-After. If supplied, wait at least that long before retrying the affected work.
  3. Inspect limit and quota headers. For example, Screenshot API (.org) documents X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, X-Quota-Remaining, and X-Quota-Reset. Header names differ among services.
  4. Check the account usage endpoint or dashboard. Reconcile your local counters with the provider’s live account data where available.
  5. Classify before retrying. A temporary rate response may succeed after waiting; an exhausted monthly allowance or invalid request will not.

Do not infer whether failed requests consume a rate bucket or quota. Providers differ: ScreenshotEngine says failed requests remain subject to rate limiting even though they do not count toward successful-capture allowance. Check the service’s own policy.

5. Pace a queue safely

For batch or recurring jobs, use a queue instead of launching every URL at once. Keep a per-account or per-key limiter aligned with the provider’s enforcement scope. If the API publishes remaining and reset values, use them to schedule starts. Maintain an independent cap on in-flight work only when the vendor documents a simultaneous-render limit or your own worker capacity requires one.

Retry policy

  • For 429 and temporary 503 responses, honor Retry-After when present.
  • If no retry delay is supplied, use bounded exponential backoff with jitter, for example a short initial delay that grows per attempt, a sensible maximum delay, and at most three retries.
  • Do not automatically retry invalid input, invalid credentials, or an exhausted quota.
  • For a client timeout, check whether the provider completed the capture before submitting another request. A retry can create a second successful, billable capture.

Python example: bounded retry and pacing

This generic example assumes the provider accepts a screenshot URL in a query parameter and returns an image body. Adapt the endpoint, authentication, successful status codes, and response parsing to the specific API. It honors Retry-After when present, retries only 429/503, adds jitter otherwise, and stops after three retries.

import random
import time
from email.utils import parsedate_to_datetime
from datetime import datetime, timezone

import requests

API_URL = "https://YOUR-SCREENSHOT-API/shot"
API_KEY = "YOUR_API_KEY"


def retry_after_seconds(value):
    if not value:
        return None
    try:
        return max(0.0, float(value))
    except ValueError:
        try:
            retry_at = parsedate_to_datetime(value)
            if retry_at.tzinfo is None:
                retry_at = retry_at.replace(tzinfo=timezone.utc)
            return max(0.0, (retry_at - datetime.now(timezone.utc)).total_seconds())
        except (TypeError, ValueError, OverflowError):
            return None


def capture(url):
    for attempt in range(4):  # initial attempt plus at most three retries
        response = requests.get(
            API_URL,
            params={"access_key": API_KEY, "url": url},
            timeout=90,
        )
        if response.status_code not in (429, 503):
            response.raise_for_status()
            return response.content, response.headers

        if attempt == 3:
            response.raise_for_status()

        delay = retry_after_seconds(response.headers.get("Retry-After"))
        if delay is None:
            delay = min(30.0, 1.0 * (2 ** attempt)) + random.uniform(0, 0.5)
        time.sleep(delay)

    raise RuntimeError("Retry loop ended unexpectedly")


if __name__ == "__main__":
    image_bytes, headers = capture("https://example.com")
    with open("capture.png", "wb") as output:
        output.write(image_bytes)
    print("Saved capture.png")
    print("Rate remaining:", headers.get("X-RateLimit-Remaining"))
    print("Quota remaining:", headers.get("X-Quota-Remaining"))

For a production queue, add a shared limiter across workers and persist scheduled retry times. A limiter local to each process can still exceed a per-account cap when several workers run at once.

cURL: inspect a response without discarding headers

Use the provider’s authentication method. This illustration writes response headers to a file so you can inspect rate and quota metadata; replace the placeholder host and parameter names with the API’s documented values.

curl -sS -D response-headers.txt -o screenshot.png \
  -G "https://YOUR-SCREENSHOT-API/shot" \
  --data-urlencode "access_key=YOUR_API_KEY" \
  --data-urlencode "url=https://example.com"

cat response-headers.txt

Node.js: bounded retry and header inspection

const endpoint = 'https://YOUR-SCREENSHOT-API/shot';
const apiKey = 'YOUR_API_KEY';

function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}

function retryDelay(response, attempt) {
  const value = response.headers.get('retry-after');
  if (value) {
    const seconds = Number(value);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1000);
    const date = Date.parse(value);
    if (Number.isFinite(date)) return Math.max(0, date - Date.now());
  }
  return Math.min(30_000, 1_000 * (2 ** attempt)) + Math.random() * 500;
}

async function capture(url) {
  const query = new URLSearchParams({ access_key: apiKey, url });
  for (let attempt = 0; attempt <= 3; attempt++) {
    const response = await fetch(`${endpoint}?${query}`);
    if (response.status !== 429 && response.status !== 503) {
      if (!response.ok) {
        throw new Error(`Screenshot API returned ${response.status}: ${await response.text()}`);
      }
      return {
        bytes: Buffer.from(await response.arrayBuffer()),
        rateRemaining: response.headers.get('x-ratelimit-remaining'),
        quotaRemaining: response.headers.get('x-quota-remaining'),
      };
    }
    if (attempt === 3) {
      throw new Error(`Retry limit reached after HTTP ${response.status}`);
    }
    await sleep(retryDelay(response, attempt));
  }
}

const result = await capture('https://example.com');
await import('node:fs/promises').then(fs => fs.writeFile('capture.png', result.bytes));
console.log({ rateRemaining: result.rateRemaining, quotaRemaining: result.quotaRemaining });

This is an application-level pattern, not a guarantee that every API uses these status codes, header names, or authentication parameters. Follow the target provider’s reference.

6. Concurrency, throughput, and performance

Rate limits constrain starts; render duration constrains how many jobs finish per unit of time. If requests take a long time, a permitted start rate can still create a large in-flight queue. Conversely, a low active-render cap can keep throughput below the request-start limit. Measure your own request durations and use the provider’s documented model to set queue size.

  • Start with a conservative queue and increase gradually while watching 429s, timeouts, remaining capacity, and completion time.
  • Use a token bucket or scheduled queue if you need to smooth bursts. A simple average requests-per-minute calculation may still violate a provider’s per-second burst rule.
  • Share pacing state among workers when limits are account-wide or key-wide.
  • Cache repeat captures when freshness requirements permit; deduplicating identical work can reduce both load and quota use where the service’s cache semantics support it.
  • Set connection and overall timeouts deliberately. A client timeout does not prove the remote render stopped.

No cross-provider throughput benchmark is established by the documentation cited here. The practical capacity depends on the provider’s enforcement rules, render duration, target pages, and your job mix.

7. Reliability and cost implications

Track rate capacity and period quota as separate budgets. Estimate period use from the number of captures you expect to complete, and account for retries only according to the provider’s stated billing rules. A rate-limit response can consume request capacity even if it does not count as a successful screenshot; ScreenshotEngine documents that distinction for its service.

Retries improve resilience for temporary overload, but they can also multiply traffic. Bound attempts, add jitter so many workers do not retry together, and record the original URL, attempt count, response code, and next retry time. Do not log API keys in query strings. Screenshot API (.net) warns that query-string keys may be exposed and recommends bearer authentication for production calls; use the target API’s documented secure method. Screenshot API (.net) documentation.

Keep enough operational metrics to answer: how many jobs are queued, how many are in flight, what are the recent 429/503 rates, how much quota remains, and when does it reset? Alert before the monthly quota is exhausted if the workload is important.

8. Troubleshooting

Symptom Likely cause Fix
429 immediately after a burst Request-start window or burst limit reached. Stop starting work, honor Retry-After or reset metadata, and smooth the queue.
429 continues after waiting Monthly quota may be depleted, or the limiter is shared across workers or keys. Read the error body and account usage; identify enforcement scope and quota reset.
429 despite low local concurrency The provider’s “concurrency” field may count request starts, or another process is using the same account. Confirm field semantics and coordinate pacing across all clients.
Requests stay below per-minute limit but quota runs out Short-window rate and monthly allowance are separate constraints. Reduce total job volume, deduplicate captures, or choose a plan with sufficient quota.
Retry loop creates more errors Invalid input, bad credentials, exhausted allowance, or unbounded retries. Retry only transient responses; fix the request or account state for permanent failures.
Duplicate capture or unexpected charge after timeout The server may have completed the render after the client gave up. Check job status or request identifiers if supported before retrying; use idempotency only if the provider documents it.
Header values appear missing The provider may not expose them on that endpoint, or a proxy may strip them. Consult docs and inspect the raw response from the API directly.

9. Or skip the browser setup

If the job is taking screenshots rather than testing rate-limit behavior itself, ScreenshotNeo provides a screenshot API and MCP server. One GET request returns an image or PDF, and the API’s response headers report whether a request was billed and the page verdict.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options and response behavior. ScreenshotNeo removes known cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account.

10. FAQ

Does a 429 always mean I sent too many requests per minute?

No. It can represent a short-window rate cap, an active-work cap, or depleted quota. Use the provider’s response details and usage data to identify which one.

Can I safely retry a screenshot request after a timeout?

Not automatically. The remote service may have completed the capture after your client timed out. Check job status or request identifiers if available, because a duplicate request may create a second capture.

Are requests per second and requests per minute interchangeable?

No. A per-second burst cap can reject a burst even when the minute-wide total is below its limit. Apply the exact windows documented by the API.

Is concurrency always the number of renders running at once?

No. ScreenshotOne documents its usage field as a request bucket. Other providers may define concurrency differently or enforce a separate active-render cap.

References