ScreenshotNeo

BlogEngineering

Screenshot API Concurrency Limits: What Happens When You Exceed Them?

Learn what screenshot API concurrency limits mean, how to identify the limit behind a 429, and how to queue and retry captures safely.

By the ScreenshotNeo team4 October 202610 min read

A screenshot API concurrency limit controls how quickly your client can start capture requests, how many renders can run at once, or both. When you exceed one, the provider may reject requests—often with HTTP 429—or delay them. The response code alone does not tell you which limit you hit. Read the machine-readable error, headers, usage endpoint, and account plan before deciding whether to wait, reduce request starts, or address a depleted quota.

For a concrete example, ScreenshotOne defines concurrency_limit_reached as an empty request bucket for the current one-minute window. That limit counts requests started in the bucket, not renders still running: a request that finishes quickly still used one start. Other providers may define concurrency differently, so treat this as provider-specific behavior, not a universal rule. ScreenshotOne’s concurrency limit documentation explains its behavior and recovery.

What happens when you exceed a screenshot API limit?

The provider rejects or throttles new work once the applicable allowance is exhausted. You may see a 429 response, a provider-specific error code, or a response that identifies a monthly quota or spending limit. Work already accepted may continue, but do not assume that a rejected request will be queued automatically. Persist pending jobs in your own queue and resume them when the provider says capacity is available.

There are three common kinds of limits:

  • Request-start rate: limits how many requests may begin in a time window. It can be exhausted even when renders finish quickly.
  • Active-render concurrency: limits the number of captures running at the same time. A slot may open when a render finishes.
  • Usage quota or spend ceiling: limits successful captures, monthly usage, account balance, or spending. Waiting for a short rate window to reset may not resolve it.

A provider may enforce more than one kind at once. Check its current documentation and dashboard for your account’s actual limits and plan; numerical ceilings are service- and plan-specific and may change.

How to diagnose the limit behind a 429

  1. Read the response body. Capture the provider’s error code and message. Do not label every 429 as a concurrency error.
  2. Inspect response headers. Look for Retry-After, remaining-request values, reset times, request IDs, or quota details. Header names vary by provider.
  3. Check the usage endpoint or account dashboard. Compare rate-window capacity with monthly successful captures, balance, and spending limits.
  4. Choose the matching recovery. Pause and retry after a temporary window resets; reduce starts if request rate is the issue; address account usage or plan limits if quota or spending is exhausted.
  5. Log enough context to investigate. Record timestamp, status, error code, request ID if present, and relevant remaining/reset fields. Redact API keys and sensitive page data.

ScreenshotOne’s documented error is concurrency_limit_reached. Its usage endpoint exposes concurrency.remaining and concurrency.reset; queue work according to remaining capacity, and wait until reset when it is zero. The vendor’s explanation uses 40 requests per minute as an example of a plan limit, not as a ceiling that applies to every account.

A 429 can also mean monthly quota exhaustion. For example, ScreenshotEngine documents separate per-minute request limits and monthly successful-capture allowances, and says both can return 429; its response body distinguishes the causes. Its published plan figures are specific to ScreenshotEngine and are not general API limits. See its rate-limit documentation and account dashboard for current account details.

Queue and retry captures safely

For bulk capture, put URLs into a queue and let a controlled number of workers start requests. Shape starts using the provider’s remaining allowance and reset time when available. Do not set the worker count from average render duration alone: a fast render still consumes a request-start allowance when the provider counts starts per window.

  1. Keep unstarted jobs pending instead of launching an unbounded burst.
  2. Before starting work, use the provider’s usage data or documented limit headers to determine whether capacity remains.
  3. On a temporary limit, honor a valid Retry-After value. If it is missing or invalid, use bounded exponential backoff with random jitter.
  4. Set both a maximum retry count and a maximum total retry time. After that, surface the job for inspection instead of retrying forever.
  5. Make retries safe for your workflow. Store job identity and output state so a network timeout does not accidentally create duplicate downstream work.
  6. For jobs that must survive process restarts, use a durable queue. ScreenshotOne’s documentation names Redis/BullMQ or SQS as examples; an in-process queue can be a starting point for simpler workloads.

Do not layer an unlimited application retry loop on top of SDK retries. Repeated unsuccessful requests can contribute to per-minute limits. If the response identifies an exhausted monthly quota, balance, or spend ceiling, stop retries and resolve that account condition first.

Runnable retry example in Python

This example shows a bounded request loop for a provider endpoint that returns an image on success and a JSON error on failure. Replace the endpoint and authentication parameter with your provider’s documented values. It honors Retry-After when it is a number of seconds, otherwise applies capped exponential backoff with jitter. If your provider returns a reset timestamp or usage endpoint, use that documented signal in place of the fallback delay.

import random
import time
import requests

API_URL = "https://api.example.com/screenshot"
MAX_ATTEMPTS = 5
MAX_DELAY_SECONDS = 60


def capture(url):
    for attempt in range(MAX_ATTEMPTS):
        response = requests.get(
            API_URL,
            params={"url": url, "api_key": "YOUR_API_KEY"},
            timeout=90,
        )
        if response.ok:
            return response.content

        if response.status_code != 429:
            response.raise_for_status()

        # Inspect the body before retrying: a quota or billing limit needs
        # account action, not another immediate request.
        try:
            detail = response.json()
        except ValueError:
            detail = {"error": response.text[:500]}
        code = detail.get("error_code") or detail.get("code")
        message = str(detail.get("message") or detail.get("error") or "")
        lowered = (str(code) + " " + message).lower()
        if any(word in lowered for word in ("quota", "balance", "spend", "billing")):
            raise RuntimeError(f"Account limit needs attention: {detail}")

        if attempt == MAX_ATTEMPTS - 1:
            raise RuntimeError(f"Retry limit reached: {detail}")

        retry_after = response.headers.get("Retry-After")
        try:
            delay = float(retry_after)
            if delay < 0:
                raise ValueError()
        except (TypeError, ValueError):
            ceiling = min(MAX_DELAY_SECONDS, 2 ** attempt)
            delay = random.uniform(0, ceiling)
        time.sleep(min(delay, MAX_DELAY_SECONDS))

    raise RuntimeError("Capture did not complete")


image_bytes = capture("https://example.org")
with open("capture.png", "wb") as output:
    output.write(image_bytes)

The example treats 429s with quota or billing clues as account errors. Adapt that classification to the provider’s documented error schema; do not infer the cause from status alone. If the API provides remaining and reset, schedule the next attempt against those values rather than guessing from a generic delay.

Equivalent handling in cURL and Node.js

For manual diagnosis, first capture headers and body instead of discarding them. These commands make one request and let you inspect the response; production retry scheduling should be handled by a bounded queue or application loop.

cURL

curl -sS -D response-headers.txt \
  -G "https://api.example.com/screenshot" \
  --data-urlencode "url=https://example.org" \
  --data-urlencode "api_key=YOUR_API_KEY" \
  -o response-body.bin

cat response-headers.txt

Check the HTTP status, response body, Retry-After, and any provider-specific remaining/reset headers. Avoid printing secrets in shared logs.

Node.js

import { setTimeout as sleep } from "node:timers/promises";

const endpoint = "https://api.example.com/screenshot";
const maxAttempts = 5;

async function capture(url) {
  for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
    const query = new URLSearchParams({
      url,
      api_key: "YOUR_API_KEY",
    });
    const response = await fetch(`${endpoint}?${query}`);

    if (response.ok) {
      return Buffer.from(await response.arrayBuffer());
    }

    const bodyText = await response.text();
    if (response.status !== 429) {
      throw new Error(`HTTP ${response.status}: ${bodyText.slice(0, 500)}`);
    }

    if (attempt === maxAttempts - 1) {
      throw new Error(`Retry limit reached: ${bodyText.slice(0, 500)}`);
    }

    const retryAfter = Number(response.headers.get("retry-after"));
    const fallback = Math.random() * Math.min(60_000, 1000 * 2 ** attempt);
    const delayMs = Number.isFinite(retryAfter) && retryAfter >= 0
      ? Math.min(60_000, retryAfter * 1000)
      : fallback;
    await sleep(delayMs);
  }
  throw new Error("Capture did not complete");
}

const bytes = await capture("https://example.org");
await import("node:fs/promises").then(({ writeFile }) =>
  writeFile("capture.png", bytes)
);

This example uses a generic endpoint because the title is provider-neutral. Replace the request and response parsing with the provider’s documented API, especially for distinguishing temporary throttling from quota exhaustion.

Common errors and fixes

Symptom Likely cause Fix
concurrency_limit_reached For ScreenshotOne, the current one-minute request bucket has no remaining starts. Queue work, check concurrency.remaining, and resume after concurrency.reset when it is zero.
HTTP 429 with an unclear message The status may represent a rate window, monthly quota, balance, or spending limit. Read the body and inspect the account usage endpoint or dashboard before retrying.
Retries keep receiving 429 Retries are immediate, synchronized across workers, or the account quota is depleted. Use Retry-After or bounded backoff with jitter; stop if the response identifies an account limit.
Worker pool stays idle after errors The client may have dropped rejected jobs instead of preserving them, or reset time was parsed incorrectly. Keep jobs pending, validate reset units and time zone, and alert on jobs that exceed their retry budget.
Throughput is lower after adding workers More simultaneous workers can create bursts that exhaust a request-start window and add retry delays. Shape request starts using provider capacity; tune from observed responses and documented limits.
429 persists after waiting A monthly quota, balance, or spend limit does not necessarily clear when a short rate window resets. Check billing and usage, then change plan or wait for the provider’s documented quota reset as appropriate.
Retries produce duplicate processing A timeout may have happened after the provider accepted or completed a request. Track job IDs and completion state; use provider idempotency support if documented.

Performance, reliability, and cost

Performance

Concurrency is only one part of throughput. A request-start limit, render-duration variation, page load time, and downstream storage can each constrain the pipeline. Increase workers gradually while watching accepted starts, throttles, queue depth, and completion time. If the provider exposes a remaining count and reset, those values are better scheduling inputs than an assumed fixed number of workers.

Reliability

Keep queue state outside a short-lived process when losing jobs on restart would matter. Apply finite retries and dead-letter or otherwise surface jobs that exhaust their retry budget. Preserve the provider response for diagnosis while redacting credentials. Treat temporary throttles and account-level limits as different states in monitoring so an exhausted quota does not trigger endless retries.

Cost

Check whether the provider bills requests started, successful captures, or another unit. A failed capture may still consume a rate-window start even when it does not consume a successful-capture allowance. ScreenshotEngine, for example, documents that requests are rate limited while failed requests do not consume its successful monthly capture allowance. Confirm current terms with your provider; do not infer billing from HTTP status.

What to compare when choosing a screenshot API

Compare these details in documentation and account settings:

  • Does the limit count request starts, active renders, successful captures, or multiple categories?
  • What is the enforcement window, and are bursts allowed within it?
  • Is there a separate monthly quota, balance, or spend ceiling?
  • Does the response expose remaining capacity, a reset time, and a request ID?
  • Do failed renders, cache hits, and retries count toward billing or usage?
  • Can bulk work be submitted asynchronously, and how are job completion and retries reported?

ScreenshotNeo is an option to try first: cookie banners, popups, and chat widgets are removed before capture, only clean shots are billed, and its lowest paid plan is $5 for 3,000 shots. Its responses identify page verdict and billing state, which helps distinguish clean captures from other outcomes. See the ScreenshotNeo API documentation for request and response details.

Or skip the browser setup

ScreenshotNeo is a screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its capture can remove cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Here is the one-call cURL request. Replace the target URL and API key. The API docs cover supported options and formats.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
await import('node:fs/promises').then(({ writeFile }) =>
  writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))
);

Sign up for ScreenshotNeo to get 1,000 screenshots a month free, with no card required.

FAQ

Does a fast screenshot render avoid the concurrency limit?

Not necessarily. ScreenshotOne’s documented limit counts requests started during its one-minute bucket, so a quickly completed request still uses a start. Another provider may count active renders instead.

Should I retry every 429?

No. First determine whether it is a temporary rate limit or an account quota, balance, or spend issue. Retrying cannot restore depleted account allowance.

Is concurrency the same as requests per minute?

Not always. Providers use the terms differently. Check whether the documented cap measures starts over time, active work, or both.

How many screenshot workers should I run?

There is no provider-neutral number. Use the current plan’s documented capacity and live remaining/reset values, then tune while monitoring throttles and queue depth.

Can a 429 mean a failed capture was still billable?

It depends on the provider’s billing rules. Check the response’s usage fields and the provider’s terms for failed captures, retries, and cache hits.