How Many Screenshot API Requests Can Run at Once?
There is no universal concurrency limit. Learn how to distinguish requests per second, in-flight renders and monthly quotas, and how to pace requests safely.
Short answer: There is no universal number of screenshot API requests that can run at once. A provider may publish a requests-per-second limit, a true in-flight concurrency limit, a monthly screenshot quota, or some combination. These measure different things. If the documentation gives only a request rate, do not treat it as a guaranteed count of simultaneous renders.
For example, a service that accepts five requests per second has not necessarily promised that exactly five browser renders can be in flight. Requests take different amounts of time, and the provider may queue work, reject excess load, or apply account-specific limits. Check the endpoint and plan you actually use.
1. The three limits to distinguish
| Limit | What it measures | What it does not tell you |
|---|---|---|
| Request rate | How many HTTP requests you may submit in a time window, such as per second or minute. | How many screenshot renders may be running at the same moment. |
| In-flight concurrency | How many requests or browser renders may be active simultaneously. | How many captures you may make over a month. |
| Monthly quota | The number of screenshots or renders included in a billing period. | How fast you may submit them. |
Some services publish all three; others publish only a subset. A batch endpoint can put several URLs inside one HTTP submission, but that does not establish that the provider waives render-rate, concurrency, or monthly quota limits. Read the batch endpoint’s sizing and quota rules separately.
2. Published examples: rates are not concurrency guarantees
The following figures illustrate why limits must be compared by unit. They are provider-specific documentation, not a general limit for screenshot APIs.
| Provider and plan or endpoint | Published rate | Monthly allowance or other detail | Explicit in-flight count in cited docs? |
|---|---|---|---|
| Screenshot API Free | 1 request/second | 100 renders/month | No |
| Screenshot API Starter | 5 requests/second | 2,000 renders/month | No |
| Screenshot API Pro | 10 requests/second | 10,000 renders/month | No |
| Screenshot API Team | 25 requests/second | 25,000 renders/month | No |
| Screenshot API Business | 50 requests/second | 100,000 renders/month | No |
| Cloudflare Workers Paid Browser Rendering REST API | 10 requests/second (600/minute), increased from 3 requests/second (180/minute) on March 4, 2026 | Browser Sessions concurrency and new-browser limits are a separate topic | The cited changelog and screenshot endpoint docs do not specify a per-account simultaneous request count |
| Screenshot API .org free batch API | 60 requests/minute | 500 screenshots/month; batch endpoint submits multiple URLs in one request | No universal simultaneous-render count is established |
Screenshot API states that its request rate is a burst-control limit and its monthly renders are an independent plan allowance; its monthly quota resets at the start of each calendar month in UTC. Cloudflare’s changelog lists the REST rate and treats Browser Sessions limits separately. The batch API’s request allowance and screenshot quota likewise use different units. Consult the provider’s current plan and endpoint documentation before relying on any example rate.
Sources: Screenshot API documentation; Cloudflare March 4, 2026 changelog and Cloudflare screenshot endpoint documentation; Screenshot API .org documentation.
3. Find the limit that applies to your account
- Identify the exact provider, endpoint, account plan, and access method. A REST API, browser session product, and batch endpoint can have different limits.
- Find separate documentation for requests per second or minute, in-flight renders or browser sessions, and monthly captures.
- Inspect response headers and the account dashboard for account-specific quotas or rate-limit information. Do not assume every service uses the same header names.
- Check the provider’s documented behavior for throttling and saturation, including whether it returns
429,503, aRetry-Afterheader, or a job identifier. - Start below the published request rate, measure response times and errors, then increase traffic gradually while observing the provider’s guidance.
If documentation gives only a requests-per-second value, report that value as a submission rate. State that the in-flight concurrency limit is not specified instead of deriving a simultaneous-render number from it.
4. Pace requests and recover from throttling
A safe client uses bounded concurrency, a request-rate limiter, and retries that respect the provider’s instructions. These controls solve different problems: a concurrency cap limits active work, while a rate limiter spaces out submissions. For long-running captures, also set a timeout appropriate to the provider and your page complexity.
Runnable Python example
This example uses a semaphore to cap simultaneous requests and a simple spacing interval to limit request starts. Set the values to the limit documented for your provider. Replace the example URL and endpoint parameters with that provider’s documented API format.
import asyncio
import time
import httpx
API_URL = "https://api.example.com/screenshot"
API_KEY = "YOUR_API_KEY"
URLS = ["https://example.com", "https://example.org"]
MAX_IN_FLIGHT = 2
REQUESTS_PER_SECOND = 1
async def main():
semaphore = asyncio.Semaphore(MAX_IN_FLIGHT)
start_lock = asyncio.Lock()
last_start = 0.0
interval = 1.0 / REQUESTS_PER_SECOND
async with httpx.AsyncClient(timeout=90) as client:
async def capture(page_url):
nonlocal last_start
async with semaphore:
async with start_lock:
delay = interval - (time.monotonic() - last_start)
if delay > 0:
await asyncio.sleep(delay)
last_start = time.monotonic()
response = await client.get(
API_URL,
params={"access_key": API_KEY, "url": page_url},
)
if response.status_code == 429:
retry_after = response.headers.get("Retry-After")
raise RuntimeError(
f"Rate limited for {page_url}; Retry-After={retry_after}"
)
response.raise_for_status()
return page_url, response.content
results = await asyncio.gather(*(capture(url) for url in URLS))
for index, (page_url, image_bytes) in enumerate(results, start=1):
with open(f"shot-{index}.png", "wb") as image_file:
image_file.write(image_bytes)
print(f"Saved screenshot for {page_url}: {len(image_bytes)} bytes")
asyncio.run(main())
The sample deliberately surfaces a 429 rather than retrying immediately in a tight loop. In production, put failed jobs back on a delayed queue, parse and honor a valid Retry-After value when present, and cap retries. Do not retry permanent input errors as though they were transient.
cURL: one request at a time
cURL sends a single request here; it does not demonstrate or control a provider’s account-wide concurrency limit. Run calls from a queue or scheduler that enforces your provider’s published rate.
curl -G "https://api.example.com/screenshot" \
-H "Authorization: Bearer YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--output screenshot.png \
--write-out "HTTP %{http_code}\n"
Runnable Node.js example
This minimal example makes one capture and reports rate limiting. A multi-URL worker should add a queue, concurrency cap, request pacing, and delayed retry policy rather than launching an unbounded Promise.all.
const endpoint = new URL('https://api.example.com/screenshot');
endpoint.search = new URLSearchParams({
url: 'https://example.com',
}).toString();
const res = await fetch(endpoint, {
headers: { Authorization: 'Bearer YOUR_API_KEY' },
signal: AbortSignal.timeout(90000),
});
if (res.status === 429) {
console.error('Rate limited. Retry-After:', res.headers.get('retry-after'));
process.exitCode = 1;
} else if (!res.ok) {
throw new Error(`Screenshot request failed: HTTP ${res.status}`);
} else {
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('screenshot.png', bytes));
}
5. Troubleshooting common limit and load errors
| Symptom | Likely cause | What to do |
|---|---|---|
| HTTP 429 | The request rate or another account limit was exceeded. | Reduce request starts, inspect response headers, and honor Retry-After when the provider supplies it. Check whether the limit is per key, account, endpoint, or time window. |
| HTTP 503 or a provider-specific “busy” response | Renderers or browser capacity are temporarily saturated; this is not necessarily the same as a rate-limit violation. | Pause briefly and retry with a bounded backoff. For Screenshot API specifically, its docs describe 503 busy as renderer saturation and advise retrying after a short pause. |
| Requests succeed but captures arrive slowly | Pages may take different times to render, or the service may queue requests. A request rate alone does not guarantee completion throughput. | Measure end-to-end latency and queue depth. Reduce in-flight work if latency or failures climb, and consult provider-specific concurrency guidance. |
| HTTP 401 or 403 | Missing, invalid, expired, or insufficiently scoped credentials; sometimes an account or endpoint is not enabled. | Verify the key and authentication format against that provider’s docs. Avoid retrying unchanged credentials. |
| HTTP 400 or 422 | Invalid URL, unsupported parameter, malformed batch, or request format mismatch. | Validate inputs and the endpoint schema; retry only after fixing the request. |
| Client timeout | The page or renderer took longer than the client’s deadline, the service is queued, or the network stalled. | Set a realistic timeout, check provider status and job semantics, and retry only if the operation is safe to repeat. Prefer asynchronous jobs for workflows supported by the provider. |
| Monthly quota exhausted | The capture allowance was used even though the request rate is within limits. | Check quota headers or dashboard, wait for the reset if appropriate, or change plan. A lower request rate does not restore monthly quota. |
6. Performance, reliability, and cost
- Measure useful throughput: Track submitted requests, successful captures, latency percentiles, timeouts, 429s, saturation errors, and queue depth. Requests accepted per second are not the same as completed screenshots per second.
- Use bounded work: Keep a queue and limit active jobs. If errors or latency rise, reduce concurrency and let the queue drain before increasing traffic again.
- Make retries deliberate: Honor provider retry guidance, use capped backoff for transient failures, and avoid retry storms. For asynchronous APIs, retain job IDs and poll or accept callbacks as documented.
- Budget in the provider’s billing unit: Confirm whether the plan charges per successful render, submitted URL, batch item, or another unit, and whether failed or cached results count. Do not infer billing from the HTTP request count.
- Account for burst and month limits separately: A workload can stay under a monthly quota but hit a per-second cap, or stay under a rate cap and exhaust its monthly captures.
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its single-request API returns an image or PDF, and the API documentation lists capture options and usage details.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
- Cookie banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include
X-Page-VerdictandX-Billedheaders. - An MCP server lets AI agents using Claude, Cursor, or other MCP clients call
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
8. FAQ
Can I calculate concurrency from requests per second?
No, not without assumptions about render time and queuing. Use an explicitly documented in-flight limit or measure your workload while staying within the published rate.
Does batching mean I can exceed the screenshot quota?
No such conclusion follows from batching. It can reduce HTTP submissions, but check how the provider counts batch items and applies rate, concurrency, and monthly limits.
Is a 429 the same as a busy renderer?
Not necessarily. A 429 commonly indicates a limit was exceeded; some providers separately report renderer saturation, such as the documented 503 busy response for Screenshot API. Follow the specific service’s response guidance.
Which number should I put in my capacity plan?
Record the documented request rate, explicit in-flight limit if provided, monthly quota, and observed completion latency as separate values. If concurrency is undocumented, label it unknown and validate with measured traffic.
Sources
- Screenshot API documentation — plan request rates, monthly render allowances, quota reset, and error guidance.
- Cloudflare Browser Rendering changelog, March 4, 2026 — REST API request-rate change.
- Cloudflare screenshot endpoint — REST API and binding access details.
- Screenshot API .org docs — batch endpoint and separate request and screenshot allowances.


