How to Deal With Rate Limits in a Bulk Website Screenshot API
Queue screenshot jobs, pace them against provider limits, and distinguish short-window throttling from a plan quota ceiling.
Put bulk screenshot work in a queue, check the provider’s current allowance, and start only as many captures as that allowance permits. When a short-window bucket is empty, retain the affected jobs and resume after its reset. When the plan quota is exhausted, pause the queue until capacity is restored or the plan changes. A bulk endpoint does not necessarily provide extra capacity, and rate-limit errors do not use one universal HTTP status.
The details below use ScreenshotOne’s documented behavior as a concrete example, with Urlbox as a second-provider comparison. Check the current documentation for whichever API you use: limits, error shapes, bulk execution, and quota rules are provider-specific.
1. Model the workload as a queue
Represent each capture as a durable job, rather than firing off every URL at once. At minimum, store:
- A stable job ID and target URL.
- The capture options needed to reproduce the request.
- An attempt count and next eligible attempt time.
- A terminal result: success, provider throttling, plan quota exhaustion, or target/render failure.
Keep unstarted jobs in the queue when the provider says to wait. Keep retryable failures there too, but classify them so a target-site failure does not get mistaken for provider throttling. For a single process, an in-memory queue can be enough for a starter integration; for durable work or multiple workers, use shared queue state so workers coordinate instead of each spending the same allowance. ScreenshotOne’s guide names Redis/BullMQ and SQS as options for multi-worker or durable retries.
2. Check provider capacity before draining work
For ScreenshotOne, the authenticated usage endpoint reports concurrency.remaining and concurrency.reset. The remaining value describes how many request starts are available in a one-minute bucket; it is not the number of renders currently active. Use the current response to cap starts, then use the reset value to decide when to resume. See the ScreenshotOne usage and rate-limit guidance and API documentation for the provider’s endpoint and response details.
The endpoint and response fields are provider-specific. Do not copy this field names or assume another service exposes the same usage API. If a provider does not expose live capacity or reset information, use its published limits, pace conservatively, and adapt based on its documented error responses.
Minimal pacing example in Python
This illustrative worker shows the control flow. Set the usage URL and authentication exactly as specified by your provider; the response keys shown are ScreenshotOne’s documented fields. It takes no more than the reported starts, and defers remaining jobs until the reset time.
import time
import requests
USAGE_URL = "SCREENSHOTONE_USAGE_ENDPOINT_FROM_CURRENT_DOCS"
CAPTURE_URL = "SCREENSHOTONE_CAPTURE_ENDPOINT_FROM_CURRENT_DOCS"
API_KEY = "YOUR_API_KEY"
jobs = [
{"id": "job-001", "url": "https://example.com"},
{"id": "job-002", "url": "https://example.org"},
]
session = requests.Session()
while jobs:
usage_response = session.get(
USAGE_URL,
headers={"Authorization": f"Bearer {API_KEY}"},
timeout=30,
)
usage_response.raise_for_status()
usage = usage_response.json()
remaining = int(usage["concurrency"]["remaining"])
reset_at = usage["concurrency"]["reset"]
if remaining <= 0:
# Parse reset_at according to the timestamp format in the API response.
# A bounded sleep loop avoids spinning while the same bucket is empty.
seconds_until_reset = max(1, seconds_until(reset_at))
time.sleep(seconds_until_reset)
continue
batch = jobs[:remaining]
del jobs[:len(batch)]
retry_later = []
for job in batch:
response = session.get(
CAPTURE_URL,
params={"access_key": API_KEY, "url": job["url"]},
timeout=90,
)
if response.ok:
save_capture(job["id"], response.content)
elif is_concurrency_limit_error(response):
retry_later.append(job)
elif is_transient_error(response):
retry_with_backoff(job)
else:
record_terminal_failure(job, response.status_code, response.text)
jobs.extend(retry_later)
seconds_until, save_capture, is_concurrency_limit_error, is_transient_error, and retry_with_backoff are application helpers; implement them against the exact timestamp and error schema in the provider’s current documentation. This outline deliberately avoids inventing a usage endpoint URL or response format beyond the documented field names.
3. Handle ScreenshotOne bulk requests correctly
ScreenshotOne’s bulk endpoint uses the same one-minute request bucket as regular screenshot requests. Wrapping many captures in one bulk call therefore does not mean unlimited capacity. Also, bulk requests are lazy-loaded by default: a screenshot is taken when its result URL is downloaded. Set execute: true when captures should execute before the bulk response, and allow enough time for that execution. An executed bulk response includes per-request status and error details, so inspect each item rather than treating a successful outer response as proof that every capture succeeded. See the bulk API documentation.
Bulk optimization applies only with execute: true and is not guaranteed, especially when a page reload is needed. Measure it on your own sites and options before depending on it for capacity planning.
4. Separate short-window limits from plan quota
| Signal | Meaning | Queue action |
|---|---|---|
concurrency_limit_reached |
ScreenshotOne’s current one-minute request bucket has no remaining starts. | Keep affected jobs and retry after concurrency.reset. |
screenshots_limit_reached |
The plan’s screenshot usage quota is exceeded. | Pause instead of retrying every minute; resolve the plan or allowed-limit configuration. |
ScreenshotOne documents concurrency exhaustion as an HTTP 400 response with its concurrency_limit_reached code. Urlbox documents HTTP 429 for too many requests or a reached rate limit. Handle the provider’s documented status and structured body; do not assume all rate limits are 429. See ScreenshotOne error documentation and Urlbox documentation.
5. Implement retries without creating a retry storm
- Retry only errors that can plausibly clear, such as a short-window limit or a transient network/service error.
- Use bounded exponential backoff with jitter for transient failures, and cap the number of attempts.
- For an explicit provider reset, schedule at that reset rather than repeatedly retrying sooner.
- For a plan quota error, stop automatic retries and alert an operator or wait for a known quota renewal event.
- Record each attempt and final result so a worker restart does not lose work or repeat completed captures.
ScreenshotOne says its API does not automatically retry requests, so retry ownership belongs in the integration. Its options documentation also says failed requests are not counted against rendering quota. Check the provider’s current rules before assuming either behavior applies elsewhere. See ScreenshotOne options and retry guidance.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Repeated concurrency_limit_reached |
The worker keeps submitting before the one-minute bucket resets, or multiple workers are spending the allowance independently. | Use the reported reset, share queue/limiter state between workers, and retain unstarted jobs. |
Repeated screenshots_limit_reached |
The plan quota is exhausted, rather than a temporary request bucket. | Pause the queue and resolve quota or plan configuration; a short sleep will not fix it. |
| HTTP 429 handling misses a provider limit | The integration assumes every provider uses 429. ScreenshotOne documents its concurrency error as HTTP 400 with a structured code. | Parse the provider’s documented status and body together. |
| Bulk call succeeds but some captures do not | The outer response was treated as the result for every item, or lazy execution was mistaken for completed rendering. | Inspect per-item status and errors; use execute: true if execution must happen before the response. |
| Bulk work seems not to start | Default lazy loading means the screenshot runs when its result URL is downloaded. | Download each result or request eager execution as documented. |
| Workers exceed allowance despite local pacing | Each process independently reads the same remaining allowance and starts work simultaneously. | Use a shared queue or distributed limiter, and reserve capacity atomically where possible. |
| A retry never improves the result | The problem is persistent at the target page, not provider throttling. | Classify target/render errors separately, inspect the response detail, and avoid retrying permanent failures indefinitely. |
7. Performance, reliability, and cost
- Throughput: Start work at the allowance the provider reports, not at an assumed active-render concurrency. Increase worker parallelism only while the shared limiter keeps starts within that allowance.
- Latency: A one-minute bucket can make a large queue wait between groups of starts. Use reset-aware scheduling so workers do useful work elsewhere instead of polling an exhausted bucket.
- Reliability: Persist jobs and per-item outcomes. A bulk envelope, an HTTP success, or a worker process completing does not by itself prove every page was captured successfully.
- Retries: Bound retries and distinguish transient provider errors from permanent target failures. ScreenshotOne states it does not retry automatically; integrations need an explicit retry policy.
- Cost: Do not infer billable usage from submitted bulk envelopes alone. Check the provider’s billing rules, especially for failed captures and lazy execution. ScreenshotOne documents that failed requests are not counted against rendering quota.
- Optimization: Treat bulk optimization as a possible efficiency, not a guaranteed capacity multiplier. ScreenshotOne says it applies with eager execution and may not help when pages reload.
8. Or skip the browser setup
If you would rather send one screenshot request than maintain a browser capture service, ScreenshotNeo is a website screenshot API and MCP server. Its one-call endpoint returns an image or PDF. For a bulk queue, still pace work according to the provider’s applicable limits and handle each response deliberately.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers say which outcome occurred. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Does a bulk endpoint bypass rate limits?
No general rule says it does. ScreenshotOne’s bulk endpoint consumes the same one-minute request bucket as regular screenshot requests.
Should every rate-limit response be retried?
No. Retry a short-window limit after its reset; pause on a plan quota ceiling. Treat persistent page failures separately.
Is HTTP 429 the standard rate-limit response?
No. Urlbox documents 429 for its rate limit, while ScreenshotOne documents concurrency exhaustion as HTTP 400 with a structured error code.
Does ScreenshotOne retry failed captures automatically?
No. Its documentation says the integration decides whether and how to retry.


