How API Quotas Work for Image Generation Services
Learn how image API quotas work, diagnose 429 errors, choose retry strategies, and plan reliable, cost-aware generation workloads.
Direct answer: image-generation API quotas are provider-, model-, account-tier- and scope-specific. A service may limit requests, input or output tokens, generated images, spending, or several of these at once over minute, daily or rolling windows. The first limit you exhaust is the one that stops the request. Check your provider’s live limits and inspect the actual error before deciding whether to wait, slow down, add credits or change configuration.
There is no universal “images per day” number. OpenAI documents request, token and, for some models, image-per-minute dimensions. Gemini documents requests per minute, input tokens per minute, requests per day and image-per-minute limits for image-capable models. Values change with model, usage tier, account standing and current capacity. OpenAI rate-limit documentation and Gemini rate-limit documentation are the authoritative places to check.
What a quota measures
| Dimension | What it counts | Typical symptom |
|---|---|---|
| Requests per minute/day | API calls, regardless of prompt size | Small requests still receive 429 responses |
| Tokens per minute/day | Input and sometimes output tokens | Large prompts exhaust capacity first |
| Images per minute | Generated images, often model-specific | Parallel image jobs are throttled |
| Spend or credits | Money or prepaid balance in a time window | Waiting does not restore access |
| Daily allowance | Requests or images in a calendar window | Recovery occurs at the documented reset time |
These dimensions can apply simultaneously. For example, a batch can remain under a request limit while exceeding an image-per-minute limit, or use few requests but too many tokens. Plan against every documented dimension for the model you selected.
Scope and reset windows
Quota scope is provider-specific. Gemini states that limits apply per project rather than per API key, and that requests-per-day quotas reset at midnight Pacific time. OpenAI exposes applicable limits and reset timing through response headers in many rate-limited responses; some limits are organization- or project-scoped. Never assume that creating another key creates another quota.
Use the dashboard for the account and project that actually sends traffic. Gemini’s active limits are shown in AI Studio. OpenAI directs users to the limits area in account settings. Document the model, project or organization, tier and timezone alongside your application configuration so operators know which limit they are observing.
How to inspect the limit that applies
- Identify the exact model and endpoint producing the error.
- Open the provider’s live limits page for the active project or organization.
- Log the HTTP status, provider error code, request ID and relevant response headers.
- Record whether the limit is request, token, image, daily or spend based.
- Check the reset value or
Retry-Afterheader before retrying.
OpenAI rate-limit responses may include headers such as x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests and corresponding token fields. Published examples such as 60 permitted, 59 remaining and a one-second reset are illustrative header examples, not default quotas.
Why a 429 does not have one meaning
A 429 can mean temporary throttling, a traffic ramp that is too fast, exhausted prepaid credits, an organization or project spend limit, an assigned usage ceiling, or a provider-specific resource-exhausted condition. Read the error body and code; do not blindly replay every 429.
| Error situation | Correct response |
|---|---|
| Temporary rate limit | Honor Retry-After; otherwise use bounded exponential backoff with jitter. |
Traffic ramp / slow_down |
Reduce concurrency and increase spacing between requests. |
| Credits exhausted | Add credits or use the documented billing action before retrying. |
| Spend or usage ceiling | Change the approved limit or workload; retries alone cannot fix it. |
| Invalid or blocked image request | Change the prompt or parameters; do not replay unchanged. |
OpenAI’s image-generation guidance says to retry transient rate-limit and server failures with backoff, and not to automatically retry quota errors or image-generation user errors that require changing the request. Gemini documents 429 RESOURCE_EXHAUSTED for spend-based limits and recommends waiting briefly, reducing expensive-request rate or requesting an increase when normal workloads repeatedly hit the limit.
Safe retry code
The following language-neutral policy works around transient throttling without turning a billing problem into a retry storm:
for attempt in 0..MAX_RETRIES:
response = send_request()
if response.ok:
return response
if response.status == 429 and is_transient(response):
delay = retry_after(response) or exponential_backoff_with_jitter(attempt)
sleep(min(delay, MAX_TOTAL_WAIT_REMAINING))
continue
if response.status in [500, 502, 503, 504]:
sleep(exponential_backoff_with_jitter(attempt))
continue
raise NonRetryableError(response)
- Treat
Retry-Afteras a minimum wait, not a suggestion to send immediately. - Bound both attempts and total elapsed retry time.
- Use one retry layer. Official SDK retries plus application retries can multiply traffic.
- Add jitter so many workers do not retry simultaneously.
- Unsuccessful requests may still contribute to per-minute limits, so failed loops need the same controls as successful traffic.
Concurrency, throughput and cost planning
Estimate capacity with every active dimension: safe_rate = min(request_limit, image_limit, token_limit converted to requests). Keep concurrency below the smallest safe rate, then add a queue rather than launching unbounded promises. Separate interactive traffic from batch work so a bulk job cannot consume the allowance needed for user requests.
Measure requests, images, tokens, latency, 429s and retry time by model and project. Cache identical generations where policy permits, resize or reduce output variants when quality requirements allow, and stop retrying when the provider reports credits, spend or quota exhaustion. A retry that cannot succeed only increases cost and pressure on the same limit.
Gemini’s published spend-rate examples are $10 per 10 minutes for Tier 1, $50 per 10 minutes for Tier 2 and $200 per 10 minutes for Tier 3 where those limits apply. They are spend limits, not image counts or guaranteed entitlements. Gemini also states that specified rate limits are not guaranteed and actual capacity may vary.
Reliability checklist
- Store provider, model, project and tier with each job.
- Capture status, error code, request ID, remaining headers and reset data.
- Use a durable queue with maximum concurrency.
- Apply one bounded retry policy with jitter.
- Alert separately on throttling, credits, spend ceilings and invalid requests.
- Provide a user-visible state such as queued, retrying, failed or action required.
- Recheck limits after changing model, tier, project or billing.
Troubleshooting common quota errors
“I get 429 immediately at low traffic.”
Check whether the project already consumed its daily allowance, whether another service shares the organization quota, and whether the error is billing or spend related. Read the provider code before reducing request rate.
“Waiting did not help.”
You may have exhausted credits or a configured usage ceiling rather than a rolling rate window. Add credits or adjust the documented limit, then retry a new request.
“A second API key did not increase capacity.”
The quota may be project-, organization- or account-scoped. Gemini explicitly applies limits per project, not per key.
“Retries made the outage worse.”
Remove nested retry loops, honor Retry-After, cap attempts and reduce concurrency. Check whether the SDK already retries.
“Only image requests fail.”
Inspect image-per-minute and model-specific limits separately from general request limits. A model change can have different dimensions and values.
“The request is rejected every time.”
It may be an invalid or blocked image-generation request rather than transient throttling. Change the request according to the error instead of replaying it.
Or skip the browser setup
If your workflow also needs screenshots of generated-image pages, ScreenshotNeo provides a single website screenshot API call. Cookie banners, newsletter popups and chat widgets are removed before capture; bot checks, blank pages and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Create a free ScreenshotNeo account with 1,000 screenshots each month and no card.
FAQ
How many images can I generate per minute?
There is no cross-provider number. Check the selected model’s live image, request and token limits.
When does an API quota reset?
It depends on the dimension. Some limits use rolling windows; Gemini says requests-per-day quotas reset at midnight Pacific; OpenAI may return reset timing in headers.
Should every 429 be retried?
No. Retry only transient throttling or server failures. Fix credits, spend ceilings and invalid requests first.
Does changing API keys create more quota?
Not when the limit is scoped to a project, organization or account.
Can a provider increase my limit?
Sometimes, depending on tier and account status. Use the provider’s documented limit-increase process and do not promise that approval or capacity is guaranteed.


