When Does an API Quota Reset? A Provider-by-Provider Guide
API quotas reset on different clocks. Learn how to identify the limit, read reset headers, and fix 429, insufficient_quota, and billing errors.
There is no universal API quota reset time. The answer depends on the provider, the quota dimension, and your account configuration. A request-rate limit may refill on a rolling window or synchronized minute; a daily quota may reset at a provider-defined timezone; a monthly usage or spend limit may reset only at the next billing cycle. Billing and prepaid-credit errors may not clear by waiting.
When a request fails, first identify the HTTP status, error type, quota dimension, and reset metadata. Honor Retry-After or provider-specific reset headers instead of guessing.
What “quota” can mean
| Quota or limit | Typical reset model | What to do when exhausted |
|---|---|---|
| Requests per minute (RPM) | Rolling or synchronized short interval | Throttle and retry after the advertised reset |
| Tokens per minute (TPM) | Rolling or synchronized short interval | Reduce concurrency, token size, or wait for reset |
| Requests per day (RPD) | Provider-defined daily boundary | Wait for that provider’s timezone boundary |
| Approved monthly usage | Monthly account cycle | Check the organization limit and billing state |
| Spend limit | Configured cap or monthly cycle | Raise or remove the cap, or wait if intentionally enforced |
| Prepaid balance | No automatic rate reset | Add credits or correct billing |
How to determine your reset time
- Capture the complete response. Save the status code, JSON body, response headers, and request ID.
- Classify the failure. Distinguish a temporary 429 or rate-limit error from insufficient quota, exhausted credits, or a spend cap.
- Read reset metadata. Look for
Retry-After, provider reset headers, or a Unix timestamp from a rate-limit endpoint. - Check scope. The limit may apply to an API key, project, organization, model, or resource group.
- Convert the boundary. Convert the provider’s timestamp or timezone to the timezone used by your operations team.
- Fix non-time-based causes. Add prepaid credits, change a spend limit, request a higher quota, or select the correct project.
Generic cURL inspection
curl -i https://api.example.com/endpoint \
-H "Authorization: Bearer $API_KEY"
Inspect Retry-After, remaining-capacity headers, reset headers, and the error body. Do not retry every 4xx response automatically.
Python: log status, headers, and error details
import requests
response = requests.get(
"https://api.example.com/endpoint",
headers={"Authorization": f"Bearer {YOUR_API_KEY}"},
timeout=30,
)
print("status:", response.status_code)
for name, value in response.headers.items():
if "rate" in name.lower() or name.lower() in {"retry-after", "x-request-id"}:
print(name, value)
try:
print(response.json())
except ValueError:
print(response.text)
Node.js: inspect reset headers
const res = await fetch('https://api.example.com/endpoint', {
headers: { Authorization: `Bearer ${process.env.API_KEY}` }
});
console.log('status:', res.status);
for (const [name, value] of res.headers) {
if (name.includes('rate') || name === 'retry-after' || name === 'x-request-id') {
console.log(name, value);
}
}
console.log(await res.text());
OpenAI API: rate limits, credits, and monthly limits
OpenAI has several separate controls. Temporary request or token rate limits expose remaining capacity and reset countdowns through headers such as x-ratelimit-remaining-requests, x-ratelimit-remaining-tokens, x-ratelimit-reset-requests, x-ratelimit-reset-tokens, and x-ratelimit-reset-project-tokens. A temporary 429 may also include Retry-After. Read the values from the response rather than assuming a fixed minute. See the OpenAI rate-limit guide.
OpenAI also separates prepaid credits, organization usage limits, and organization or project spend limits. OpenAI states that each organization has an approved monthly usage limit. An insufficient_quota-style error can therefore mean the account has exhausted credits or a configured spend limit, rather than an ordinary burst-rate limit. Retrying will not restore access; check the organization Limits page, add prepaid credits when the balance is exhausted, or change the relevant spend limit.
OpenAI handling pattern
import time
import requests
url = "https://api.openai.com/v1/responses"
headers = {"Authorization": f"Bearer {YOUR_API_KEY}"}
payload = {"model": "gpt-4.1-mini", "input": "ping"}
r = requests.post(url, headers=headers, json=payload, timeout=60)
print(r.status_code, r.json() if r.content else "")
if r.status_code == 429:
retry_after = r.headers.get("retry-after")
reset_requests = r.headers.get("x-ratelimit-reset-requests")
reset_tokens = r.headers.get("x-ratelimit-reset-tokens")
print({"retry_after": retry_after, "reset_requests": reset_requests, "reset_tokens": reset_tokens})
Gemini API: RPM, TPM, and RPD
Gemini measures limits independently as requests per minute (RPM), tokens per minute (TPM), and requests per day (RPD). Exceeding one dimension can produce a rate-limit error even when the other dimensions have capacity. Limits apply per project rather than per API key. Google documents that RPD quotas reset at midnight Pacific Time. See the Gemini rate-limits documentation.
For an RPM or TPM failure, wait for the short-term limit shown by the response or dashboard. For RPD, convert midnight Pacific Time to your operating timezone. If the project still fails after the documented boundary, verify that the request is using the intended project and model and inspect current usage in Google AI Studio.
Google Cloud APIs: service-specific intervals
Google Cloud does not have one reset rule for every API. Quota intervals are predefined per service. Compute Engine uses synchronized one-minute intervals for rate quotas. If a project reaches its limit at 10:00:15, capacity can return at the next boundary, such as 10:01:00, rather than exactly 60 seconds later. The documented failure can be a 403 with reason rateLimitExceeded. See Compute Engine rate quotas.
gcloud compute regions describe REGION \
--format="yaml(quotas)"
Use the quota page for the specific Google Cloud service, because another service may use a different interval or scope.
GitHub API: read the resource reset timestamp
GitHub exposes resource-specific reset times through its rate-limit endpoint. REST and GraphQL use separate rate-limit systems, so inspect the resource and API family involved. The reset value is a Unix timestamp.
curl -H "Authorization: Bearer $GITHUB_TOKEN" \
https://api.github.com/rate_limit
python - <<'PY'
import os, time, requests
r = requests.get(
"https://api.github.com/rate_limit",
headers={"Authorization": f"Bearer {os.environ['GITHUB_TOKEN']}"},
timeout=30,
)
data = r.json()
for resource, values in data.get("resources", {}).items():
print(resource, values.get("remaining"), time.ctime(values.get("reset", 0)))
PY
Wait until the resource’s reset timestamp, and make sure you are not confusing core REST capacity with search, GraphQL, or another resource bucket.
Why waiting sometimes does not help
- You hit a monthly or spend limit. A short retry delay cannot clear an account cap.
- Your prepaid balance is empty. Add credits before retrying.
- You are checking the wrong scope. A different project, organization, model, or resource may be limited.
- The boundary is not your local midnight. Gemini RPD uses Pacific Time; synchronized cloud intervals use provider boundaries.
- Retries are consuming the remaining capacity. Unbounded concurrency can keep the limiter exhausted.
- Dashboard charts are delayed or aggregate data. Confirm with the live response and the provider’s usage view.
Reliable retry design
- Retry only transient rate-limit responses, not authentication, validation, billing, or hard-quota errors.
- Honor
Retry-Afterwhen present. - Otherwise use the provider reset value; add small random jitter so workers do not retry together.
- Use exponential backoff with a maximum delay and a maximum attempt count.
- Reduce concurrency and token/request size when TPM or RPM is the constrained dimension.
- Record request IDs, status, error code, project, model, and reset metadata for diagnosis.
import random, time
def retry_delay(response, attempt):
retry_after = response.headers.get("retry-after")
if retry_after:
try:
return float(retry_after)
except ValueError:
pass
return min(60, 2 ** attempt) + random.uniform(0, 0.5)
for attempt in range(6):
response = make_request()
if response.status_code != 429:
break
time.sleep(retry_delay(response, attempt))
Performance, reliability, and cost notes
- Throttle before the provider does: a token bucket or semaphore prevents bursts and improves tail latency.
- Track RPM, TPM, RPD, error rate, and spend separately; one green metric does not prove all quotas are available.
- Cache safe, repeatable responses to avoid spending quota on identical work.
- Use queues for daily quotas and schedule work after the documented boundary.
- Do not assume a reset makes capacity infinite; the next burst can exhaust it again.
- Billing errors need a billing action. Repeated retries add load and can increase costs where requests are billable.
Or skip the browser setup
If your workflow needs website screenshots while you are already managing API quotas, ScreenshotNeo provides a single GET request that returns a PNG, JPEG, WebP, or PDF. The API accepts the cookie or consent banner like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account.
Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| 429 with a reset header | Temporary RPM or TPM limit | Honor the reset value, reduce concurrency, retry with jitter |
| 429 after the daily boundary | Wrong project, model-specific quota, or dashboard lag | Verify project and model, inspect live usage, then contact provider support if needed |
insufficient_quota |
Approved monthly usage, spend cap, or credits | Check limits and billing; waiting alone may not work |
403 rateLimitExceeded on Google Cloud |
Synchronized service quota exhausted | Wait for the next service interval or request more quota |
| GitHub calls still fail after reset | Different resource or API family is limited | Inspect each resource in /rate_limit |
FAQ
Does every API reset at midnight?
No. Midnight applies only when the provider documents a daily boundary, and the timezone may not be yours.
How long should I wait after a 429?
Use Retry-After or the provider’s reset countdown. If neither is present, apply bounded exponential backoff and reduce load.
Why do I still get an error after waiting?
You may have exhausted credits, a monthly usage allowance, or a spend cap. Check billing and account limits before retrying.
Is a quota attached to my API key?
It depends on the provider. Gemini documents project-level limits, while GitHub exposes resource buckets and OpenAI has organization and project controls.
Can I predict a reset from the time of my last request?
Only when the provider documents a rolling window. Synchronized intervals and daily or monthly boundaries require the provider’s own timestamp or timezone.


