What Is Rate Limiting? Everything You Need to Know
Learn how rate limiting works, choose the right algorithm and identity key, return HTTP 429 correctly, and build reliable limits for distributed APIs.
Rate limiting restricts how many requests a client, user, API key, IP address, tenant or other counting key can make during a period. It protects capacity, allocates access fairly and reduces abuse. When a client exceeds a limit, the usual HTTP response is 429 Too Many Requests.
There is no universal limit value or single best algorithm. Choose the counting identity, time model, enforcement location and response behavior for the resource you are protecting.
What is rate limiting?
Rate limiting is a policy that accepts, delays or rejects requests according to how frequently an identified key has used a service. A limit can apply to one endpoint, an entire API, a user account, an API key, an IP address, a tenant or a combination of dimensions.
RFC 6585 defines 429 as the response for a user that has sent too many requests in a given amount of time. The RFC does not prescribe how a server identifies a user or counts requests. A server may count per resource, across a server or across a group of servers, and may identify a caller with credentials or a stateful cookie.
Why rate limiting matters
- Capacity protection: expensive endpoints cannot consume all workers, database connections or third-party quota.
- Fairness: one customer cannot monopolize a shared service.
- Abuse resistance: limits slow scraping, credential attacks, spam and resource exhaustion.
- Cost control: request volume can drive metered infrastructure or provider charges.
- Predictability: clients receive a defined policy instead of random timeouts during overload.
How rate limiting works
- Identify the request with one or more keys, such as an API key, account ID, IP address, HTTP method and route.
- Look up the key’s current usage in a counter, window, queue or token balance.
- Decide whether the request is allowed, delayed or rejected.
- Atomically update the state when necessary.
- Return a success response or a 429 response with useful retry information.
Rate limit terminology
| Term | Meaning |
|---|---|
| Rate | How quickly capacity replenishes, such as 10 requests per second. |
| Burst | How many requests may pass immediately above the steady rate. |
| Window | The interval used to count requests. |
| Quota | A total allowance over a longer period, such as a day or billing cycle. |
| Key | The identity whose requests share a counter or token balance. |
| Retry-After | A response header telling a client when it may try again. |
What does HTTP 429 mean?
429 Too Many Requests means the server is throttling a caller because it has exceeded a configured request frequency. The response should explain the condition. RFC 6585 says a response may include Retry-After, either as seconds or an HTTP date, and 429 responses must not be stored by a cache.
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 12
{"error":"rate_limited","message":"Try again later."}
Do not use 429 for every overload condition. Use it when the caller’s request frequency is the reason for rejection. Return an appropriate 5xx response for an internal failure or an unavailable dependency.
Rate-limiting algorithms
Fixed window
A fixed window increments one counter for a defined interval and resets it at the boundary. It is simple and inexpensive, but it permits a boundary burst: a client can spend its full allowance at the end of one window and again at the start of the next.
window_id = floor(current_time / 60 seconds)
key = "api:" + client_id + ":" + window_id
count = increment(key)
set_expiry_if_new(key, 60 seconds)
allow if count <= 100
Sliding window
A sliding window counts requests over the previous N seconds, reducing fixed-window boundary artifacts. An exact implementation stores timestamps; an approximate implementation combines adjacent buckets. Exact tracking uses more storage and work than a single counter.
Token bucket
A token bucket replenishes a finite balance at a steady rate. Each request consumes one or more tokens. The bucket capacity controls bursts while the refill rate controls the long-term average. AWS API Gateway documents token-bucket throttling with rate and burst settings.
elapsed = now - last_refill
tokens = min(capacity, tokens + elapsed * refill_rate)
if tokens < request_cost:
reject
else:
tokens -= request_cost
allow
Leaky bucket
A leaky bucket places work in a queue and releases it at a controlled pace. It is useful when smoothing output matters more than allowing bursts. The queue needs an explicit maximum; otherwise waiting work can consume memory and increase latency.
| Algorithm | Burst behavior | Strength | Trade-off |
|---|---|---|---|
| Fixed window | Boundary bursts possible | Very simple | Less precise short-term control |
| Sliding window | More even | Accurate moving-period enforcement | More state or approximation logic |
| Token bucket | Controlled bursts | Clear rate and burst controls | Requires careful atomic updates |
| Leaky bucket | Queue smooths output | Predictable processing pace | Queues add latency and need bounds |
How to choose the counting key
The key should match the abuse or fairness problem. Use an authenticated identity for customer quotas, an API key for developer access, an endpoint key for expensive operations and an IP address as a coarse signal for unauthenticated traffic.
Login protection
OWASP guidance recommends independent controls for attempts against each username and attempts from each source IP (or IP plus ASN). A single combined IP-plus-username counter can let an attacker spread attempts across many usernames.
IP and NAT effects
IP-only limits can group many legitimate people behind one corporate proxy, mobile carrier or home router. Cloudflare warns that shared NAT addresses can create false positives. Combine suitable dimensions and monitor legitimate rejection rates.
| Key | Good use | Risk |
|---|---|---|
| API key or user ID | Customer quotas and billing | Unauthenticated callers have no stable key |
| IP address | Anonymous abuse signal | NAT and proxies group unrelated users |
| Tenant ID | Fairness between organizations | One noisy user can affect a tenant |
| Route plus identity | Different limits for cheap and expensive operations | More counters and configuration |
| Username plus source controls | Login attack resistance | Requires separate counters and careful privacy handling |
Where to enforce a rate limit
- Edge or WAF: blocks abusive traffic before it reaches your application. Cloudflare rate-limiting rules can match request characteristics, count traffic and take an action after a threshold.
- API gateway: applies consistent policies by API, stage, method, account, region or API key. AWS API Gateway supports these scopes, but its documentation describes throttles and quotas as best-effort targets rather than guaranteed ceilings.
- Application middleware: understands business identity and response outcomes, making it suitable for user and endpoint-specific rules.
- Shared datastore: Redis or another atomic store coordinates counters across service instances and regions.
- Local process memory: is easy to start with, but load-balanced instances can each grant the full limit unless traffic is consistently routed.
For distributed services, update counters atomically. Redis documents fixed-window, sliding-window and token-bucket patterns, including Lua scripts for read-decide-update operations that prevent concurrent requests from double-spending capacity.
Implement a simple token bucket in Python
This runnable example limits one process to 5 requests per second with a burst of 10. Production deployments should store state in a shared system when multiple instances serve the same clients.
import time
from threading import Lock
class TokenBucket:
def __init__(self, rate, capacity):
self.rate = rate
self.capacity = capacity
self.tokens = capacity
self.updated = time.monotonic()
self.lock = Lock()
def allow(self, cost=1):
with self.lock:
now = time.monotonic()
self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)
self.updated = now
if self.tokens < cost:
return False
self.tokens -= cost
return True
bucket = TokenBucket(rate=5, capacity=10)
for request_number in range(20):
if bucket.allow():
print(request_number, "allowed")
else:
print(request_number, "429 Too Many Requests")
time.sleep(0.05)
Client retry behavior
A client should honor Retry-After, apply exponential backoff with jitter and stop after a bounded number of attempts. Immediate retries create a feedback loop that increases load.
async function fetchWithBackoff(url, options = {}, maxAttempts = 5) {
for (let attempt = 0; attempt < maxAttempts; attempt++) {
const response = await fetch(url, options);
if (response.status !== 429) return response;
const retryAfter = response.headers.get('retry-after');
const serverDelay = retryAfter && /^\d+$/.test(retryAfter)
? Number(retryAfter) * 1000
: 0;
const jitter = Math.random() * 250;
const delay = Math.max(serverDelay, 250 * 2 ** attempt) + jitter;
await new Promise(resolve => setTimeout(resolve, delay));
}
throw new Error('Rate limit persisted after retries');
}
Configuration checklist
- Define the protected resource and the failure you are preventing.
- Select stable identity dimensions and account for NAT, proxies and credential rotation.
- Set separate limits for cheap reads, expensive writes, authentication and background jobs.
- Choose burst tolerance and whether requests should queue or fail immediately.
- Make distributed updates atomic.
- Return 429 with a safe explanation and, when useful,
Retry-After. - Log decisions without exposing secrets or detailed attacker-useful counter state.
- Measure allowed requests, rejected requests, latency, queue depth and false positives.
- Document whether a threshold is a best-effort target or a contractual guarantee.
Performance, reliability and cost
A local counter has low latency but does not coordinate instances. A shared counter adds a network round trip and datastore cost, while improving consistency. Sliding-window timestamps consume more memory than fixed counters. Token buckets usually offer a practical balance when controlled bursts are acceptable.
Place cheap coarse controls at the edge and reserve detailed business rules for the application. Protect the limiter itself: bound key cardinality, expire inactive state, use timeouts and define behavior when the datastore is unavailable. Failing open preserves availability but can permit abuse; failing closed protects the dependency but may reject legitimate traffic. Choose per endpoint and record the decision.
There is no standards-based universal number. As one vendor-specific example, Cloudflare’s API limits page, updated August 25, 2026, listed 1,200 client API requests per five-minute period per user or account token. Treat that as a product quota example, not a recommendation for every API.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Clients receive 429 immediately after a deploy | Counters reset or are shared incorrectly | Use a stable key and verify window or token initialization. |
| Traffic doubles around minute boundaries | Fixed-window boundary burst | Use sliding windows or token buckets. |
| Users behind one office are blocked | IP-only identity | Add authenticated identity or another fair dimension. |
| Each instance allows the full quota | Process-local counters behind a load balancer | Use a shared atomic datastore or enforce at a coordinated gateway. |
| Retries make an outage worse | Immediate or synchronized retries | Honor Retry-After and add exponential backoff with jitter. |
| Attackers bypass login controls | One combined IP-plus-username bucket | Track username and source-IP controls independently. |
| Limiter latency spikes | Slow datastore or unbounded key state | Set datastore timeouts, expire keys and bound cardinality. |
Rate limiting screenshot workloads
Screenshot generation is a useful example because one URL capture can consume browser, CPU, memory and network capacity. Limit by API key or account for customer fairness, and add stricter endpoint limits for PDF, full-page and bulk jobs. Queue asynchronous work when a short burst should be accepted without starting every browser task at once.
Or skip the browser setup
If your goal is reliable screenshots rather than operating browser workers, ScreenshotNeo provides a GET endpoint that returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the full option set: full-page and element capture, device presets, dark mode, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agent, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture and usage data. Its MCP server provides take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
Plans include 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Is rate limiting the same as throttling?
The terms overlap. Rate limiting usually means rejecting or controlling requests after a threshold; throttling can also mean slowing or queuing work to shape traffic.
Should every 429 include Retry-After?
It is optional under RFC 6585, but it is useful when the server can calculate a safe retry time. Clients should still use bounded backoff when it is absent.
Can I rate-limit by both user and IP?
Yes. Independent dimensions often provide better abuse resistance and fairness than one combined counter. Apply the dimensions that match the endpoint’s threat model.
Are gateway limits exact guarantees?
Not always. AWS describes its throttles and quotas as best-effort targets. Read the documentation for the enforcement layer you operate and avoid promising stronger behavior than it provides.
What should happen when the limiter datastore is down?
Decide per endpoint whether availability or abuse protection has priority, use a short timeout and monitor the fallback path. Document the choice so an outage does not produce surprising behavior.


