Guide
API Rate Limiting: Algorithms, Headers and 429
Rate limit an HTTP API: fixed window, sliding window, token bucket and leaky bucket compared, plus 429, Retry-After, RateLimit headers, nginx and Express.
On this page
- Algorithms and their failure modes
- Who to count: key, IP or route
- Where to enforce
- The response contract
- Advertising the limit: X-RateLimit-* versus RateLimit
- Real configuration
- nginx: limit_req
- Express: express-rate-limit v7
- Redis token bucket
- Cloudflare rate limiting rules
- Client side: honor, then back off
- Verify
TL;DR: Pick the algorithm by what you want to do with bursts (token bucket allows them, leaky bucket smooths them, fixed window lets 2x through at the boundary), identify callers by API key before IP, reject with
429plusRetry-After, and make clients back off with jitter. In nginx, rememberlimit_req_status 429;because the default is503.
Algorithms and their failure modes
Every limiter answers one question per request: has this caller used more than its allowance in the recent past? The algorithms differ in how they define “recent” and what they remember.
| Algorithm | State per caller | Allows bursts? | Characteristic flaw |
|---|---|---|---|
| Fixed window | one counter + window start | Yes, up to 2x at a boundary | Boundary burst |
| Sliding window log | timestamp of every request | No | Memory grows with the limit |
| Sliding window counter | two counters | Slightly | Approximation, not exact |
| Token bucket | token count + last refill time | Yes, up to bucket size | Burst size is a second knob to tune |
| Leaky bucket | queue depth + last drain time | No, it queues instead | Adds latency, or drops when the queue is full |
Fixed window. Count requests in aligned windows (for example, per calendar minute) and reject once the count passes the limit. It is one INCR plus EXPIRE in Redis. The flaw is the boundary: with a limit of 100 per minute, a client can send 100 requests at 12:00:59 and another 100 at 12:01:00. Both windows are within the limit, yet 200 requests arrived within two seconds. If your backend capacity is sized for 100 per minute, the effective peak is double.
Sliding window log. Store a timestamp for every accepted request, drop those older than the window, and compare the remaining count to the limit. It is exact and has no boundary burst, but memory per caller is proportional to the limit, which gets expensive at 10,000 requests per hour per key.
Sliding window counter. Keep the current and previous fixed-window counts and estimate the sliding count as previous * (1 - elapsed_fraction) + current. At 25% into the current minute with 80 requests last minute and 10 so far, the estimate is 80 * 0.75 + 10 = 70. Two integers per caller and no boundary spike; the error comes from assuming the previous window’s requests were evenly spread.
Token bucket. A bucket holds up to capacity tokens and refills at rate tokens per second. Each request spends a token (or more, for expensive endpoints). A caller idle for a while accumulates a full bucket and can burst up to capacity, then is held to rate. You tune two numbers independently: sustained rate and burst size. This is what most API providers expose, and it is the easiest model to explain to customers.
Leaky bucket. Requests enter a queue that drains at a constant rate; a full queue rejects. The output is perfectly smooth, which protects fragile backends, at the cost of added latency for queued requests. The nginx documentation describes limit_req as the leaky bucket method.
Who to count: key, IP or route
- Per API key or user ID is the right identity for authenticated APIs. It is stable across devices and is the unit you bill and support.
- Per IP is the only option for unauthenticated routes (login, signup, password reset), but one office or mobile carrier NAT can host thousands of users. Set per-IP limits high, and behind a proxy make sure you key on the real client address rather than the proxy’s. In Express that means configuring
trust proxycorrectly, or every client appears to come from the load balancer. - Per route (or per cost class) protects expensive endpoints. A search or export endpoint deserves its own tighter bucket, and a login route deserves a per-account limit on top of the per-IP one, otherwise a botnet rotating IPs can still brute-force one account.
Layer them: a generous global per-IP limit at the edge, a per-key limit at the gateway, and a tight per-route limit where the work is expensive.
Where to enforce
- Edge or CDN (Cloudflare, CloudFront with WAF). Rejects abusive traffic before it reaches your network. Counting is coarse and keyed on request properties, not your business identity.
- Gateway or reverse proxy (nginx, Envoy, Kong, API Gateway). Cheap, language-independent, and can key on a header such as
X-Api-Key. Several proxy nodes need a shared store, or you accept that each node counts separately. - Application. The only layer that knows plan tiers, tenants and endpoint cost. It spends a request that already reached your process, so it should sit behind an edge limit, not replace it.
The response contract
Reject with 429 Too Many Requests (RFC 6585 section 4) and tell the client when to come back:
HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 30
RateLimit-Policy: "burst";q=100;w=60
RateLimit: "burst";r=0;t=30
{"type":"https://api.example.com/problems/rate-limited","title":"Rate limit exceeded","status":429,"detail":"100 requests per 60 seconds. Retry in 30 seconds."}
Retry-After takes either delay seconds (30) or an HTTP-date (Mon, 04 Jan 2027 12:00:00 GMT), per RFC 9110 section 10.2.3. Seconds are easier for clients to get right because they do not depend on clock sync. See Retry-After for details, and 429 Too Many Requests for the status code itself. Use 503 only when the server as a whole is overloaded rather than one caller being over quota; see 503 Service Unavailable.
Advertising the limit: X-RateLimit-* versus RateLimit
Two header families exist, and neither is an RFC.
X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset are a convention popularised by large API providers. There is no specification, so Reset may be epoch seconds on one API and seconds-until-reset on another. Read each provider’s docs and do not write one parser that assumes. See X-RateLimit headers.
RateLimit and RateLimit-Policy come from the IETF httpapi working group draft draft-ietf-httpapi-ratelimit-headers. It is an Internet-Draft, not an RFC, and was at revision 11 (May 2026) when this page was reviewed. In that revision both are Structured Fields lists whose members carry a quoted policy name and parameters:
RateLimit-Policy: "burst";q=100;w=60,"daily";q=1000;w=86400
RateLimit: "burst";r=50;t=30
q is the quota, w the window in seconds, r the remaining quota and t the seconds until it resets. The syntax has been reshaped between draft revisions (earlier ones used separate RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset headers, and one revision used limit=, remaining=, reset= parameters on a single RateLimit header), so if you emit these headers, note which revision you implemented and expect to change it.
Real configuration
nginx: limit_req
limit_req_zone defines the key and the rate; limit_req applies it to a location.
http {
# 10 MB shared zone; 1 MB holds about 16,000 states (32-bit) or 8,000 (64-bit)
limit_req_zone $binary_remote_addr zone=perip:10m rate=10r/s;
# Default is 503. Without this line your throttling looks like an outage.
limit_req_status 429;
limit_req_log_level warn;
server {
listen 80;
server_name api.example.com;
location /v1/ {
limit_req zone=perip burst=20 nodelay;
error_page 429 = @rate_limited;
proxy_pass http://app_backend;
}
location @rate_limited {
default_type application/problem+json;
add_header Retry-After 1 always;
return 429 '{"type":"about:blank","title":"Too Many Requests","status":429}';
}
}
}
How burst and nodelay interact:
- No
burst(default 0): anything arriving faster than 10 per second is rejected immediately. Two requests 50 ms apart means the second one is rejected. burst=20withoutnodelay: excess requests are queued and released at the configured rate, so the 21st simultaneous request waits about two seconds. Clients see latency instead of errors.burst=20 nodelay: excess requests up to 20 are served immediately, but the slots they occupy free up at the configured rate, so a sustained flood is still held to 10 per second. This is the closest nginx gets to token bucket behaviour.delay=N(alongsideburst): the firstNexcess requests go through immediately and the rest are delayed.
Rejections appear in the error log at the level set by limit_req_log_level (default error; delays log one level lower):
limiting requests, excess: 20.450 by zone "perip", client: 203.0.113.7, server: api.example.com, request: "GET /v1/items HTTP/1.1", host: "api.example.com"
limit_req_dry_run on; counts and logs violations without rejecting, which is the safe way to size a new limit against real traffic. The $limit_req_status variable (PASSED, DELAYED, REJECTED, plus _DRY_RUN variants) can go in your access log format.
add_header skips most error statuses unless you add always, which is why the named location above includes it.
Express: express-rate-limit v7
In v7 the request cap option is limit, not the old max.
import express from 'express'
import { rateLimit } from 'express-rate-limit'
const app = express()
app.set('trust proxy', 1) // number of proxies in front of the app; get this right
const apiLimiter = rateLimit({
windowMs: 60 * 1000,
limit: 100, // per key per window
standardHeaders: 'draft-7', // RateLimit + RateLimit-Policy; 'draft-6' and 'draft-8' also accepted
legacyHeaders: false, // turn off the X-RateLimit-* headers
keyGenerator: (req) => req.get('x-api-key') ?? req.ip
})
app.use('/v1/', apiLimiter)
The library responds with 429 by default. standardHeaders takes 'draft-6' (separate RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset headers), 'draft-7' (combined RateLimit and RateLimit-Policy headers) or 'draft-8' (adds a named policy via the identifier option). Because the IETF text is still moving, the library exposes the drafts as separate values instead of one “standard”. The default store is in memory and per process; with more than one instance, pass a shared store (a Redis store package, for example) or each instance counts on its own.
Redis token bucket
A token bucket needs a read-modify-write that must be atomic, so run it as a Lua script. Redis runs a script without interleaving other commands.
-- KEYS[1] bucket key
-- ARGV[1] capacity, ARGV[2] refill tokens per second, ARGV[3] cost
local t = redis.call('TIME')
local now_ms = t[1] * 1000 + math.floor(t[2] / 1000)
local capacity = tonumber(ARGV[1])
local rate = tonumber(ARGV[2])
local cost = tonumber(ARGV[3])
local data = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(data[1])
local ts = tonumber(data[2])
if tokens == nil then
tokens = capacity
ts = now_ms
end
tokens = math.min(capacity, tokens + (now_ms - ts) / 1000 * rate)
local allowed = 0
local retry_ms = 0
if tokens >= cost then
tokens = tokens - cost
allowed = 1
else
retry_ms = math.ceil((cost - tokens) / rate * 1000)
end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now_ms)
-- an idle bucket is full again after capacity/rate seconds, so it can expire
redis.call('PEXPIRE', KEYS[1], math.ceil(capacity / rate * 1000) + 1000)
return { allowed, math.floor(tokens), retry_ms }
The caller maps the result: allowed == 0 becomes 429 with Retry-After: ceil(retry_ms / 1000), and the remaining count feeds your rate limit headers. Reading Redis TIME inside the script avoids clock skew between application servers. Decide in advance what happens when Redis is unreachable; for most APIs, failing open with an alert is less damaging than failing closed.
Cloudflare rate limiting rules
Cloudflare evaluates rules in the WAF, in front of your origin. A rule has these parts:
- Expression: which requests the rule matches, for example
(http.request.uri.path starts_with "/api/"). - Characteristics: what to count by. IP is the default, and higher plans can count by header, cookie, query parameter and more. “IP with NAT support” uses session cookies to avoid lumping users behind one NAT together.
- Period and requests: the counting window, chosen from 10 seconds, 1 minute, 2 minutes, 5 minutes, 10 minutes or 1 hour (availability depends on plan), and the number of requests allowed within it.
- Action and mitigation timeout: block, a challenge, or log only. Once triggered, the action applies for the mitigation timeout, which ranges up to 24 hours depending on plan.
- Response: blocked requests get
429by default; a custom status in the 400 range is supported.
A custom counting expression does not inherit the rule’s match expression. If you only want to count failed logins, restate those conditions in the counting expression yourself.
Client side: honor, then back off
Clients should treat 429 as an instruction, not an error:
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms))
async function fetchWithBackoff(url, options, { retries = 5, baseMs = 500, capMs = 30000 } = {}) {
for (let attempt = 0; ; attempt++) {
const res = await fetch(url, options)
if ((res.status !== 429 && res.status !== 503) || attempt >= retries) return res
const retryAfter = res.headers.get('retry-after')
let delayMs
if (retryAfter !== null) {
const seconds = Number(retryAfter)
delayMs = Number.isFinite(seconds)
? seconds * 1000
: Math.max(0, Date.parse(retryAfter) - Date.now())
} else {
// full jitter: uniform in [0, min(cap, base * 2^attempt)]
delayMs = Math.random() * Math.min(capMs, baseMs * 2 ** attempt)
}
await sleep(delayMs)
}
}
Jitter matters because without it every client throttled at the same instant retries at the same instants and keeps re-tripping the limit. Only retry requests that are safe to repeat: GET, PUT and DELETE are idempotent by definition, but a POST needs an idempotency key (see REST API design with HTTP semantics). Also cap the total time spent; a worker that sleeps for an hour on a Retry-After is rarely what the caller wanted.
Verify
Send a burst against your limiter and watch the status codes and headers:
for i in $(seq 1 30); do
curl -s -o /dev/null -w "%{http_code} " https://api.example.com/v1/items
done; echo
curl -si https://api.example.com/v1/items | grep -iE '^(HTTP|retry-after|ratelimit|x-ratelimit)'
If you see 503 instead of 429 from nginx, limit_req_status is missing. If you see a 429 your application never sent, an edge layer produced it; 429 Too Many Requests fix explains how to tell who sent it.
Frequently asked questions
Which rate limiting algorithm should I use for an API?
Use a token bucket when you want to allow short bursts but cap the sustained rate, which suits most public APIs. Use a leaky bucket (nginx limit_req) when the backend needs a smooth, constant arrival rate. Use a sliding window counter when you must advertise a hard "N per minute" and cannot tolerate the fixed-window boundary burst. A plain fixed window is acceptable for coarse daily quotas.
Why does nginx return 503 instead of 429 when I rate limit?
limit_req_status defaults to 503. Add limit_req_status 429; in the http, server or location block to return 429 Too Many Requests. Until you do, clients and monitoring will read your throttling as a server outage.
Is Retry-After required on a 429 response?
No. RFC 6585 says a 429 response MAY include Retry-After, and RFC 9110 section 10.2.3 defines its value as either a number of seconds or an HTTP-date. It is still the single most useful header you can add, because it lets well-behaved clients wait the right amount instead of guessing.
Are the RateLimit and RateLimit-Policy headers a standard?
Not yet. They come from the IETF httpapi working group draft draft-ietf-httpapi-ratelimit-headers, which was still an Internet-Draft (revision 11, May 2026) at the time of writing, and its field syntax has changed between revisions. The X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset headers are a de facto convention with no specification, and their exact semantics (for example whether Reset is epoch seconds or seconds remaining) differ per provider.
Should I rate limit per IP or per API key?
Per authenticated key or user when you have one, because it is the only identity that survives NAT, mobile carriers and shared proxies. Use per-IP limits as a coarse outer layer for unauthenticated routes such as login and signup, and set the limit high enough that a corporate NAT with many users behind one address is not locked out.
How should a client retry after a 429?
Honor Retry-After when it is present. When it is absent, wait a random time between zero and min(cap, base * 2^attempt), which is exponential backoff with full jitter, and give up after a bounded number of attempts. Retry non-idempotent requests such as POST only if they carry an idempotency key.
Sources
- RFC 6585 section 4: 429 Too Many Requestsrfc-editor.org
- RFC 9110 section 10.2.3: Retry-Afterrfc-editor.org
- IETF draft: RateLimit header fields for HTTPdatatracker.ietf.org
- MDN: 429 Too Many Requestsdeveloper.mozilla.org
- nginx: ngx_http_limit_req_modulenginx.org
- express-rate-limit documentationgithub.com
- Cloudflare: Rate limiting rulesdevelopers.cloudflare.com
Related
HTTP 429 Too Many Requests: Causes, Retry-After and Fixes
Fix HTTP 429 Too Many Requests: read Retry-After and RateLimit headers, back off with jitter, and set up express-rate-limit v7, nginx limit_req and Cloudflare.
429 Too Many Requests: Fix It with Retry-After and Backoff
Fix 429 Too Many Requests as an API client: honor Retry-After, back off exponentially with jitter, read RateLimit headers, handle GitHub and LLM limits.
Retry-After
Learn how the Retry-After header tells clients how long to wait before retrying a request. Understand its use with 503, 429, and 301 status codes.
X-RateLimit Headers
Learn how X-RateLimit headers inform API clients about rate limits, remaining requests, and reset times. Implement proper rate limiting in your applications.