Debug guide · you're seeing
429 Too Many RequestsHTTP/1.1 429 Too Many RequestsYou have exceeded a secondary rate limit. Please wait a few minutes before you try again.Rate limit reached for gpt-4o in organization org-xxxx on requests per min (RPM): Limit 3, Used 3, Requested 1. Please try again in 20s.Error 1015: You are being rate limited
429 Too Many Requests: Fix It with Retry-After and Backoff
Fix 429 Too Many Requests as an API client: honor Retry-After, back off exponentially with jitter, read RateLimit headers, handle GitHub and LLM limits.
On this page
TL;DR: The server is throttling your client. If the response has
Retry-After, wait at least that long. If not, retry with exponential backoff and full jitter, cap the attempts, and lower your steady request rate so you stop hitting the limit. Immediate retries make it worse and can turn a rate limit into a ban.
What it means
429 was added by RFC 6585 section 4 and is still the status servers use for “this user has sent too many requests in a given amount of time”. The response may include a Retry-After header and should explain the limit in the body:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 30
RateLimit-Policy: "burst";q=100;w=60
RateLimit: "burst";r=0;t=30
{"error":"rate_limited","message":"Rate limit exceeded. Try again in 30s."}
Throttling can be applied per API key, per IP address, per account, per endpoint or per token count, and several limits can apply at once. A client that was well below its per-minute limit can still hit a per-second burst limit.
Who sent it?
Several different layers return 429, and they need different fixes:
- The API’s own application layer: JSON body naming the limit, often with
Retry-AfterandX-RateLimit-*orRateLimit-*headers. Follow the vendor’s documentation. - An API gateway (AWS API Gateway throttling, Kong, Apigee): a generic “Too Many Requests” body, often without a
Retry-After. Gateways commonly allow bursts and a steady rate; slow your average. - A WAF or CDN in front of the origin: a Cloudflare page “Error 1015: You are being rate limited” with
Server: cloudflareand aCF-Ray; the origin never saw the request. These rules often key on IP and path, so changing nothing but your IP or user agent changes nothing about the real problem. - Your own nginx:
limit_reqreturns 503 by default, so a 429 from nginx means someone setlimit_req_status 429.
Print only the headers that matter:
curl -si https://api.example.com/items -H "Authorization: Bearer $TOKEN" \
| grep -iE '^(HTTP/|retry-after|ratelimit|x-ratelimit|x-github|server|cf-ray|via)'
Fix it, in order of likelihood
- Stop retrying immediately. Honor
Retry-After. - Add capped exponential backoff with jitter for the case where there is no
Retry-After. - Cap the number of attempts and the total wait, and surface the failure instead of looping forever.
- Reduce the request rate itself: limit client concurrency, batch calls, and cache.
- Check whether the 429 means “slow down” or “you have no quota”, and do not retry the latter.
- Share one limiter across all workers and processes that use the same credential.
Honor Retry-After
Retry-After is either a non-negative integer number of seconds or an HTTP date (RFC 9110 section 10.2.3). Handle both:
Retry-After: 120
Retry-After: Mon, 04 Oct 2027 09:30:00 GMT
Exponential backoff with full jitter
Delay for attempt n (starting at 0) is a random value between 0 and min(cap, base * 2^n). This is the “full jitter” strategy from the AWS Architecture Blog and performs better than plain exponential delay because it breaks the synchronization between clients.
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms))
function parseRetryAfter(value) {
if (!value) return null
if (/^\d+$/.test(value.trim())) return Number(value) * 1000 // delta-seconds
const date = Date.parse(value) // HTTP-date
if (!Number.isNaN(date)) return Math.max(0, date - Date.now())
return null
}
async function fetchWithRetry(url, init = {}, { retries = 5, baseMs = 500, capMs = 30_000 } = {}) {
for (let attempt = 0; ; attempt++) {
const response = await fetch(url, init)
if (response.status !== 429 || attempt >= retries) return response
const serverDelay = parseRetryAfter(response.headers.get('retry-after'))
const backoff = Math.random() * Math.min(capMs, baseMs * 2 ** attempt)
// Never wait less than the server asked for; add jitter on top of it
const delay = serverDelay !== null ? serverDelay + Math.random() * 1000 : backoff
await response.body?.cancel() // free the connection before waiting
await sleep(delay)
}
}
Requests with a streaming body cannot be replayed, so only use this wrapper with strings, Blob, FormData or URLSearchParams bodies. If your callers already pass an AbortSignal, pass it through init; an aborted fetch throws and exits the loop.
Python with requests and urllib3 already implements most of this. Retry honors Retry-After for 413, 429 and 503 by default and sleeps backoff_factor * 2**(n-1) seconds otherwise. Its delays are not jittered unless you set backoff_jitter (urllib3 2.x), and the sleep is not “full jitter” as described above:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
retry = Retry(
total=5,
status_forcelist=[429, 503],
backoff_factor=0.5,
backoff_jitter=0.5, # urllib3 2.x; without it the delays are deterministic
respect_retry_after_header=True, # the default
allowed_methods=None, # retry all methods; by default POST is excluded
)
session = requests.Session()
session.mount('https://', HTTPAdapter(max_retries=retry))
Setting allowed_methods=None retries POST on those statuses. For a 503 that can be unsafe if the server partially processed the request; use it only when the API is idempotent or accepts an idempotency key.
Slow down on purpose
Backoff reacts to a limit you already hit. Staying below it needs client-side pacing:
- Cap concurrency: a pool of 4 to 8 workers rather than firing 500
Promise.allcalls. - Use a token bucket sized to the documented rate, shared across processes (for example in Redis) if several workers use the same API key.
- Spread scheduled jobs; ten cron jobs starting at
:00produce a burst each hour. - Replace polling with webhooks or longer intervals, and use conditional requests (
If-None-Match) where the API supports them. Some APIs, including GitHub’s, document that a304 Not Modifiedconditional response does not count against the primary rate limit; check your API’s rules. - Cache responses that do not change and request only the fields you need.
Read the rate limit headers
Two generations of headers exist. The older X-RateLimit-* headers are widely used but never standardized:
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1790000000
The meaning of Reset varies: GitHub uses a Unix timestamp in seconds, other APIs use seconds remaining. The IETF HTTPAPI working group’s draft (draft-ietf-httpapi-ratelimit-headers, still a draft, not an RFC) defines structured-field headers instead:
RateLimit-Policy: "burst";q=100;w=60
RateLimit: "burst";r=40;t=15
The quoted string is a policy name that ties a RateLimit item to its RateLimit-Policy item. In the policy, q is the quota, w the window in seconds and the optional qu the unit (requests, content-bytes or concurrent-requests). In RateLimit, r is the remaining quota and t the effective window in seconds: the draft warns clients not to assume the whole quota is restored when t elapses. If a response carries both RateLimit and Retry-After, Retry-After takes precedence. Earlier revisions used RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset, which is a different syntax. Both are in the wild. When r is small, wait instead of sending another request.
Provider specifics
GitHub has two separate limits. Exceeding the primary limit returns 403 or 429 with x-ratelimit-remaining: 0; x-ratelimit-reset is the UTC epoch second when the window resets. Secondary limits (concurrency and request-rate abuse protection) also answer 403 or 429, with a message like You have exceeded a secondary rate limit. GitHub’s documented guidance is: if retry-after is present, wait that many seconds; if x-ratelimit-remaining is 0, wait until x-ratelimit-reset; otherwise wait at least a minute, and increase the wait on repeated failures. GitHub also tells clients to avoid concurrent requests and to pace requests that create content. Continuing to send requests while limited can get the integration banned.
LLM and AI APIs generally apply several limits at once: requests per minute, tokens per minute, and sometimes concurrent requests or a spending cap. Responses usually carry the remaining counts and reset time for each in provider-specific headers (for example the x-ratelimit-*-requests and x-ratelimit-*-tokens families, or anthropic-ratelimit-* headers together with retry-after). Two points apply to all of them. Retry the 429 only when it indicates a rate limit; some providers use the same status for an exhausted quota or billing limit, and no amount of backoff fixes that. And count tokens, not just requests: a handful of very large requests can exhaust a token-per-minute budget that many small ones would not.
Reproduce and verify
To confirm your client behaves, point it at a server you control that always returns 429. Any local server works. With Node:
// node 429-server.js
import http from 'node:http'
let hits = 0
http.createServer((req, res) => {
hits += 1
console.log(new Date().toISOString(), 'hit', hits)
if (hits < 4) {
res.writeHead(429, { 'Retry-After': '2', 'Content-Type': 'text/plain' })
return res.end('slow down')
}
res.writeHead(200, { 'Content-Type': 'text/plain' })
res.end('ok')
}).listen(8080)
curl -si http://localhost:8080/ | head -n 4
The log should show your client’s requests spaced at least two seconds apart (the Retry-After value) for the first three, and a success on the fourth. Spacing that is shorter than the header, or perfectly regular, means the code is ignoring Retry-After or has no jitter.
If you operate the API and want to produce a correct 429 yourself, return Retry-After and a body that names the limit, and log the key that triggered it so clients can tell their own throttling from a shared one.
Related
- 429 Too Many Requests covers the status itself and how servers implement rate limiting.
- Retry-After documents both value forms and its use with 503 and 3xx.
- X-RateLimit headers describes the legacy trio and how to read it.
- 503 Service Unavailable is the server-wide counterpart and also supports
Retry-After.
Frequently asked questions
How long should I wait after a 429?
Wait for the time in the Retry-After header if the response has one. It is either a number of seconds or an HTTP date, and RFC 9110 section 10.2.3 defines both forms. Only when it is absent should you fall back to your own exponential backoff with jitter.
Why add jitter to exponential backoff?
Without jitter, every client that was throttled at the same moment retries at the same moments (1 second later, 2 seconds later, 4 seconds later), so the retries arrive in synchronized waves and keep tripping the limit. Randomizing each delay spreads the load. The "full jitter" variant picks a random value between zero and the exponential ceiling.
Should I retry a POST after a 429?
Usually yes, because 429 means the server rejected the request without acting on it, so a retry is safe even for non-idempotent methods. Confirm in the API documentation, and for gateways that may return 429 after partially processing, send an Idempotency-Key header if the API supports one.
What are the RateLimit and RateLimit-Policy headers?
They come from an IETF HTTPAPI working group draft (draft-ietf-httpapi-ratelimit-headers, revision 11 as of May 2026), which is not an RFC and may still change. RateLimit-Policy advertises the quota policies (quota q over window w seconds) and RateLimit reports the remaining quota r and the effective window t in seconds. Many APIs still send the older, non-standard X-RateLimit-* headers with different semantics, so always read the specific API documentation.
Why does retrying a 429 from an LLM API sometimes never work?
Some providers return 429 for two different conditions: a short-term rate limit that clears in seconds, and an exhausted quota or billing limit that never clears until you change the plan. The response body usually distinguishes them. Retry the first kind with backoff and surface the second to a human.
Is a 429 the same as a 503?
No. 429 says this client is sending too much and may retry later. 503 says the service is unavailable or overloaded for everyone. Both may carry Retry-After and both should be retried with the same backoff logic, but a 429 is also a signal to reduce your own request rate.
Sources
- MDN: 429 Too Many Requestsdeveloper.mozilla.org
- RFC 6585 Section 4: 429 Too Many Requestsrfc-editor.org
- RFC 9110 Section 10.2.3: Retry-Afterrfc-editor.org
- IETF Draft: RateLimit header fields for HTTPdatatracker.ietf.org
- GitHub Docs: Rate limits for the REST APIdocs.github.com
- AWS Architecture Blog: Exponential Backoff And Jitteraws.amazon.com
Related
API Rate Limiting: Algorithms, Headers and 429
Rate limit an HTTP API: fixed window, sliding window, token bucket and leaky bucket compared, plus 429, Retry-After, RateLimit headers, nginx and Express.
HTTP 503 Service Unavailable: Causes, Fixes and Retry-After
Fix HTTP 503 Service Unavailable: nginx no live upstreams, Kubernetes endpoints, ALB healthy hosts, Cloudflare, and a maintenance page with Retry-After.
Retry-After
Learn how the Retry-After header tells clients how long to wait before retrying a request. Understand its use with 503, 429, and 301 status codes.
X-RateLimit Headers
Learn how X-RateLimit headers inform API clients about rate limits, remaining requests, and reset times. Implement proper rate limiting in your applications.