How HTTP Works

Debug guide · you're seeing

  • 429 Too Many Requests
  • HTTP/1.1 429 Too Many Requests
  • You have exceeded a secondary rate limit. Please wait a few minutes before you try again.
  • Rate limit reached for gpt-4o in organization org-xxxx on requests per min (RPM): Limit 3, Used 3, Requested 1. Please try again in 20s.
  • Error 1015: You are being rate limited

429 Too Many Requests: Fix It with Retry-After and Backoff

Fix 429 Too Many Requests as an API client: honor Retry-After, back off exponentially with jitter, read RateLimit headers, handle GitHub and LLM limits.

Reviewed 7 min readintermediate6 sourcesTry itMarkdown
On this page

TL;DR: The server is throttling your client. If the response has Retry-After, wait at least that long. If not, retry with exponential backoff and full jitter, cap the attempts, and lower your steady request rate so you stop hitting the limit. Immediate retries make it worse and can turn a rate limit into a ban.

What it means

429 was added by RFC 6585 section 4 and is still the status servers use for “this user has sent too many requests in a given amount of time”. The response may include a Retry-After header and should explain the limit in the body:

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 30
RateLimit-Policy: "burst";q=100;w=60
RateLimit: "burst";r=0;t=30

{"error":"rate_limited","message":"Rate limit exceeded. Try again in 30s."}

Throttling can be applied per API key, per IP address, per account, per endpoint or per token count, and several limits can apply at once. A client that was well below its per-minute limit can still hit a per-second burst limit.

Who sent it?

Several different layers return 429, and they need different fixes:

  • The API’s own application layer: JSON body naming the limit, often with Retry-After and X-RateLimit-* or RateLimit-* headers. Follow the vendor’s documentation.
  • An API gateway (AWS API Gateway throttling, Kong, Apigee): a generic “Too Many Requests” body, often without a Retry-After. Gateways commonly allow bursts and a steady rate; slow your average.
  • A WAF or CDN in front of the origin: a Cloudflare page “Error 1015: You are being rate limited” with Server: cloudflare and a CF-Ray; the origin never saw the request. These rules often key on IP and path, so changing nothing but your IP or user agent changes nothing about the real problem.
  • Your own nginx: limit_req returns 503 by default, so a 429 from nginx means someone set limit_req_status 429.

Print only the headers that matter:

curl -si https://api.example.com/items -H "Authorization: Bearer $TOKEN" \
  | grep -iE '^(HTTP/|retry-after|ratelimit|x-ratelimit|x-github|server|cf-ray|via)'

Fix it, in order of likelihood

  1. Stop retrying immediately. Honor Retry-After.
  2. Add capped exponential backoff with jitter for the case where there is no Retry-After.
  3. Cap the number of attempts and the total wait, and surface the failure instead of looping forever.
  4. Reduce the request rate itself: limit client concurrency, batch calls, and cache.
  5. Check whether the 429 means “slow down” or “you have no quota”, and do not retry the latter.
  6. Share one limiter across all workers and processes that use the same credential.

Honor Retry-After

Retry-After is either a non-negative integer number of seconds or an HTTP date (RFC 9110 section 10.2.3). Handle both:

Retry-After: 120
Retry-After: Mon, 04 Oct 2027 09:30:00 GMT

Exponential backoff with full jitter

Delay for attempt n (starting at 0) is a random value between 0 and min(cap, base * 2^n). This is the “full jitter” strategy from the AWS Architecture Blog and performs better than plain exponential delay because it breaks the synchronization between clients.

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms))

function parseRetryAfter(value) {
  if (!value) return null
  if (/^\d+$/.test(value.trim())) return Number(value) * 1000 // delta-seconds
  const date = Date.parse(value) // HTTP-date
  if (!Number.isNaN(date)) return Math.max(0, date - Date.now())
  return null
}

async function fetchWithRetry(url, init = {}, { retries = 5, baseMs = 500, capMs = 30_000 } = {}) {
  for (let attempt = 0; ; attempt++) {
    const response = await fetch(url, init)

    if (response.status !== 429 || attempt >= retries) return response

    const serverDelay = parseRetryAfter(response.headers.get('retry-after'))
    const backoff = Math.random() * Math.min(capMs, baseMs * 2 ** attempt)
    // Never wait less than the server asked for; add jitter on top of it
    const delay = serverDelay !== null ? serverDelay + Math.random() * 1000 : backoff

    await response.body?.cancel() // free the connection before waiting
    await sleep(delay)
  }
}

Requests with a streaming body cannot be replayed, so only use this wrapper with strings, Blob, FormData or URLSearchParams bodies. If your callers already pass an AbortSignal, pass it through init; an aborted fetch throws and exits the loop.

Python with requests and urllib3 already implements most of this. Retry honors Retry-After for 413, 429 and 503 by default and sleeps backoff_factor * 2**(n-1) seconds otherwise. Its delays are not jittered unless you set backoff_jitter (urllib3 2.x), and the sleep is not “full jitter” as described above:

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

retry = Retry(
    total=5,
    status_forcelist=[429, 503],
    backoff_factor=0.5,
    backoff_jitter=0.5,     # urllib3 2.x; without it the delays are deterministic
    respect_retry_after_header=True,  # the default
    allowed_methods=None,   # retry all methods; by default POST is excluded
)

session = requests.Session()
session.mount('https://', HTTPAdapter(max_retries=retry))

Setting allowed_methods=None retries POST on those statuses. For a 503 that can be unsafe if the server partially processed the request; use it only when the API is idempotent or accepts an idempotency key.

Slow down on purpose

Backoff reacts to a limit you already hit. Staying below it needs client-side pacing:

  • Cap concurrency: a pool of 4 to 8 workers rather than firing 500 Promise.all calls.
  • Use a token bucket sized to the documented rate, shared across processes (for example in Redis) if several workers use the same API key.
  • Spread scheduled jobs; ten cron jobs starting at :00 produce a burst each hour.
  • Replace polling with webhooks or longer intervals, and use conditional requests (If-None-Match) where the API supports them. Some APIs, including GitHub’s, document that a 304 Not Modified conditional response does not count against the primary rate limit; check your API’s rules.
  • Cache responses that do not change and request only the fields you need.

Read the rate limit headers

Two generations of headers exist. The older X-RateLimit-* headers are widely used but never standardized:

X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1790000000

The meaning of Reset varies: GitHub uses a Unix timestamp in seconds, other APIs use seconds remaining. The IETF HTTPAPI working group’s draft (draft-ietf-httpapi-ratelimit-headers, still a draft, not an RFC) defines structured-field headers instead:

RateLimit-Policy: "burst";q=100;w=60
RateLimit: "burst";r=40;t=15

The quoted string is a policy name that ties a RateLimit item to its RateLimit-Policy item. In the policy, q is the quota, w the window in seconds and the optional qu the unit (requests, content-bytes or concurrent-requests). In RateLimit, r is the remaining quota and t the effective window in seconds: the draft warns clients not to assume the whole quota is restored when t elapses. If a response carries both RateLimit and Retry-After, Retry-After takes precedence. Earlier revisions used RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset, which is a different syntax. Both are in the wild. When r is small, wait instead of sending another request.

Provider specifics

GitHub has two separate limits. Exceeding the primary limit returns 403 or 429 with x-ratelimit-remaining: 0; x-ratelimit-reset is the UTC epoch second when the window resets. Secondary limits (concurrency and request-rate abuse protection) also answer 403 or 429, with a message like You have exceeded a secondary rate limit. GitHub’s documented guidance is: if retry-after is present, wait that many seconds; if x-ratelimit-remaining is 0, wait until x-ratelimit-reset; otherwise wait at least a minute, and increase the wait on repeated failures. GitHub also tells clients to avoid concurrent requests and to pace requests that create content. Continuing to send requests while limited can get the integration banned.

LLM and AI APIs generally apply several limits at once: requests per minute, tokens per minute, and sometimes concurrent requests or a spending cap. Responses usually carry the remaining counts and reset time for each in provider-specific headers (for example the x-ratelimit-*-requests and x-ratelimit-*-tokens families, or anthropic-ratelimit-* headers together with retry-after). Two points apply to all of them. Retry the 429 only when it indicates a rate limit; some providers use the same status for an exhausted quota or billing limit, and no amount of backoff fixes that. And count tokens, not just requests: a handful of very large requests can exhaust a token-per-minute budget that many small ones would not.

Reproduce and verify

To confirm your client behaves, point it at a server you control that always returns 429. Any local server works. With Node:

// node 429-server.js
import http from 'node:http'
let hits = 0
http.createServer((req, res) => {
  hits += 1
  console.log(new Date().toISOString(), 'hit', hits)
  if (hits < 4) {
    res.writeHead(429, { 'Retry-After': '2', 'Content-Type': 'text/plain' })
    return res.end('slow down')
  }
  res.writeHead(200, { 'Content-Type': 'text/plain' })
  res.end('ok')
}).listen(8080)
curl -si http://localhost:8080/ | head -n 4

The log should show your client’s requests spaced at least two seconds apart (the Retry-After value) for the first three, and a success on the fourth. Spacing that is shorter than the header, or perfectly regular, means the code is ignoring Retry-After or has no jitter.

If you operate the API and want to produce a correct 429 yourself, return Retry-After and a body that names the limit, and log the key that triggered it so clients can tell their own throttling from a shared one.

Frequently asked questions

How long should I wait after a 429?

Wait for the time in the Retry-After header if the response has one. It is either a number of seconds or an HTTP date, and RFC 9110 section 10.2.3 defines both forms. Only when it is absent should you fall back to your own exponential backoff with jitter.

Why add jitter to exponential backoff?

Without jitter, every client that was throttled at the same moment retries at the same moments (1 second later, 2 seconds later, 4 seconds later), so the retries arrive in synchronized waves and keep tripping the limit. Randomizing each delay spreads the load. The "full jitter" variant picks a random value between zero and the exponential ceiling.

Should I retry a POST after a 429?

Usually yes, because 429 means the server rejected the request without acting on it, so a retry is safe even for non-idempotent methods. Confirm in the API documentation, and for gateways that may return 429 after partially processing, send an Idempotency-Key header if the API supports one.

What are the RateLimit and RateLimit-Policy headers?

They come from an IETF HTTPAPI working group draft (draft-ietf-httpapi-ratelimit-headers, revision 11 as of May 2026), which is not an RFC and may still change. RateLimit-Policy advertises the quota policies (quota q over window w seconds) and RateLimit reports the remaining quota r and the effective window t in seconds. Many APIs still send the older, non-standard X-RateLimit-* headers with different semantics, so always read the specific API documentation.

Why does retrying a 429 from an LLM API sometimes never work?

Some providers return 429 for two different conditions: a short-term rate limit that clears in seconds, and an exhausted quota or billing limit that never clears until you change the plan. The response body usually distinguishes them. Retry the first kind with backoff and surface the second to a human.

Is a 429 the same as a 503?

No. 429 says this client is sending too much and may retry later. 503 says the service is unavailable or overloaded for everyone. Both may carry Retry-After and both should be retried with the same backoff logic, but a 429 is also a signal to reduce your own request rate.

Sources

  1. MDN: 429 Too Many Requestsdeveloper.mozilla.org
  2. RFC 6585 Section 4: 429 Too Many Requestsrfc-editor.org
  3. RFC 9110 Section 10.2.3: Retry-Afterrfc-editor.org
  4. IETF Draft: RateLimit header fields for HTTPdatatracker.ietf.org
  5. GitHub Docs: Rate limits for the REST APIdocs.github.com
  6. AWS Architecture Blog: Exponential Backoff And Jitteraws.amazon.com

Keep going

Browse /search