Rate limits
A rate limit caps how much you can use an AI service in a set stretch of time, such as how many requests you send each minute.
A rate limit caps how much you can use an AI service in a set stretch of time. Providers use them to stop floods of requests. They also help share the service fairly, so it stays smooth for everyone.
Services usually count several things at once: requests each minute (RPM), tokens each minute (TPM) and requests each day (RPD). Tokens are the small chunks of text a model reads and writes. Token limits cap how much text you send, and some providers also cap how much comes back. Go over any single limit and you get an error. Picture a limit of 20 requests a minute. A 21st request inside that minute fails, even if each request was tiny.
The error is the web status code 429, which means too many requests in a short window. The reply may carry a retry-after header. A header is a short note attached to the reply, and this one says how long to wait. Its value can be a calendar date or a count of seconds. The usual response is to wait at least that long, then retry. If it fails again, wait a little longer each time, and add a short random pause so many apps do not all retry at once. Resending instantly in a loop does not work, because failed requests still count toward the per-minute limit.
For engineers: OpenAI replies carry headers such as x-ratelimit-remaining-requests, which show how many requests you have left. Anthropic’s API uses a token bucket. Think of a bucket that slowly refills up to the top, instead of being emptied and refilled at fixed times. A cache keeps text the model has already read, so it can be reused. On most Claude models, input tokens served from the cache are left out of the input-token limit. On the Gemini API, the daily request quota starts fresh at 12 a.m. Pacific time.
Without limits, one heavy user could slow a shared service for everyone.
Follow one request as it meets the limit.
- 1 · sendYour app sends a request, which counts toward request limits and token limits.
- 2 · checkThe service compares your usage with each limit, and going over any one of them causes an error.
- 3 · rejectIf you are over, it replies with a 429 error that may carry a retry-after header saying how long to wait.
- 4 · retryYour app waits at least that long, adds a short random pause, then tries again.
Treat a 429 as wait, then try again, not as a crash.
| Who | What they ask | What it works with |
|---|---|---|
| Chatbot startup | “Why do some replies fail at lunchtime rush?” | Requests per minute across all users |
| Data team | “Can we summarise ten thousand reports tonight?” | Tokens per minute for a long job |
| Student project | “Why did my free key stop working after lots of tests?” | Requests per day on a free tier |
- It stops one client from flooding a shared service.
- It gives many users a fair share of the capacity.
- It helps keep the service smooth and steady for everyone.
- It is not the same as a spend limit, which is a separate control on cost.
- The limits are maximums, not a promised amount of capacity.
- Resending straight away does not help, because failed requests still count.
- On Anthropic's API, hitting a spend cap also returns a rate-limit error, and retrying fails until access resumes.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsRate limits, OpenAI · read 28 Sept 2026
- docsRate limits, Anthropic · read 28 Sept 2026
- docsRate limits, Google AI for Developers · read 28 Sept 2026
- officialRFC 6585: Additional HTTP Status Codes, IETF · read 28 Sept 2026
- officialRFC 9110: HTTP Semantics, IETF · read 28 Sept 2026