Building with AI

Rate limits

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A rate limit caps how much you can use an AI service in a set stretch of time, such as how many requests you send each minute.

1 · What it is

A rate limit caps how much you can use an AI service in a set stretch of time. Providers use them to stop floods of requests. They also help share the service fairly, so it stays smooth for everyone.

Services usually count several things at once: requests each minute (RPM), tokens each minute (TPM) and requests each day (RPD). Tokens are the small chunks of text a model reads and writes. Token limits cap how much text you send, and some providers also cap how much comes back. Go over any single limit and you get an error. Picture a limit of 20 requests a minute. A 21st request inside that minute fails, even if each request was tiny.

The error is the web status code 429, which means too many requests in a short window. The reply may carry a retry-after header. A header is a short note attached to the reply, and this one says how long to wait. Its value can be a calendar date or a count of seconds. The usual response is to wait at least that long, then retry. If it fails again, wait a little longer each time, and add a short random pause so many apps do not all retry at once. Resending instantly in a loop does not work, because failed requests still count toward the per-minute limit.

For engineers: OpenAI replies carry headers such as x-ratelimit-remaining-requests, which show how many requests you have left. Anthropic’s API uses a token bucket. Think of a bucket that slowly refills up to the top, instead of being emptied and refilled at fixed times. A cache keeps text the model has already read, so it can be reused. On most Claude models, input tokens served from the cache are left out of the input-token limit. On the Gemini API, the daily request quota starts fresh at 12 a.m. Pacific time.

2 · Why it exists

Without limits, one heavy user could slow a shared service for everyone.

Deliberate floodingSomeone could flood the service with requests to try to overload it.
Unfair shareOne person or company sending far too much could bog it down for others.
Sudden surgesA sharp jump in traffic could strain the servers and cause slowdowns.
3 · How it works

Follow one request as it meets the limit.

Every request is checked against each limit. Over any one, the reply is a 429 and the app waits before trying again.
  1. 1 · sendYour app sends a request, which counts toward request limits and token limits.
  2. 2 · checkThe service compares your usage with each limit, and going over any one of them causes an error.
  3. 3 · rejectIf you are over, it replies with a 429 error that may carry a retry-after header saying how long to wait.
  4. 4 · retryYour app waits at least that long, adds a short random pause, then tries again.

Treat a 429 as wait, then try again, not as a crash.

4 · Where it's used
WhoWhat they askWhat it works with
Chatbot startup“Why do some replies fail at lunchtime rush?”Requests per minute across all users
Data team“Can we summarise ten thousand reports tonight?”Tokens per minute for a long job
Student project“Why did my free key stop working after lots of tests?”Requests per day on a free tier
5 · What it solves, and what it doesn't
solves
  • It stops one client from flooding a shared service.
  • It gives many users a fair share of the capacity.
  • It helps keep the service smooth and steady for everyone.
doesn't solve
  • It is not the same as a spend limit, which is a separate control on cost.
  • The limits are maximums, not a promised amount of capacity.
  • Resending straight away does not help, because failed requests still count.
  • On Anthropic's API, hitting a spend cap also returns a rate-limit error, and retrying fails until access resumes.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsRate limits, OpenAI · read 28 Sept 2026
  2. docsRate limits, Anthropic · read 28 Sept 2026
  3. docsRate limits, Google AI for Developers · read 28 Sept 2026
  4. officialRFC 6585: Additional HTTP Status Codes, IETF · read 28 Sept 2026
  5. officialRFC 9110: HTTP Semantics, IETF · read 28 Sept 2026