LLM gateway
An LLM gateway is one server that sits between your apps and AI model providers, adding limits, caching, fallbacks and logs to model calls.
Picture five school clubs sharing one laptop cart. If every club grabs laptops freely, one club can take the whole cart and nobody knows who broke what. So the school puts one person in charge of the cart. That person checks who is asking, hands out laptops fairly and keeps a sign-out sheet. An LLM gateway plays that role for AI models.
Without a gateway, each app calls each model provider directly. LiteLLM calls its gateway a proxy server. Your app talks only to the gateway, and the gateway talks to the models. It lets you reach many models in the same way. It can also track spending for each key. Those models can come from several providers, for example models in Microsoft Foundry or on Amazon Bedrock.
Four jobs happen inside that one stop. The first is access. The gateway can check who is calling before anything reaches a model. LiteLLM, one gateway product, hands out virtual keys. These are its own passwords for the gateway, not the real provider keys, and it tracks spend for each one. It can also give a key a spending limit. The real provider keys can stay in the gateway’s own settings file.
The second job is limits. Providers often set quotas in tokens per minute. Microsoft’s example sets a limit of 500 tokens per minute for each subscription key. Microsoft’s gateway can also count a prompt’s tokens itself, so it can stop a prompt that is already over the limit before it reaches the model.
The third job is routing, which means choosing where each request goes. A fallback list is the simplest version. In Cloudflare’s example, the request goes first to a Llama 3.1 8B model on Workers AI, and if that fails, the gateway sends it to OpenAI instead. A response header called cf-aig-step then reports which step answered. You can add as many backups to the list as you need. A load balancer can also spread requests across several model endpoints.
The fourth job is memory and records. A cache is a store of answers the gateway has already given. Cloudflare labels each saved answer with a fingerprint of the request, made with a hashing method called SHA-256. Changing a single word in the message makes a separate cache entry, so only exact repeats get the saved answer. Some gateways offer semantic caching, which reuses answers for prompts with a similar meaning. The gateway can also log prompts and completions (the model’s replies) and count tokens for each caller, so questions about cost have one place to look.
A gateway has limits too. It does not add model capacity: the models behind it still need to be scaled for the load. It also comes in different shapes. Microsoft’s AI gateway is a feature inside Azure API Management, not a separate product. LiteLLM instead gives you a proxy server to run yourself, for example from a Dockerfile. The idea is the same in each: one front door, clear rules and one place to look when something goes wrong.
When several apps share the same AI models, sharing quotas and surviving outages gets hard.
Follow one request through a gateway; the fallback part follows Cloudflare's example.
- 1 · sendThe app sends its request to one gateway address instead of to each provider's own API.
- 2 · checkThe gateway checks the caller's limit, which can count tokens per minute for each key.
- 3 · cacheIf caching is on and an identical request was answered before, the stored reply comes back without calling a model.
- 4 · routeOtherwise the gateway sends the request to the first model on its list, and moves to the next one if that call fails.
- 5 · logIt records requests, tokens and cost, so spending can be tracked for each key or team.
The app keeps one address; the gateway decides which model actually answers.
| Who | What they ask | What it works with |
|---|---|---|
| Platform team | “How do we give ten product teams model access, each with its own budget?” | One virtual key per team, with a spending limit on each |
| Support chatbot team | “How do we keep answering when the main model provider is down?” | An ordered fallback list of models from two providers |
| FAQ bot with fixed buttons | “Why pay twice for the exact same question?” | An exact-match response cache |
| Finance | “Which team spent the most on tokens last month?” | Per-key spend records kept by the gateway |
- Your code talks to one endpoint in one request format, and the gateway translates it for each provider.
- Token limits per key stop one app from using up a shared quota.
- Fallbacks retry a failed request on another model or provider.
- Requests, tokens and cost are counted in one place instead of in every app.
- Exact-match caching helps only when requests are identical, which suits apps with a small set of fixed prompts.
- Fallbacks react to failed calls. Cloudflare starts one on a request error or after a timeout you set.
- It does not add model capacity. The model services behind it still need to be scaled for the load.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsLiteLLM AI Gateway (LLM Proxy), LiteLLM · read 28 Sept 2026
- docsFallbacks (Provider Failover), LiteLLM · read 28 Sept 2026
- docsVirtual Keys, LiteLLM · read 28 Sept 2026
- docsCloudflare AI Gateway, Cloudflare · read 28 Sept 2026
- docsFallbacks, Cloudflare · read 28 Sept 2026
- docsCaching, Cloudflare · read 28 Sept 2026
- docsAI gateway capabilities in Azure API Management, Microsoft · read 28 Sept 2026