Model routing
Model routing sends each prompt to the model that suits it, so simple requests go to cheaper models and hard ones to stronger models.
Model routing puts a small decision step in front of several AI models. Each prompt goes to the router first. The router guesses which model will do a good enough job, then passes the prompt on. Easy prompts go to a smaller, cheaper model. Hard ones go to a larger, stronger model.
Think of a homework help app. A question like “what is 7 times 8?” does not need the most expensive model. A tricky proof might. A router lets the app send each question to the right one, without the user choosing. In the RouteLLM research system, a trained model guesses the chance that the stronger model would give the better answer. You pick a threshold, a cut-off number, that turns this chance into a choice between the two models. The routers learned from human preference data.
There is a second design, called a cascade. It asks a cheap model first and scores the answer. If the score is too low, it asks the next, stronger model. The FrugalGPT paper used this idea. In its tests it matched the best single model at up to 98% lower cost.
Several platforms now offer ready-made routers. Azure’s model router judges prompts by complexity, reasoning and task type, and offers Balanced, Cost and Quality modes. Amazon Bedrock predicts how good each model’s answer will be. It starts from a fallback, or default, model and switches only when another is expected to do clearly better. OpenRouter’s Auto Router sorts each prompt into one of about 30 task types. It then picks from the models its users spend most on for that kind of task.
When an app uses one model for everything, it pays too much or gets weaker answers.
Follow one prompt through a two-model router.
- 1 · sendThe app sends its prompt to the router instead of to one fixed model.
- 2 · predictA small trained model estimates how likely the stronger model is to beat the cheaper one on this prompt.
- 3 · decideA threshold you set turns that estimate into a choice between the cheaper and the stronger model.
- 4 · answerThe chosen model writes the reply, and the app never has to pick a model itself.
The router's job is to predict, before any model answers, whether the expensive one is worth it.
| Who | What they ask | What it works with |
|---|---|---|
| Any app with mixed requests | “Why pay top prices for easy questions?” | Smaller, cheaper models when they are enough, larger ones for complex tasks |
| Coding assistant team | “Which requests really need a reasoning model?” | Reasoning models only for tasks that need complex reasoning |
| Platform team on Azure | “Can we get one deployment that picks the model per request?” | A model router in Balanced, Cost or Quality mode |
| Team on Amazon Bedrock | “When is the bigger model clearly better than our default one?” | A fallback model plus a response quality difference setting |
- Easy prompts go to smaller, cheaper models, and harder prompts still reach larger ones.
- One router or deployment can stand in front of several models.
- A threshold or routing mode lets you choose how much quality to trade for lower cost.
- In the RouteLLM tests, trained routers cut costs by more than half without lowering response quality.
- A router is only as good as its training data, and it can misjudge unusual or specialised prompts.
- Bedrock's router is tuned for English prompts only.
- In Azure's router, the usable context window is limited by the smallest model in the pool.
- You still need to test the router against your current setup to know it helps.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperRouteLLM: Learning to Route LLMs with Preference Data, Ong et al., UC Berkeley, Anyscale and Canva · read 28 Sept 2026
- repoRouteLLM, LMSYS · read 28 Sept 2026
- paperFrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, Chen, Zaharia and Zou, Stanford University · read 28 Sept 2026
- docsModel router for Microsoft Foundry, Microsoft Learn · read 28 Sept 2026
- docsUnderstanding intelligent prompt routing in Amazon Bedrock, Amazon Web Services · read 28 Sept 2026
- docsAuto Router, OpenRouter · read 28 Sept 2026