Building with AI

Model routing

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Model routing sends each prompt to the model that suits it, so simple requests go to cheaper models and hard ones to stronger models.

1 · What it is

Model routing puts a small decision step in front of several AI models. Each prompt goes to the router first. The router guesses which model will do a good enough job, then passes the prompt on. Easy prompts go to a smaller, cheaper model. Hard ones go to a larger, stronger model.

Think of a homework help app. A question like “what is 7 times 8?” does not need the most expensive model. A tricky proof might. A router lets the app send each question to the right one, without the user choosing. In the RouteLLM research system, a trained model guesses the chance that the stronger model would give the better answer. You pick a threshold, a cut-off number, that turns this chance into a choice between the two models. The routers learned from human preference data.

There is a second design, called a cascade. It asks a cheap model first and scores the answer. If the score is too low, it asks the next, stronger model. The FrugalGPT paper used this idea. In its tests it matched the best single model at up to 98% lower cost.

Several platforms now offer ready-made routers. Azure’s model router judges prompts by complexity, reasoning and task type, and offers Balanced, Cost and Quality modes. Amazon Bedrock predicts how good each model’s answer will be. It starts from a fallback, or default, model and switches only when another is expected to do clearly better. OpenRouter’s Auto Router sorts each prompt into one of about 30 task types. It then picks from the models its users spend most on for that kind of task.

2 · Why it exists

When an app uses one model for everything, it pays too much or gets weaker answers.

Strong models cost morePowerful models give better answers but are expensive. Smaller models are cheaper but less capable.
Prices vary widelyOne study found that fees for popular model APIs could vary about a hundredfold from cheapest to priciest.
Teams build their own rulesWithout a router, teams compare models one by one and write their own rules to balance quality, cost and speed.
3 · How it works

Follow one prompt through a two-model router.

The router makes one choice per prompt. Moving the threshold trades quality against cost.
  1. 1 · sendThe app sends its prompt to the router instead of to one fixed model.
  2. 2 · predictA small trained model estimates how likely the stronger model is to beat the cheaper one on this prompt.
  3. 3 · decideA threshold you set turns that estimate into a choice between the cheaper and the stronger model.
  4. 4 · answerThe chosen model writes the reply, and the app never has to pick a model itself.

The router's job is to predict, before any model answers, whether the expensive one is worth it.

4 · Where it's used
WhoWhat they askWhat it works with
Any app with mixed requests“Why pay top prices for easy questions?”Smaller, cheaper models when they are enough, larger ones for complex tasks
Coding assistant team“Which requests really need a reasoning model?”Reasoning models only for tasks that need complex reasoning
Platform team on Azure“Can we get one deployment that picks the model per request?”A model router in Balanced, Cost or Quality mode
Team on Amazon Bedrock“When is the bigger model clearly better than our default one?”A fallback model plus a response quality difference setting
5 · What it solves, and what it doesn't
solves
  • Easy prompts go to smaller, cheaper models, and harder prompts still reach larger ones.
  • One router or deployment can stand in front of several models.
  • A threshold or routing mode lets you choose how much quality to trade for lower cost.
  • In the RouteLLM tests, trained routers cut costs by more than half without lowering response quality.
doesn't solve
  • A router is only as good as its training data, and it can misjudge unusual or specialised prompts.
  • Bedrock's router is tuned for English prompts only.
  • In Azure's router, the usable context window is limited by the smallest model in the pool.
  • You still need to test the router against your current setup to know it helps.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperRouteLLM: Learning to Route LLMs with Preference Data, Ong et al., UC Berkeley, Anyscale and Canva · read 28 Sept 2026
  2. repoRouteLLM, LMSYS · read 28 Sept 2026
  3. paperFrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, Chen, Zaharia and Zou, Stanford University · read 28 Sept 2026
  4. docsModel router for Microsoft Foundry, Microsoft Learn · read 28 Sept 2026
  5. docsUnderstanding intelligent prompt routing in Amazon Bedrock, Amazon Web Services · read 28 Sept 2026
  6. docsAuto Router, OpenRouter · read 28 Sept 2026