LLM Model Routing: Send Easy Prompts to Small Models
Most prompts hitting an LLM app are simple. Route those to a small model and escalate only the hard ones to a large model to cut cost and keep quality.
Most LLM apps send every request to one model, usually the biggest one the budget allows, because that setup never embarrasses you in a demo. Then the bill arrives. The uncomfortable part is that the traffic paying for that big model is not uniform. Take these two questions:
- “which page documents the reprocessing config”
- “reconcile these three log excerpts and tell me why the job died”
Fed to the same endpoint, they cost exactly the same, even though the first is a lookup and the second is real work.
Routing means picking the cheapest model that will still get a given request right, instead of paying the top rate for the whole stream. This post is for engineers who already run an LLM feature in production, have watched the monthly cost climb, and suspect most of what they are paying for is easy. I cover the two shapes routing takes, a small predictive router you can actually read, how to build the classifier that drives it, and the failure modes that make routing quietly worse than doing nothing.
I care about this because I built Archi, a RAG copilot for CMS computing operations at CERN, and the question mix there is lopsided. A lot of it is “where is the doc for X” and “what does this error string mean,” which a small model answers fine. A smaller slice needs a model that can hold several log snippets in its head and reason across them. Sending both kinds to the same expensive model is the obvious thing, and also the wasteful thing.
Two shapes: decide before the call, or after
There are two honest ways to route, and they differ on when you commit money.
Predictive routing: classify the prompt first
Predictive routing looks at the prompt and picks a model before calling anything. A cheap classifier reads the request, guesses whether it is easy or hard, and sends it down one lane. This is the pattern Anthropic describes as routing in their agent workflows writeup: a classifier directs input to a specialized handler.
The appeal is that the happy path is one model call. The risk is that the classifier is guessing. When it sends a hard prompt to the small model, nobody finds out until a user complains.
Cascade routing: try the small model, then check
Cascade routing runs the small model first, every time, and then checks the answer. If the answer passes a verifier, you ship it and never call the big model. If it fails, you escalate.
FrugalGPT formalized this shape in 2023. The paper reports matching a top model’s quality on its benchmarks at much lower cost, by chaining cheaper models this way. The tradeoff is that hard prompts pay twice: once for the small model that failed, and once for the big model that fixed it.
Choosing between them
Which one fits depends on whether you trust a classifier more than a verifier. Predictive routing has to judge a prompt cold, before any answer exists. Cascade routing gets to look at a real answer before deciding. That is more information, at the cost of always paying for the first call.
The research framing for the predictive side is RouteLLM from LMSYS. It trains a router on human preference data to hit a target like “90% of the strong model’s quality” while pushing as much traffic as possible to the weak model. You do not need their trained router to start. You need to understand which decision you are making, and when.
A predictive router you can read
Here is the core of a predictive router. It tries the cheapest signals first and only reaches for a model to classify when those signals give up. The first piece is a set of rules that either return a model or return None to defer:
SMALL = "small-model"
LARGE = "large-model"
def route_by_rules(prompt: str) -> str | None:
"""Cheap heuristics. Return a model, or None to defer."""
words = len(prompt.split())
lowered = prompt.lower()
# Long context almost always needs the bigger model.
if words > 600:
return LARGE
# Words that signal multi-step reasoning, not lookup.
if any(w in lowered for w in ("step by step", "prove", "reconcile", "compare")):
return LARGE
# Short, specific questions are usually easy.
if words < 40:
return SMALL
return None # undecidedRules like these are unglamorous, and they carry more traffic than people expect, because a large share of real requests are short lookups or obviously long analyses. When the rules defer, the router falls through to something with more judgment, here a small model asked to classify difficulty:
async def route(prompt: str) -> str:
decided = route_by_rules(prompt)
if decided is not None:
return decided
# Fall through: ask a small, cheap model to classify difficulty.
verdict = await classify_difficulty(prompt) # returns "easy" or "hard"
return SMALL if verdict == "easy" else LARGEThe important property is the ordering. You spend a model call on classification only for the ambiguous middle, not for the whole stream. If your classifier costs as much as the model you are trying to avoid, you have built a more complicated way to spend the same money.
The difficulty classifier is the whole game
Everything in predictive routing rests on the difficulty classifier. There are three levels of effort, each with a different price and a different accuracy ceiling.
Heuristics are the cheapest: the route_by_rules function above. They look at length, keywords, whether the request includes code, and whether it is a follow-up in a long conversation. They cost nothing, and they are surprisingly hard to beat on the obvious cases. Their weakness is the middle: a short prompt can still be a hard reasoning question, and a long one can be a trivial summary.
An embedding classifier is the middle option. Embed a few dozen example prompts you have hand-labeled easy or hard. Then embed the incoming prompt and route by nearest neighbors, or by a small logistic model on top of the embeddings. This is cheap per request, because an embedding call is far cheaper than a generation call, and it generalizes better than keyword rules. It needs a labeled set, which means you have to look at your own traffic, and that is worth doing anyway.
A small LLM acting as the classifier is the most flexible option: the classify_difficulty call above. You prompt a cheap model with “is this a simple lookup or does it need multi-step reasoning” and read one token back. It handles cases the other two miss, and it costs a real (if small) call each time. This is the same idea as using an LLM as a judge, pointed at difficulty instead of quality.
In practice you stack them: rules catch the easy majority for free, and the model classifier handles what is left. The point is not to pick the fanciest classifier. It is to spend classification effort in proportion to how ambiguous the prompt is.
In a cascade, the verifier is the hard part
Cascade routing moves the hard decision to after the answer exists. That sounds easier, and it is not, because now you need something that can look at a draft answer and decide whether it is good enough to ship. That verifier is the whole design. The cascade itself is a few lines, and everything interesting hides inside good_enough:
async def answer(prompt: str) -> str:
draft = await call(SMALL, prompt)
if await good_enough(prompt, draft):
return draft
return await call(LARGE, prompt)The tempting shortcut is to trust the model’s own confidence, and it is a trap. A small model that is wrong is often wrong confidently, so its self-reported certainty is not a reliable escalation signal.
What works better is a check tied to the actual task:
- In a RAG system, I can ask a cheap, concrete question: did the draft cite a document that exists in the retrieved set, or did it answer from nothing?
- If the answer is supposed to be JSON, I can just try to parse it, which is the same discipline as getting reliable JSON out of an LLM in the first place. A failed parse is a clean, honest escalation trigger, with no second model needed to detect it.
When there is no cheap structural check, a separate small judge model can score the draft, but keep it cheap. A verifier as expensive as the large model erases the saving, since every hard prompt then pays for the small model, the judge, and the large model. The verifier only earns its place if it is much cheaper than the escalation it prevents.
Where routing quietly makes things worse
Routing looks like free money. Then several things go wrong, in ways that are easy to miss because the system keeps returning answers.
Misroutes are not symmetric. Sending an easy prompt to the big model wastes a little money. Sending a hard prompt to the small model ships a wrong answer to a user, which costs trust. Those are not the same size of mistake, so the router should not treat them as a coin flip. Bias toward escalation: when the classifier is unsure, spend the money. A router tuned for maximum saving is a router tuned to be confidently wrong on your hardest requests.
Cascade only wins if easy prompts dominate. Because hard prompts pay twice under a cascade, the math only works when the easy majority is large enough for its savings to cover the double-paid minority. If your traffic is half hard, a cascade can cost more than sending everything to the big model. Measure the split before you build it, not after.
The router’s own cost is real. Every classification call and every verifier call is spend you did not have before. If you route or verify with a model as expensive as the one you are avoiding, the ledger does not move. Route with something at least an order of magnitude cheaper than the model you are protecting. The FrugalGPT paper notes that prices across available models differ by up to two orders of magnitude. That gap is exactly what makes routing worth doing, and it is also the gap you erase if your router is heavy.
The traffic distribution drifts. Yesterday’s easy question is today’s hard one when a new subsystem ships and users start asking about it. A classifier tuned on last quarter’s traffic slowly rots. Log every routing decision alongside the outcome, so you can see the easy/hard boundary moving and retune before the misroute rate climbs.
You cannot tell if it is working without evaluation. Cost is easy to measure and quality is not. So the failure mode is a router that cuts the bill in half and also quietly drops answer quality on the hard tail. You need a labeled evaluation set and a per-tier quality number, or you are flying blind on the half of the tradeoff that actually matters.
What I would measure before turning it on
Before routing anything, I would answer one question with real data: what fraction of my traffic is genuinely easy? Pull a few hundred recent requests, label them, and look at the split.
- If it is 90/10 easy to hard, routing is a large win and worth the complexity.
- If it is closer to even, the savings shrink while the misroute risk stays. You might be better off spending the effort on prompt caching or semantic caching instead, which cut cost without any risk of sending a request to a weaker brain.
Once it is live, measure three things:
- the actual easy/hard split the router produces,
- the answer quality per tier, against your labeled set,
- the cost per thousand requests, before and after.
If quality on the small-model tier holds and cost drops, you were right about your traffic. If quality on that tier sags, your classifier is too eager, and you tighten it toward escalation.
Routing is one lever on the same bill I have written about from other angles. Caching removes repeated work; routing removes overqualified work. For something like Archi they stack: cache the answers that recur, send the new-but-easy ones to a small model, and keep the expensive model for the hard operator questions that actually need it. None of these tricks is clever by itself. Stacked, they are what keeps an LLM feature alive through a budget review.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.