What to Look for in Model Routing
Model routing is a policy that uses a cheaper model for simple steps and a stronger model only when the task needs it — not a dropdown labelled “best.”
Model routing is the policy that maps a task class to a model class before inference runs. It is not a brand preference. It is not a dropdown labelled “best.”
Someone using the flagship model to label a ticket is how you pay frontier prices for work a compact model could have finished in a second. That is not a moral failing. It is a missing policy. The product either chooses before the call, or a person chooses in a menu, or the default is the largest model “for quality.” Only the first is routing.
OpenAI’s API pricing and Anthropic’s pricing make the same point in public: compact and frontier models are not the same invoice line. Stanford HAI’s 2025 AI Index has tracked how fast inference cost and capability moved — which is exactly why “best model” is not a routing policy. Best for a memo is not best for a classify step. Best last quarter is not best this quarter.
Words you’ll hear
- Frontier / flagship model. The strongest (and usually most expensive) model a lab currently sells. Reserved for synthesis, hard reasoning, and novel language. Not for labelling.
- Compact / small model. Faster and cheaper. Often enough for extract, classify, and summarise. “Small” is a cost and latency class, not an insult.
- Task class. The kind of step: classify, retrieve, forecast, synthesise. If the platform cannot name the class, it cannot route. It can only default.
- NTU. A metered unit of useful work so you can quote and cap a loop. See What is AI token economics. Seats hide routing. NTU makes it visible.
- Model-agnostic. The platform can call more than one provider. That is a menu, not a policy, until it chooses by task class. Extra logos with a hidden flagship default is lock-in with branding.
- Always-flagship. Marketing for “we use the best model.” A classify job does not need a long-context reasoner. Quality theatre is a cost event.
Routing is also not fine-tuning (changing a model’s weights), and it is not orchestration (what steps exist). You can orchestrate a brilliant graph and still send every node to the flagship. You can fine-tune a compact model and still need a policy that sends classify there. Evaluate them separately.
Why you should care
Teams do not wake up and choose waste. They inherit a default.
- Single-model shop. Every label, every search, every memo calls the same flagship. Finance sees one invoice and cannot split labelling from reasoning. You cannot cap what you cannot see.
- User-picked dropdown. Power users pick the most expensive option “to be safe.” New hires copy that habit. Routing is now a training problem. Training problems do not survive quarter-end.
- Always-flagship as quality theatre. Best for whom? Best for a forecast interpolation is often a time-series path, not a frontier model inventing a number that looks fine in a short demo.
McKinsey’s 2025 State of AI survey keeps showing the operational gap: regular use, then a struggle to scale because cost and workflow were never treated as a system. Routing is that system for inference. Without it, scale is a token bill.
Gartner’s AI TRiSM framing implies you can see which model ran, on which data, at what cost. A platform that cannot show that is not ready, regardless of its red-team slides. DORA and NIS2 change the evidence question: you should understand ICT dependencies. “We are not sure which model ran last Tuesday” is a dependency you cannot explain.
Choosing a model is choosing a brain. Choosing tools is choosing hands. Decide them separately. A compact model that extracts a refund still cannot write it without a quoted named signer if that is the workstream policy. Routing does not replace governance. Governance does not replace routing. You need both.
Red flags: “we support many providers” with no task-class map; seat pricing that includes unlimited flagship; users pick the model in production; classify and memo share a model id in the demo; no NTU quote before a run expands; fallback is “switch the dropdown”; graph does not record model id per step; forecasting done by an LLM in the demo on a short series.
What to look for
- They can refuse the flagship. Run a classify-only job and show the model id. If classify used the same model as the memo, routing is a slide. Refusal is the proof. Support for many models is not.
- Task-class map you can read. Classify → compact. Retrieve → embeddings, not stuffing hundreds of tickets into a long window. Forecast → a time-series path, not a frontier model interpolating a spreadsheet. Reason → frontier. If they cannot name the classes, they cannot route them.
- NTU quotes before the run expands. Operators see an estimate and can set a workstream cap. Seat licences hide routing. Unlimited flagship under a seat is always-flagship with a predictable opex line. See Total cost of ownership for enterprise AI.
- Fallback is a logged promotion, not “users will switch the dropdown.” Compact models fail on novel schemas and policy-edge language. Temporarily raise that class, budget-aware, then revert. A promotion without a log is a silent cost change. A dropdown is a training problem.
- The graph records the model id per step. Six months later you can answer “which model drafted this?” without grepping provider dashboards. That is audit as well as cost. See How to evaluate AI audit and observability.
Ask for four artefacts from one workstream run: model id per step; NTU per task class; a classify job that did not use the flagship; a forecast that did not use an LLM as the estimator. If the vendor can only show a chat transcript and a blended token total, routing is not in the product.
Why each artefact matters: model id is Measure in NIST language. NTU per class is how Finance splits labelling from reasoning. Classify-without-flagship is the refusal test. Forecast-without-LLM is whether they know the difference between narration and estimation. Demos are short series. Production is seasonality, holidays, and missing days.
What a live demo should prove
Do not accept a provider logo wall.
- Run one workstream with extract, classify, retrieve, and a memo.
- Show model id per step. Classify and extract are compact. The memo may be frontier.
- Show NTU (or tokens) per task class, quoted before the run grows, with a cap on the workstream.
- Force a compact failure on a novel schema. Show a logged promotion, then a revert — not a user switching a dropdown.
- Show a forecast path that is not an LLM interpolating a sheet. Narration can still be frontier.
- Query the graph: which model drafted this payload? Answer from the ledger, not from a provider console.
- Confirm a named-role gate still applies regardless of which model drafted. Routing chooses the brain. Governance still decides the write.
If they pass by opening a playground and picking “best,” you evaluated a dropdown.
How this shows up in Nimbus
Nimbus treats routing as an operating decision tied to workstream steps: task type, sensitivity, and cost — not “best everywhere.” Compact models handle extract. Frontier models are reserved for synthesis. Spend is NTU-metered, quoted per workstream, visible per step.
Release gates apply regardless of which model drafted the payload. Connectors stay read-only by default. Routing decides which brain reads them. Governance still decides whether anything writes.
See Models. For the unit of account, What is AI token economics. Score the four artefacts above. The product claim is the policy, not the catalogue.
Questions people actually ask
Is “model-agnostic” the same as routing?
No. Model-agnostic means more than one provider. Routing means it chooses by task class, with a default that is cheap where cheap is correct. A hidden always-flagship default is lock-in with extra logos. Ask what happens if the operator never touches a dropdown. If the answer is flagship, you have your policy.
Should operators ever pick a model?
Rarely, and as an override. Production operators should brief outcomes. If quality depends on each user knowing which model is good at JSON, you have staffed a routing department by accident. Overrides should be logged, budget-aware, and exceptional. A dropdown on every run is how always-flagship returns through the side door.
Why not put forecasting in the LLM if the numbers look fine in the demo?
Demos are short series. Production is seasonality, holidays, and missing days. Keep narration on the frontier model and estimation on a time-series path. A fluent number is not a control. The warehouse or the statistical path already owns the number. RAG plus a frontier model is for policy language, not for revenue by region.
What if legal requires a single approved model vendor?
Routing still applies inside that vendor’s catalogue: compact vs frontier vs embedding. Single-vendor is a contracting constraint, not an excuse to max tokens. Model-agnostic is nice. Task-class mapping inside one catalogue is the control. Do not skip routing because the RFP named one lab.
Does NIS2 or DORA change the routing question?
They change the evidence question. DORA and NIS2 expect you to understand ICT dependencies. “We are not sure which model ran last Tuesday” is a dependency you cannot explain. Record model id per step on the graph. That is enough to start. You do not need a new product category. You need Measure.
Related reading
How to evaluate AI workstream platforms, How to evaluate an enterprise AI operating system, and Total cost of ownership for enterprise AI.
Sources
See what governed AI looks like on your stack.
Connect your tools, run a workstream, and keep every decision on your ledger - free for 7 days.