An application that sends every request to one large model pays that model's rates for all of it, from a two-word classification to a multi-document synthesis. Moving routine calls to a smaller model by hand means writing routing logic and keeping it current as models change.
Microsoft Foundry's model router packages that choice as a single deployment. You call it like any chat deployment, with the deployment name in the model parameter, and it selects an underlying model for each request. Several of its documented constraints affect how the calling application has to be built.
Routing Modes and What Each One Trades
According to how model router works, the router reads the full request, including the system message, tool definitions and conversation history, and picks a model based on the configured routing mode and model subset.
There are three modes. Balanced is the default: it considers the underlying models within a small quality range of the highest-quality model for that prompt, which the overview gives as an example of 1 to 2 percent, and picks the most cost-effective. Cost mode widens that range, with 5 to 6 percent given as the example, and the design page says it aggressively favors cheaper models, accepting slightly lower quality on complex prompts. Quality mode picks the highest-rated model for the prompt and ignores cost.
Those percentages are the router's own estimates for a prompt, stated as examples. They are not a measured quality guarantee for your workload. The mode belongs to the deployment, so a nightly enrichment job and a customer-facing assistant can use two deployments with different modes. Changes to the mode or to the model subset can take up to five minutes to take effect, according to the how-to guide, which also lists deployment prerequisites such as where Claude models must be deployed before a subset can name them.
The pool itself moves. The current router version, 2025-11-18, is updated in place as Microsoft adds models, and with auto-update selected the set of underlying models can change, which the overview notes could affect performance and cost.
The Smallest Model in the Pool Sets the Context Limit
The overview marks this as important: the effective context window of a router deployment is limited by the smallest underlying model. Its Limitations section spells out the runtime consequence. Other underlying models accept larger contexts, so an API call with a larger context will succeed only if the prompt happens to be routed to a model that can hold it. The troubleshooting table lists "context exceeded" as a failure to expect.
Behind a router, prompt size can decide whether a call succeeds, and the application does not make the routing decision. A retrieval step that sometimes returns many more chunks, because someone uploaded a long document, can turn into intermittent errors.
The documentation offers two responses. One is to keep prompts under the smallest window by summarizing, truncating to the relevant parts, or retrieving sections through document embeddings, which is the discipline a retrieval-augmented application needs anyway. The other is to raise the floor with a model subset that contains only models supporting the context length you need.
A subset has consequences documented elsewhere on the page. Your configured subset is also your failover set, and the guidance is to select at least two models; a single-model subset is listed as a practice to avoid because it removes routing optimization, cost savings and failover. New base models are not added to a subset unless you include them explicitly, which is why Microsoft suggests the subset as a compliance gate, and also why someone should review it when the supported model list changes.
Parameters and Caching Behind a Router
Some inference parameters apply only when a request lands on certain models. When the router selects an o-series reasoning model, it ignores Temperature and Top_P and drops stop, presence_penalty, frequency_penalty, logit_bias and logprobs, so an application that relies on a stop sequence to keep output parseable is relying on it conditionally.
Prompt caching is supported because the underlying models support it, and the overview states the condition: because routing decisions might vary, caching benefits apply only when the same model handles consecutive requests with overlapping prompt prefixes. A poor cache hit rate under a router may reflect the routing distribution. A preview session affinity setting for stateless Chat Completions conversations asks the router to try the same eligible model across related turns, and the documentation says it does not guarantee a cache hit.
Which Model Answered
The selected model is returned in the model field of each response, and the design page calls that field your primary observability signal: log it and track which models handle your traffic. When someone reports that answers changed last week, that column is how you tell whether the model mix moved. Azure Monitor metrics for the deployment can be split by underlying model, and Cost analysis can be filtered on the Deployment tag.
Microsoft's evaluation guidance asks for a benchmark against your current model on quality, cost and latency before production traffic, with prompts grouped by workload category so an average does not hide a regression in one of them.
Related Reading
Adding AI to the .NET Application You Already Run covers the application these calls live inside. Why a Task-Specific Agent Beats a Chatbot Bolted Onto Your App covers scoping what a model is asked to do.