Policy-based LLM routing
Agents don't hardcode providers or model names. They run with a capability and a runtime context, and the platform resolves the effective model from a catalog — deterministically, and auditable.
Agents don't pick models
An agent definition declares an id, a role, tools, and behavior — but no provider, no model name, not even a capability. When it runs, the runtime asks a routing resolver for the effective model, keyed by the capability it needs and which agent is asking. The resolver returns a model profile; a model factory builds the concrete client.
Concepts
| Concept | Meaning |
|---|---|
| capability | The technical model family the runtime needs — one of chat, language, embedding, image. The runtime sets this (chat for ReAct/Graph); the author never declares it. |
| model profile | A named, reusable model configuration — profile_id + capability + model (provider, name, settings). |
| agent override | An optional, per-agent profile pick that takes precedence over the capability's catalog default. |
Source: fred-runtime model_routing/contracts.py (ModelCapability, ModelProfile)
The model catalog
Routing is loaded from config/models_catalog.yaml at pod startup (override with
FRED_MODELS_CATALOG_FILE). The catalog is mandatory. Its keys:
common_model_settings— global defaults merged into every profile.common_model_settings_by_capability— defaults merged per capability.default_profile_by_capability— required: the fallback profile for each capability.profiles— the concrete model profiles.
Settings merge deterministically: common_model_settings → common_model_settings_by_capability[capability] → profile.model.settings (later wins).
version: v1
common_model_settings:
temperature: 0.0
max_retries: 0
default_profile_by_capability:
chat: default.chat
profiles:
- profile_id: default.chat
capability: chat
model:
provider: openai
name: gpt-4o-mini
- profile_id: chat.fast
capability: chat
model: { provider: openai, name: gpt-4o-mini }
- profile_id: chat.quality
capability: chat
model: { provider: openai, name: gpt-4o }
Source: fred-runtime model_routing/catalog.py · example apps/fred-agents/config/models_catalog.yaml
Providers
A profile's provider selects the client the model factory builds. Supported providers:
| provider | Backend |
|---|---|
| openai | OpenAI models — and any OpenAI-compatible endpoint (e.g. Mistral, or a local mock) by setting a base_url in settings. |
| azure-openai | Azure OpenAI deployments. |
| azure-apim | Azure OpenAI fronted by API Management. |
| ollama | Locally / self-hosted models via Ollama. |
| vertex-ai | Google Vertex AI (Gemini). |
| vertex-ai-model-garden | Vertex Model Garden (Mistral / Llama / Claude). |
| anthropic | Anthropic Claude models. |
provider: openai
with a Mistral base_url. Provider credentials come from environment variables
(OPENAI_API_KEY, ANTHROPIC_API_KEY, GCP ADC, …), never from the catalog.
Source: fred-core model/models.py (ModelProvider) · model/factory.py
Resolution
For one chat turn, the effective model is decided in a fixed order, highest precedence first:
- Platform-wide binding — an operator can pin the model every agent on the deployment uses, overriding everything below it.
- Agent override — a team can point a specific agent at a different profile than the capability's catalog default.
- Capability default —
default_profile_by_capability[capability]from the catalog, used when nothing more specific applies.
The decision is made once per turn and returned as a result carrying the resolved
profile_id, the concrete model, and which of the three levels decided
it — deterministic and auditable.
Source: fred-runtime model_routing/resolver.py, provider.py
Runtime behavior
- ReAct and Graph agents both resolve their model once per turn, not once per step — the same call, the same timing, regardless of shape.
- Reasoning is a separate, boolean on/off dimension a user can request per turn, layered on top of whichever model resolution picked — not a way to pick a different model, and not a low/medium/high effort choice (effort itself stays admin-configured). See Reasoning selector for the chat-side control.
- Embeddings for ingestion are configured separately in the Knowledge Flow config, not through the chat routing path.
Source: fred-runtime react_runtime.py, graph_runtime.py
Where routing lives
Two layers with a clean split:
- Decide which profile —
fred-runtime'smodel_routingpackage: a pure, deterministic resolver over the loaded policy (no I/O, no model construction). - Build the client —
fred-core's model factory turns the chosen profile's configuration into the concrete LangChain client.
The agent pod wires these together at startup, exposing a model-factory port to the runtime; the agent only ever asks for "a model for this turn".
Source: fred-runtime model_routing/provider.py, app/agent_app.py · fred-core model/factory.py
Governance position
The posture is policy-first: model choice is governed from the catalog, not chosen by end users at runtime — there is no end-user model picker acting as the routing authority. (The one place a user selects a model profile is the evaluation judge, a separate surface.) This keeps behavior predictable and aligned with team governance.
Model routing is the model domain of Fred's governance. Tools live in a sibling
mcp_catalog.yaml loaded the same way; conversation retention and scheduling have
their own policy catalogs. They are parallel per-domain catalogs rather than one monolith.
Source: fred-agents config/models_catalog.yaml, config/mcp_catalog.yaml