System Overview
Agent harnesses call a stable model endpoint. Semantic Router applies policy to select one model or coordinate a bounded multi-model path over configured backends. The data plane handles requests; the control plane configures and operates it.
Architecture
The default standalone frontend accepts client traffic directly. Recipe routing and native System One serving compose in the same instance. An Envoy-based ExtProc gateway is an alternative ingress, not a required component. See Component Architecture for the protocol path, model-runtime replicas, and explicit native recipes for System One auto.
Data plane
- Agent harness owns the task loop, tool execution, and durable task state. Each inference call crosses the Router boundary; the response returns to the harness for the next step.
- Frontend accepts client traffic, enforces listener access, and adapts
supported protocols. With
--gateway extproc, an external gateway instead owns ingress and forwarding while calling the Router through ExtProc. - Semantic Router extracts signals, evaluates policy, applies route-specific behavior, and selects or coordinates model candidates.
- Model runtime runs decision, classifier, embedding, and reranking models needed by configured consumers. Its workers can be managed or attached.
- Chat backends are configured model services or provider endpoints. Their operators own the weights and generation capacity.
Control plane
- Canonical YAML is the portable source of routing behavior.
- Entrypoints map one or more public model aliases to a recipe.
- Recipes are complete policy and runtime-state isolation boundaries. One or more entrypoints can select the same recipe.
- CLI and Dashboard support local setup, validation, model discovery, configuration, and operation.
- Helm and the Operator deploy the Router into Kubernetes environments.
- Evaluation and observability expose route outcomes so operators can test and improve policy.
Core objects
| Object | Purpose |
|---|---|
| Entrypoint | A mapping from one or more public model aliases to a recipe. |
| Recipe | A complete routing-policy and runtime-state isolation boundary. |
| Signal | A named fact about the request, identity, conversation, or content. |
| Projection | A reusable score, partition, or band derived from signals. |
| Decision | A policy rule that chooses an eligible route and candidate set. |
| Plugin | Route-specific processing such as request controls, memory, retrieval, or response handling. |
| Algorithm | The method used to select or coordinate candidate models. |
| Provider model | A physical inference endpoint available to one or more recipes. |
Reuse detection across policies, change policy independently of model selection, and evolve the physical pool behind a stable public entrypoint.
Request lifecycle
- A harness sends a request using OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages.
- The standalone frontend, or an ExtProc gateway, presents it to the Router.
- The requested model resolves to an entrypoint and its recipe.
- The Router extracts relevant signals and computes projections.
- Decisions enforce constraints and choose an eligible candidate set.
- The route's algorithm selects one model or executes a bounded multi-model strategy.
- Route plugins run at their configured request, execution, or response hook.
- The standalone upstream client, or the external gateway, sends the provider-shaped request to the selected backend and returns the normalized response.
This lifecycle covers a model call inside the harness's task loop. Router plugins can filter the tools exposed to a model or process request context; the harness and tool services still own tool execution and authorization. Configured multi-model algorithms coordinate model calls within this boundary.
Explicit physical model names can still be exposed when an operator wants direct selection. Those requests pass through without recipe signals, decisions, route plugins, cache, learning, or session routing. Virtual model names are useful when clients should choose an objective while the Router owns the physical route. If no decision matches inside the selected recipe, the configured default provider model is used.
Protocol and deployment boundaries
Semantic Router serves its own listeners by default, or integrates through ExtProc with an external gateway. The same routing policy applies in local Docker, Kubernetes, and hybrid environments. Chat backend provisioning and capacity remain the responsibility of the chosen inference platform; the built-in model runtime separately manages its own decision-model replicas.
The Router can consider request semantics and configured runtime observations; it does not replace a Chat backend's scheduler. A deployment may therefore use Semantic Router to choose a model class and an Inference Router to choose a healthy replica of that model. When an AI Gateway also fronts the stack, a request crosses three routing layers:
agent harness
-> AI Gateway (e.g. Agent Router / LiteLLM / agentgateway)
-> Semantic Router ExtProc
-> Inference Router / pool scheduler (e.g. llm-d / vLLM Router / AIBrix gateway)
-> model replica
| Layer | What it owns | Examples |
|---|---|---|
| AI Gateway | Client ingress, provider translation, credentials, rate limits, and traffic policy. | Agent Router (formerly Envoy AI Gateway), LiteLLM, agentgateway |
| Semantic Router | Logical model or model pool selection from request intent and policy, through recipes and decisions. The choice is written to x-selected-model. | vLLM Semantic Router |
| Inference Router | Healthy replica or endpoint selection inside the selected pool. | llm-d, vLLM Router, AIBrix gateway |
Agent Router and agentgateway call Semantic Router through ExtProc. Kubernetes Gateways and Inference Platforms list the integrations this project maintains.
The client and selected backend do not need to use the same wire format. See Protocol Compatibility for the supported client endpoints, backend formats, and pairwise translation matrix.
Next
- Use Cases for practical deployment patterns.
- Routing Pipeline for the policy layers.
- Mixture of Models for virtual models and multi-model execution.
- Quickstart to run the local stack.