Skip to main content
Version: Latest (unreleased)

Configuration

Semantic Router uses one canonical YAML document across the CLI, Dashboard, Helm, and Operator. The top-level structure is:

version:
listeners:
providers:
evaluation:
routing:
entrypoints:
recipes:
global:

Most deployments begin with version, listeners, providers, and one top-level routing profile. Add entrypoints and recipes when one deployment needs several isolated policies. Add global settings only for shared services or runtime behavior that differs from the built-in defaults.

What belongs where​

SectionOwns
versionCanonical schema version. Use v0.3.
listenersPublic Router listeners: address, port, idle timeout, optional client API keys, an optional Chat model allowlist (models; empty accepts all), separate native System One grants (systemone.models; omission publishes none), optional one-way TLS (tls.cert_file, tls.key_file), and the identity sources a listener trusts (identity.trust_headers, identity.trusted_peers; by default none), which the Router honors in standalone mode.
providersLogical provider models, physical backend endpoints, pricing, capabilities, and defaults.
evaluationOptional operator-owned benchmark definitions, versioned index DAGs, and model-linked records.
routingThe default recipe: model cards, signals, projections, decisions, strategy, algorithms, and route plugins.
entrypointsPublic virtual model aliases mapped to the default or a named recipe.
recipesAdditional isolated routing profiles that share providers and global infrastructure.
globalRouter services, stores, integrations, observability, learning, and router-owned model assets.

Keep these boundaries clear:

  • signals detect facts;
  • projections combine evidence;
  • decisions define eligibility and route policy;
  • algorithms choose or coordinate candidate models;
  • plugins add behavior at route-specific hook points; and
  • providers bind logical model names to inference endpoints.

Provider pricing belongs beside each concrete model under providers.models[].pricing. It accepts an optional uppercase three-letter currency plus non-negative prompt_per_1m, completion_per_1m, cached_input_per_1m, and cache_write_per_1m rates. Routing model cards do not repeat deployment prices or credentials.

Evaluation measurements belong in evaluation.records[] and reference a canonical Model Card identity through model. Built-in benchmark IDs work directly; define new benchmark semantics and indices beside the records under evaluation. See Custom evaluations.

Use Protocol Compatibility to choose the model's backend api_format. Then see Backend Target Compatibility before moving its bindings between Docker, Helm, the Operator, and Dashboard workflows. The target matrix distinguishes canonical pass-through from Kubernetes discovery and records which URL, path, weight, and provider fields each surface preserves.

Router-wide debugging surfaces stay closed by default. global.services.observability.profiling serves Go pprof endpoints, and only when it is explicitly enabled; it then binds 127.0.0.1:6060 so profiles never reach a routable interface without an explicit bind change. The switch is read once at startup, so changing it requires a Router restart. See API and Observability.

Built-in category/domain classification runs Vela Domain in the model runtime when no remote backend is configured. To call a named external classifier, attach a backend under global.model_catalog.modules.classifier.domain and resolve its model from global.model_catalog.external[] with model_role: classification. The shared backend fields are protocol, contract, model, and optional deadline_ms; category currently supports http_classify with the full label_distribution.v1 response contract. Omit backend to keep the runtime-served model. The earlier variant, use_modernbert and use_mmbert_32k selectors are gone; vllm-sr config migrate removes them.

Complexity attaches the same block under global.model_catalog.modules.complexity, beside prototype_scoring. It reads two contracts, so contract cannot be defaulted and must be stated: score.v1 for a regression model, where each rule turns the score into a verdict through its own hard_above/easy_below boundaries (or hard_below/easy_above for a score that falls as difficulty rises), and label_distribution.v1 for a model that returns hard/easy/medium directly. threshold remains the symmetric shorthand for the local signed margin. score.v1 reports no confidence, so decisions gated on those rules rank on the engine's structural default; the Router warns at startup. The remote call is visible through llm_remote_connector_* and llm_complexity_* metrics, and a scorer failure is recorded on every complexity rule's signal errors rather than dropped.

PII attaches the same block under global.model_catalog.modules.classifier.pii. It reads one contract, token_spans.v1, so contract may be omitted. The remote model returns entity spans as code-point offsets into the exact request string it was sent, and every label it returns must exist in the configured pii_mapping_path; a response the contract rejects is a backend failure rather than a clean "no PII" result. on_error beside the backend selects what such a failure, or a provider-declared truncated_at, does to the rule that consumed it: allow (the default) treats the content as not matching, block matches it as classification_error. Spans returned before a declared truncation still count under both policies.

The Routing Pipeline explains the design. Capability pages under Capabilities document each signal, projection, decision, algorithm, plugin, and global block.

Capability catalog​

Use this catalog to choose a reusable building block, then open its guide for configuration details. The inventory comes from config/fragments/; each one-line goal comes from the matching guide's Overview. The documentation build regenerates this block and fails if the checked-in catalog has drifted.

Signals​

Family and typeUse it toReusable fragmentGuide
action — heuristic signalaction labels each request with the operation it asks for: generate, explain, fix, refactor, test, or other.config/fragments/signal/action/Guide
authz — heuristic signalauthz turns identity and policy bindings into reusable routing inputs under routing.signals.role_bindings.config/fragments/signal/authz/Guide
classifier — learned signalclassifier exposes reusable label scores from a local native sequence classifier, a remote sequence classifier, or a configured external LLM.config/fragments/signal/classifier/Guide
complexity — learned signalcomplexity estimates whether a request is easy, medium, or hard by comparing it with configured example sets.config/fragments/signal/complexity/Guide
context — heuristic signalcontext detects requests that need a larger effective context window.config/fragments/signal/context/Guide
conversation — heuristic signalconversation routes on chat structure and protocol facts, such as message count, developer instructions, available tools, explicit tool-use constraints, or an active tool loop.config/fragments/signal/conversation/Guide
decision — learned signaldecision asks a decision model a typed question about the request and turns the answer into a routing fact.config/fragments/signal/decision/Guide
domain — learned signaldomain classifies the request topic family.config/fragments/signal/domain/Guide
embedding — learned signalembedding matches requests by semantic similarity to representative examples.config/fragments/signal/embedding/Guide
event — heuristic signalevent routes structured event-like requests by event type, severity, urgency, or domain-specific action code.config/fragments/signal/event/Guide
fact-check — learned signalfact-check decides whether a prompt should be treated as evidence-sensitive traffic.config/fragments/signal/fact-check/Guide
hallucination — learned signalhallucination checks the model's answer against the grounding context the request carried, such as tool results or retrieved documents, and reports the claims that context does not support.config/fragments/signal/hallucination/Guide
input-modality — heuristic signalinput_modality deterministically matches which kinds of input — text, image, audio, or video — are present in the parsed request.config/fragments/signal/input-modality/Guide
jailbreak — learned signaljailbreak detects prompt-injection and jailbreak attempts before the Router commits to a route.config/fragments/signal/jailbreak/Guide
kb — learned signalkb binds routing signals to the output of a named knowledge base instance.config/fragments/signal/kb/Guide
keyword — heuristic signalkeyword matches explicit words and phrases in the request.config/fragments/signal/keyword/Guide
language — heuristic signallanguage detects the request language and exposes it as a routing signal.config/fragments/signal/language/Guide
metadata — heuristic signalmetadata matches bounded string values supplied by the caller in request metadata.config/fragments/signal/metadata/Guide
modality — learned signalmodality detects whether a request should stay in text generation, switch into image generation, or support both.config/fragments/signal/modality/Guide
pii — learned signalpii detects sensitive personal data in requests.config/fragments/signal/pii/Guide
preference — learned signalpreference infers response-style preferences from examples and classifier settings.config/fragments/signal/preference/Guide
reask — learned signalreask detects when the current user turn semantically repeats recent user turns in the same conversation.config/fragments/signal/reask/Guide
safety — learned signalThe safety signal predicts content risks.config/fragments/signal/safety/Guide
structure — heuristic signalstructure detects request-shape facts such as many explicit questions, ordered workflow markers, or dense constraint phrasing.config/fragments/signal/structure/Guide
user-feedback — learned signaluser-feedback detects correction, dissatisfaction, or escalation feedback from the conversation.config/fragments/signal/user-feedback/Guide

Selection algorithms​

Family and typeUse it toReusable fragmentGuide
automix — selection algorithmautomix is an experimental selector that ranks candidate models by configured quality and cost plus internal verification and escalation estimates.config/fragments/algorithm/selection/automix.yamlGuide
decision — selection algorithmdecision asks a decision model which of a routing decision's modelRefs should answer the request.config/fragments/algorithm/selection/decision.yamlGuide
hybrid — selection algorithmhybrid combines Elo ratings, Router-DC description similarity, AutoMix's one-model value estimate, and cost into one weighted candidate score.config/fragments/algorithm/selection/hybrid.yamlGuide
kmeans — selection algorithmkmeans sends a request to the model assigned to its nearest learned cluster.config/fragments/algorithm/selection/kmeans.yamlGuide
knn — selection algorithmknn chooses a candidate from the models that performed well on the most similar recorded requests.config/fragments/algorithm/selection/knn.yamlGuide
latency-aware — selection algorithmlatency_aware ranks eligible candidates using observed TTFT and TPOT percentiles and selects the lowest relative-latency score.config/fragments/algorithm/selection/latency-aware.yamlGuide
mlp — selection algorithmmlp runs a trained neural classifier on CPU to map a request to a candidate model.config/fragments/algorithm/selection/mlp.yamlGuide
multi-factor — selection algorithmmulti_factor chooses one candidate from quality, latency, cost, and load.config/fragments/algorithm/selection/multi-factor.yamlGuide
prompt — selection algorithmprompt uses a concrete helper model to select exactly one model from the matched decision's modelRefs.config/fragments/algorithm/selection/prompt.yamlGuide
router-dc — selection algorithmrouter_dc embeds the request and each model description, then selects the candidate with the strongest semantic similarity.config/fragments/algorithm/selection/router-dc.yamlGuide
static — selection algorithmstatic provides deterministic model choice without metrics or learned state.config/fragments/algorithm/selection/static.yamlGuide
svm — selection algorithmsvm uses a trained linear or RBF support-vector classifier to map request features to a candidate model.config/fragments/algorithm/selection/svm.yamlGuide

Looper algorithms​

Family and typeUse it toReusable fragmentGuide
confidence — looper algorithmconfidence tries candidate models in order and stops when response confidence reaches a configured threshold.config/fragments/algorithm/looper/confidence.yamlGuide
fusion — looper algorithmfusion asks several models to answer a request and a judge model to synthesize one final answer.config/fragments/algorithm/looper/fusion.yamlGuide
ratings — looper algorithmratings calls every candidate model and returns one OpenAI-compatible choice per successful model. max_concurrent limits parallel work; it does not limit the total number of candidates executed.config/fragments/algorithm/looper/ratings.yamlGuide
remom — looper algorithmremom runs several candidate models across bounded rounds and synthesizes their responses into one answer.config/fragments/algorithm/looper/remom.yamlGuide
workflows — looper algorithmworkflows runs a bounded, multi-step Router Flow behind one OpenAI-compatible model name.config/fragments/algorithm/looper/workflows.yamlGuide

Native System One algorithms​

Family and typeUse it toReusable fragmentGuide
cascade — native algorithmcascade tries declared decision models in order and returns a complete System One response when its acceptance rules pass.config/fragments/algorithm/native/cascade.yamlGuide

Plugins and bundles​

Family and typeUse it toReusable fragmentGuide
content-safety — plugin bundleContent Safety combines supported route-local safety plugins into one reusable policy.config/fragments/plugin/content-safety/Guide
context-compression — route pluginUse context_compression on a decision when old conversation text or large tool outputs cost more tokens than the answer needs.config/fragments/plugin/context-compression/Guide
fast-response — route pluginfast_response is a route-local plugin that returns a deterministic fallback message immediately.config/fragments/plugin/fast-response/Guide
hallucination — route pluginhallucination is a route-local plugin for fact-checking and response-quality screening after the decision already matched.config/fragments/plugin/hallucination/Guide
header-mutation — route pluginheader_mutation is a route-local plugin for adding, updating, or deleting downstream headers.config/fragments/plugin/header-mutation/Guide
memory — route pluginmemory is a route-local plugin for retrieving and storing conversation memory.config/fragments/plugin/memory/Guide
prompt-cache — route pluginprompt_cache is a route-local plugin that inserts Anthropic prompt-cache breakpoints on a matched route.config/fragments/plugin/prompt-cache/Guide
rag — route pluginrag retrieves external context for a matched route before generation.config/fragments/plugin/rag/Guide
request-params — route pluginrequest_params is a route-local plugin that validates and trims OpenAI Chat Completions request bodies before they are forwarded to backends.config/fragments/plugin/request-params/Guide
response-cache — route pluginresponse_cache is the route-local plugin for reusing exact or semantically compatible prior responses.config/fragments/plugin/response-cache/Guide
response-jailbreak — route pluginresponse_jailbreak is a route-local plugin for screening the model response before it is returned.config/fragments/plugin/response-jailbreak/Guide
router-replay — route pluginUse Router Replay to inspect requests in Dashboard Insights: the selected route, model, token usage, response, and tool trajectory.config/fragments/plugin/router-replay/Guide
shadow-dispatch — route pluginshadow_dispatch is a route-local plugin that sends a bounded, sampled copy of the approved request to a secondary model and records the outcome without changing or delaying the primary response.config/fragments/plugin/shadow-dispatch/Guide
system-prompt — route pluginsystem_prompt is a route-local plugin for inserting or modifying the system prompt on matched traffic.config/fragments/plugin/system-prompt/Guide
tool-selection — route plugintool_selection is a decision plugin that controls how tools are chosen for a matched route.config/fragments/plugin/tool-selection/Guide
tools — route plugintools is a route-local plugin for tool filtering and semantic tool selection.config/fragments/plugin/tools/Guide

Minimal example​

version: v0.3

listeners:
- name: http-8899
address: 0.0.0.0
port: 8899
timeout: 300s

providers:
defaults:
model: local/general
models:
- name: local/general
provider_model_id: my-served-model
backend_refs:
- name: primary
endpoint: host.docker.internal:8000
protocol: http
provider: vllm

routing:
strategy: priority
modelCards:
- name: local/general
modality: text
capabilities: [chat]
signals:
keywords:
- name: needs_explanation
operator: OR
keywords: ["explain", "walk me through"]
decisions:
- name: explanatory_answer
description: Prefer an explanatory answer when the request asks for one.
priority: 100
rules:
operator: AND
conditions:
- type: keyword
name: needs_explanation
modelRefs:
- model: local/general

global:
services:
observability:
metrics:
enabled: true

Model configuration​

Models can inherit identity and reasoning from the built-in catalog or define a private model locally. This catalog-backed example lets the selected Provider mapping supply the native model ID, protocol, reasoning transport, and request path:

providers:
defaults:
model: production
reasoning_effort: medium
models:
- name: production
catalog: openai/gpt-5.6-sol
backend_refs:
- provider: openai
api_key_env: OPENAI_API_KEY

The name remains the local Router alias. A private or newly released model omits catalog and can optionally define a Model Card and reasoning contract under that alias. Start with Configure models, then use Model configuration patterns to compare the catalog, custom, reasoning, Provider, and replica combinations. The Model and provider Day-0 guide is for contributors adding reusable support to the repository catalog.

Classifier backend failures remain Unknown while the complete boolean tree is evaluated. Set rules.on_unknown to no_match, match, or fail_request to resolve an undetermined terminal result. Omitting it preserves the existing classifier-family error behavior.

Requests using an automatic model alias enter the default routing profile. A concrete provider model name is a direct pass-through request and bypasses recipe signals, decisions, route plugins, cache, learning, and session routing.

Validate and serve​

vllm-sr config validate --config config.yaml
vllm-sr serve --config config.yaml

Validation catches schema errors, unresolved references, incompatible recipe boundaries, invalid provider bindings, and unsupported plugin or algorithm settings before the Router starts.

For portable model-free Recipes, set routing.decisions[].algorithm.minimum_candidates to the smallest pool that preserves the decision's intended behavior. Empty built-in assets remain valid, while a published Entrypoint is rejected if its concrete assignments do not meet the declared cardinality.

Environment references and secrets​

Keep credentials outside the YAML file:

api_key: ${MODEL_API_KEY}

Supported string substitutions are:

  • ${VAR} and $VAR;
  • ${VAR:-default} when VAR is unset or empty;
  • ${VAR-default} when VAR is unset; and
  • $$ for a literal $.

For a custom Recipe, authorize required host variables explicitly with --recipe-env NAME. Kubernetes deployments place sensitive environment values in Secrets rather than ConfigMaps or Helm values. See Security Hardening.

Entrypoints and recipes​

Without an explicit entrypoint for recipe: default, the top-level routing recipe is published as vllm-sr/auto. To replace that name, declare its model_names in entrypoints with recipe: default. Named recipes require their own entrypoints. MoM has no implicit special meaning.

An entrypoint maps one or more public model aliases to a recipe. A recipe owns its signal, projection, decision, algorithm, plugin, cache, learning, and routing state. Providers, stores, and router-owned classifier assets may be shared without allowing policy state to cross recipe boundaries.

Set max_response_bytes on external LLM classifier entries and the MCP classifier module to cap one upstream classifier response.

In the schema, entrypoints[].model_names lists the public aliases, entrypoints[].recipe selects a named recipe, and recipes[].routing contains that recipe's policy.

global.router.strategy and global.router.fallback provide shared defaults. Top-level routing configures the default recipe only; each named recipes[].routing resolves its own strategy and fallback independently from global.router. Named recipes never inherit the default recipe's routing settings. A decision's fallback overrides its own recipe's effective fallback. Sparse overrides such as fallback: {enabled: false} preserve the remaining shared defaults in both default and named recipes. The runtime strategy default is priority.

If no decision matches, the recipe uses providers.defaults.model. The virtual entrypoint name never reaches a backend.

See Models, Entrypoints, and Serving for built-in virtual models, CLI serving, backend binding, forking, packaging, and migration. See Virtual Models for the complete schema.

Recipe candidate requirements​

Set optional candidate requirements inside the default or a named recipe's routing block:

candidate_requirements:
capabilities: declared
context: known_limits

capabilities: declared requires the assigned model to declare support for the request's task, including tools and image input, as well as a compatible provider protocol. context: known_limits checks estimated input demand plus the effective output reserve against declared model limits. The request must supply an output bound, or its decision must configure request_params.default_max_tokens as a positive integer or auto (see below). That default applies only when the caller omits the bound; max_tokens_limit then caps it as usual. A model's maximum output capacity is not a request default. Missing required model facts or an effective output bound make a candidate ineligible. Input accounting remains estimated, especially for multimodal content; this is not an exact provider token-capacity guarantee. Omit a field to retain that dimension's existing compatibility behavior.

For example, a decision can supply the bound through its existing plugin:

plugins:
- type: request_params
configuration:
default_max_tokens: 4096
max_tokens_limit: 8192

To use each model's available output capacity when the caller omits a limit:

algorithm:
type: multi_factor
multi_factor:
expected_output_tokens: 4096
plugins:
- type: request_params
configuration:
default_max_tokens: auto

auto requires an explicit vllm provider, one backend per model, the OpenAI Chat format, and declared context and output limits. Enable the backend's /v1/chat/completions/render API with vllm serve --enable-scale-out on a vLLM version that supports it. The Router preprocesses each candidate's complete provider request without generating an answer, then repeats that check after final request changes. It uses the actual templated input length to allocate min(max_output_tokens, context_window - input_tokens). A smaller backend limit than declared is a configuration error; align the deployment and model card. Explicit caller limits keep their existing behavior and need no render call. max_tokens_limit still caps the resulting output budget.

expected_output_tokens is a positive cost forecast, not an output limit. It is required for multi_factor with auto, and is bounded by the caller's limit or each candidate's available capacity. Different maximum context windows do not by themselves make a model's forecast more expensive. Without this field, non-automatic requests retain their existing cost calculation.

A request that fits a candidate is not truncated based on a text-length estimate. If every compatible candidate reports input overflow, the Router applies the configured context compression policy once and verifies the result with the backend. Compression can be conservative; final capacity checking is exact. Without a suitable compression policy, overflow remains an error. An unavailable or invalid render API fails clearly and never triggers compression. Preview reports execution_required because it does not have the complete provider request. Automatic budgets currently support text and tool requests, excluding multimodal input, provider truncation, LoRA, Looper, and shadow dispatch.

For multi-factor selection, latency_metric: ttft compares time to first token; tpot compares time per output token. Omission preserves the existing TPOT-then-TTFT fallback. Pair the metric with explicit quality evidence and a lexicographic objective when quality is a floor rather than a score to trade away.

Discover the current contract with vllm-sr config schema --section routing.candidate_requirements. DSL ROUTING blocks support the same object. Kubernetes CRD emission preserves these requirements for the default routing profile; named recipes and entrypoints require canonical YAML and are rejected by CRD emission rather than discarded.

Replay capture defaults and decision overrides​

global.services.router_replay owns shared storage, retention, enablement and capture defaults. A decision's router_replay plugin overrides only explicitly configured capture fields. Omitted fields inherit; enabled: false disables capture for that decision, while enabled: true can opt in when the global default is disabled. There is no routing.data_policy or recipe-level Replay configuration.

Set capture_personal_data: false globally or in the decision plugin to retain routing evidence while omitting content when PII is detected or its status is unknown. Missing PII detectors suppress content conservatively rather than preventing startup. Rejected requests without a selected decision use global capture defaults. These controls do not govern other stores or backend retention. See Router Replay for the complete fields and examples, or inspect vllm-sr config schema --section global.services.router_replay.

Configuration workflows​

The canonical document can be authored or applied through several interfaces:

  • local CLI and YAML;
  • Dashboard setup and visual routing tools;
  • Helm or vllm-sr serve --target k8s;
  • the Kubernetes Operator; and
  • the routing DSL.

Configuration Workflows explains which interface owns which part of the document and how to avoid competing sources of truth. Configuration Contract describes the generated machine-readable schema, Router discovery and validation APIs, and the safe authoring loop for tools and agents. Configuration Management explains how a change activates on a running Router, how a rejected change is reported, and how to list and roll back versions.

Reference sources​

Avoid copying the exhaustive example as an application config. Start with the smallest document that describes the deployment, then add only the capabilities and services it uses.