Algorithms
Overview
An algorithm runs after a decision matches. It either selects one model from
the decision's modelRefs, coordinates Chat models through the Looper, or
executes a native System One cascade.
It does not decide whether the route is eligible; signals and decisions do
that first.
Key Advantages
- Keeps route eligibility separate from model choice.
- Makes selection and orchestration policy reviewable per decision.
- Supports both stateless policies and bounded multi-model execution.
What Problem Does It Solve?
A matched route may have several valid model candidates. Algorithms make the choice explicit: fixed ordering, semantic fit, observed latency, multiple runtime factors, a learned selector, or multi-model orchestration.
When to Use
Add an algorithm when a decision has more than one candidate or deliberately runs a multi-model workflow. With one candidate, omit the algorithm unless the chosen Looper supports and needs a single-model execution plan.
Configuration
Algorithms are decision-local:
routing:
decisions:
- name: responsive-route
description: Prefer the model with the best observed latency.
priority: 100
rules:
operator: AND
conditions: []
modelRefs:
- model: small-model
- model: large-model
algorithm:
type: latency_aware
minimum_candidates: 2
latency_aware:
tpot_percentile: 90
ttft_percentile: 95
Choose an algorithm from the inventory below, then follow its guide for the required fields and dependencies.
minimum_candidates is a common algorithm field. A model-free Recipe may
declare it before any Models are bound; once an Entrypoint materializes the
Recipe, validation requires that many distinct modelRefs. The same boundary
is checked again after request-time context eligibility filtering, so a panel,
cascade, or selector does not silently run with a smaller pool than its policy
declares.
Algorithm Inventory
Selection Algorithms
Selection algorithms return one candidate model. For a full inference request, the Router filters exact candidate references by context, backend wire support, and declared model task capabilities before scoring. The configured algorithm compares the surviving pool; it does not choose an incapable winner and then replace it with the first compatible sibling.
Capability checks include Router-retained conversation content and preview the decision's request-parameter and no-tools policies without executing those plugins twice. Unannotated models retain wire-only compatibility checks. Hard quality/SLO constraints still apply, and Router Learning checks the same capabilities when considering additional candidates.
An empty pool or a violated minimum_candidates requirement fails closed.
Dispatch validates the final request again after mutations: a late capability
mismatch returns an error rather than restarting decision evaluation, rescoring,
or silently changing the selected candidate. Explicitly pinned models use the
same final validation but are not replaced by another model.
| Type | Status | Goal | Main dependency | Guide |
|---|---|---|---|---|
decision | supported | Ask a typed choice question over the matched candidates | Decision deployment with Choice support | Decision |
static | supported | Use declared order or fixed domain scores | None | Static |
router_dc | supported | Match request semantics to model descriptions | Embedding runtime and useful model cards | Router DC |
latency_aware | supported | Prefer the candidate with the best observed TTFT/TPOT | Per-process latency observations | Latency Aware |
multi_factor | supported | Balance quality, latency, cost, and load with optional SLO filters | Model metadata and live local metrics | Multi Factor |
hybrid | supported | Blend several selector scores | Component selector inputs | Hybrid |
automix | experimental | Optimize an estimated cost-quality value | Candidate pricing and quality metadata | AutoMix |
gmtrouter | experimental | Personalize an intelligence-seeded model rank | Model evidence and user feedback | GMT Router |
prompt | experimental | Let a bounded helper model choose from declared candidates | OpenAI-compatible helper model | Prompt |
knn | experimental | Follow similar labeled examples | Trained selector artifact and embeddings | KNN |
kmeans | experimental | Route through learned traffic clusters | Trained selector artifact and embeddings | KMeans |
svm | experimental | Apply a learned decision boundary | Trained selector artifact and embeddings | SVM |
mlp | experimental | Apply a learned nonlinear classifier | Trained selector artifact | MLP |
Looper Algorithms
Looper algorithms make additional model calls. The Router makes each call in
process: the call runs the matched decision's plugins and goes to the called
model's providers.models[].backend_refs, so every model a Looper calls needs
a backend. The calls increase latency and token usage, and intermediate content
is sent to every configured worker involved in the run.
global.integrations.looper.endpoint is deprecated and ignored, because Looper
calls no longer loop back through the gateway. Remove it, or run
vllm-sr config migrate.
| Type | Status | Goal | Guide |
|---|---|---|---|
confidence | supported | Escalate sequentially until confidence clears a threshold | Confidence |
ratings | supported | Return one choice from each candidate with bounded concurrency | Ratings |
remom | supported | Explore several reasoning paths over multiple rounds, then synthesize | ReMoM |
fusion | experimental | Run an analysis panel and judge/synthesis pass | Fusion |
workflows | experimental | Execute a bounded static or planner-generated worker flow | Router Flow |
Treat experimental algorithms as evaluation features: validate them on your traffic before using them for production routing.
Operational Boundaries
- Where an algorithm accepts repeated model references, the candidate includes its LoRA and reasoning controls, not just its model name. Scoring, composition, and dispatch retain the exact winning reference. A legacy model-only result that matches multiple different candidates is rejected rather than resolved to the first reference. For a LoRA candidate, the serving frontend routes by the selected base model while the provider request names the adapter.
- Router Learning session memory retains the selected candidate's controls. Protection can hold that exact choice across tool-loop continuations even when a later base selection prefers another effort of the same model. If that exact owner is excluded during an active tool loop or nonportable continuation, routing rejects the request rather than restoring the owner or falling back to a different effort. Portable turns may select a new eligible candidate; observe and bypass modes do not enforce the protection decision.
- Candidate model names must resolve through
routing.modelCardsandproviders.modelsin a complete config. - Learned selectors need artifacts produced for the same embedding dimension and candidate labels used at runtime.
- Latency and load observations are local to a Router process; they are not a cluster-wide scheduler.
- Looper algorithms share request content with their configured workers. Apply privacy and provider-boundary decisions before choosing them.
- Looper-generated planner, worker, verifier, judge, and synthesis prompts are checked against each target Model's known context window before dispatch. Missing context metadata remains eligible for compatibility.
- Validate a complete config with
vllm-sr config validate --config config.yaml.
Native System One Algorithms
A recipe published with api: systemone uses native execution. Its matched
decision runs a cascade, preserving complete typed responses. Each cascade
has its own algorithm.budget for its stages and transport retries. Signals
run first with their own timeouts and request cancellation.
| Type | Status | Goal | Guide |
|---|---|---|---|
cascade | experimental | Try declared native models in authored order until acceptance passes | Cascade |
A one-stage cascade can target one model; a longer cascade can try a small model before a stronger one. Several decisions can choose different cascades within one recipe. Native execution uses an explicit candidate roster and quality contract; it does not run Chat plugins or silently fall back to a Chat model.