Skip to main content
Version: Latest (unreleased)

Algorithms

Overview​

An algorithm runs after a decision matches. It either selects one model from the decision's modelRefs, coordinates Chat models through the Looper, or executes a native System One cascade. It does not decide whether the route is eligible; signals and decisions do that first.

Key Advantages​

  • Keeps route eligibility separate from model choice.
  • Makes selection and orchestration policy reviewable per decision.
  • Supports both stateless policies and bounded multi-model execution.

What Problem Does It Solve?​

A matched route may have several valid model candidates. Algorithms make the choice explicit: fixed ordering, semantic fit, observed latency, multiple runtime factors, a learned selector, or multi-model orchestration.

When to Use​

Add an algorithm when a decision has more than one candidate or deliberately runs a multi-model workflow. With one candidate, omit the algorithm unless the chosen Looper supports and needs a single-model execution plan.

Configuration​

Algorithms are decision-local:

routing:
decisions:
- name: responsive-route
description: Prefer the model with the best observed latency.
priority: 100
rules:
operator: AND
conditions: []
modelRefs:
- model: small-model
- model: large-model
algorithm:
type: latency_aware
minimum_candidates: 2
latency_aware:
tpot_percentile: 90
ttft_percentile: 95

Choose an algorithm from the inventory below, then follow its guide for the required fields and dependencies.

minimum_candidates is a common algorithm field. A model-free Recipe may declare it before any Models are bound; once an Entrypoint materializes the Recipe, validation requires that many distinct modelRefs. The same boundary is checked again after request-time context eligibility filtering, so a panel, cascade, or selector does not silently run with a smaller pool than its policy declares.

Algorithm Inventory​

Selection Algorithms​

Selection algorithms return one candidate model. For a full inference request, the Router filters exact candidate references by context, backend wire support, and declared model task capabilities before scoring. The configured algorithm compares the surviving pool; it does not choose an incapable winner and then replace it with the first compatible sibling.

Capability checks include Router-retained conversation content and preview the decision's request-parameter and no-tools policies without executing those plugins twice. Unannotated models retain wire-only compatibility checks. Hard quality/SLO constraints still apply, and Router Learning checks the same capabilities when considering additional candidates.

An empty pool or a violated minimum_candidates requirement fails closed. Dispatch validates the final request again after mutations: a late capability mismatch returns an error rather than restarting decision evaluation, rescoring, or silently changing the selected candidate. Explicitly pinned models use the same final validation but are not replaced by another model.

TypeStatusGoalMain dependencyGuide
decisionsupportedAsk a typed choice question over the matched candidatesDecision deployment with Choice supportDecision
staticsupportedUse declared order or fixed domain scoresNoneStatic
router_dcsupportedMatch request semantics to model descriptionsEmbedding runtime and useful model cardsRouter DC
latency_awaresupportedPrefer the candidate with the best observed TTFT/TPOTPer-process latency observationsLatency Aware
multi_factorsupportedBalance quality, latency, cost, and load with optional SLO filtersModel metadata and live local metricsMulti Factor
hybridsupportedBlend several selector scoresComponent selector inputsHybrid
automixexperimentalOptimize an estimated cost-quality valueCandidate pricing and quality metadataAutoMix
gmtrouterexperimentalPersonalize an intelligence-seeded model rankModel evidence and user feedbackGMT Router
promptexperimentalLet a bounded helper model choose from declared candidatesOpenAI-compatible helper modelPrompt
knnexperimentalFollow similar labeled examplesTrained selector artifact and embeddingsKNN
kmeansexperimentalRoute through learned traffic clustersTrained selector artifact and embeddingsKMeans
svmexperimentalApply a learned decision boundaryTrained selector artifact and embeddingsSVM
mlpexperimentalApply a learned nonlinear classifierTrained selector artifactMLP

Looper Algorithms​

Looper algorithms make additional model calls. The Router makes each call in process: the call runs the matched decision's plugins and goes to the called model's providers.models[].backend_refs, so every model a Looper calls needs a backend. The calls increase latency and token usage, and intermediate content is sent to every configured worker involved in the run.

global.integrations.looper.endpoint is deprecated and ignored, because Looper calls no longer loop back through the gateway. Remove it, or run vllm-sr config migrate.

TypeStatusGoalGuide
confidencesupportedEscalate sequentially until confidence clears a thresholdConfidence
ratingssupportedReturn one choice from each candidate with bounded concurrencyRatings
remomsupportedExplore several reasoning paths over multiple rounds, then synthesizeReMoM
fusionexperimentalRun an analysis panel and judge/synthesis passFusion
workflowsexperimentalExecute a bounded static or planner-generated worker flowRouter Flow

Treat experimental algorithms as evaluation features: validate them on your traffic before using them for production routing.

Operational Boundaries​

  • Where an algorithm accepts repeated model references, the candidate includes its LoRA and reasoning controls, not just its model name. Scoring, composition, and dispatch retain the exact winning reference. A legacy model-only result that matches multiple different candidates is rejected rather than resolved to the first reference. For a LoRA candidate, the serving frontend routes by the selected base model while the provider request names the adapter.
  • Router Learning session memory retains the selected candidate's controls. Protection can hold that exact choice across tool-loop continuations even when a later base selection prefers another effort of the same model. If that exact owner is excluded during an active tool loop or nonportable continuation, routing rejects the request rather than restoring the owner or falling back to a different effort. Portable turns may select a new eligible candidate; observe and bypass modes do not enforce the protection decision.
  • Candidate model names must resolve through routing.modelCards and providers.models in a complete config.
  • Learned selectors need artifacts produced for the same embedding dimension and candidate labels used at runtime.
  • Latency and load observations are local to a Router process; they are not a cluster-wide scheduler.
  • Looper algorithms share request content with their configured workers. Apply privacy and provider-boundary decisions before choosing them.
  • Looper-generated planner, worker, verifier, judge, and synthesis prompts are checked against each target Model's known context window before dispatch. Missing context metadata remains eligible for compatibility.
  • Validate a complete config with vllm-sr config validate --config config.yaml.

Native System One Algorithms​

A recipe published with api: systemone uses native execution. Its matched decision runs a cascade, preserving complete typed responses. Each cascade has its own algorithm.budget for its stages and transport retries. Signals run first with their own timeouts and request cancellation.

TypeStatusGoalGuide
cascadeexperimentalTry declared native models in authored order until acceptance passesCascade

A one-stage cascade can target one model; a longer cascade can try a small model before a stronger one. Several decisions can choose different cascades within one recipe. Native execution uses an explicit candidate roster and quality contract; it does not run Chat plugins or silently fall back to a Chat model.