Algorithms
Overview
An algorithm runs after a decision matches. It either selects one model from
the decision's modelRefs or coordinates several of them through the Looper.
It does not decide whether the route is eligible; signals and decisions do
that first.
Key Advantages
- Keeps route eligibility separate from model choice.
- Makes selection and orchestration policy reviewable per decision.
- Supports both stateless policies and bounded multi-model execution.
What Problem Does It Solve?
A matched route may have several valid model candidates. Algorithms make the choice explicit: fixed ordering, semantic fit, observed latency, multiple runtime factors, a learned selector, or multi-model orchestration.
When to Use
Add an algorithm when a decision has more than one candidate or deliberately runs a multi-model workflow. With one candidate, omit the algorithm unless the chosen Looper supports and needs a single-model execution plan.
Configuration
Algorithms are decision-local:
routing:
decisions:
- name: responsive-route
description: Prefer the model with the best observed latency.
priority: 100
rules:
operator: AND
conditions: []
modelRefs:
- model: small-model
- model: large-model
algorithm:
type: latency_aware
minimum_candidates: 2
latency_aware:
tpot_percentile: 90
ttft_percentile: 95
Choose an algorithm from the inventory below, then follow its guide for the required fields and dependencies.
minimum_candidates is a common algorithm field. A model-free Recipe may
declare it before any Models are bound; once an Entrypoint materializes the
Recipe, validation requires that many distinct modelRefs. The same boundary
is checked again after request-time context eligibility filtering, so a panel,
cascade, or selector does not silently run with a smaller pool than its policy
declares.
Algorithm Inventory
Selection Algorithms
Selection algorithms return one candidate model. For a full inference request, the Router filters exact candidate references by context, backend wire support, and declared model task capabilities before scoring. The configured algorithm compares the surviving pool; it does not choose an incapable winner and then replace it with the first compatible sibling.
Capability checks include Router-retained conversation content and preview the decision's request-parameter and no-tools policies without executing those plugins twice. Unannotated models retain wire-only compatibility checks. Hard quality/SLO constraints still apply, and Router Learning checks the same capabilities when considering additional candidates.
An empty pool or a violated minimum_candidates requirement fails closed.
Dispatch validates the final request again after mutations: a late capability
mismatch returns an error rather than restarting decision evaluation, rescoring,
or silently changing the selected candidate. Explicitly pinned models use the
same final validation but are not replaced by another model.
| Type | Status | Goal | Main dependency | Guide |
|---|---|---|---|---|
static | supported | Use declared order or fixed domain scores | None | Static |
router_dc | supported | Match request semantics to model descriptions | Embedding runtime and useful model cards | Router DC |
latency_aware | supported | Prefer the candidate with the best observed TTFT/TPOT | Per-process latency observations | Latency Aware |
multi_factor | supported | Balance quality, latency, cost, and load with optional SLO filters | Model metadata and live local metrics | Multi Factor |
hybrid | supported | Blend several selector scores | Component selector inputs | Hybrid |
automix | experimental | Optimize an estimated cost-quality value | Candidate pricing and quality metadata | AutoMix |
gmtrouter | experimental | Personalize an intelligence-seeded model rank | Model evidence and user feedback | GMT Router |
prompt | experimental | Let a bounded helper model choose from declared candidates | OpenAI-compatible helper model and Looper endpoint | Prompt |
knn | experimental | Follow similar labeled examples | Trained selector artifact and embeddings | KNN |
kmeans | experimental | Route through learned traffic clusters | Trained selector artifact and embeddings | KMeans |
svm | experimental | Apply a learned decision boundary | Trained selector artifact and embeddings | SVM |
mlp | experimental | Apply a learned nonlinear classifier | Trained selector artifact | MLP |
Looper Algorithms
Looper algorithms make additional model calls through
global.integrations.looper.endpoint. They increase latency and token usage,
and intermediate content is sent to every configured worker involved in the
run.
| Type | Status | Goal | Guide |
|---|---|---|---|
confidence | supported | Escalate sequentially until confidence clears a threshold | Confidence |
ratings | supported | Return one choice from each candidate with bounded concurrency | Ratings |
remom | supported | Explore several reasoning paths over multiple rounds, then synthesize | ReMoM |
fusion | experimental | Run an analysis panel and judge/synthesis pass | Fusion |
workflows | experimental | Execute a bounded static or planner-generated worker flow | Router Flow |
Treat experimental algorithms as evaluation features: validate them on your traffic before using them for production routing.
Operational Boundaries
- Where an algorithm accepts repeated model references, the candidate includes its LoRA and reasoning controls, not just its model name. Scoring, composition, and dispatch retain the exact winning reference. A legacy model-only result that matches multiple different candidates is rejected rather than resolved to the first reference.
- Router Learning session memory retains the selected candidate's controls. Protection can hold that exact choice across tool-loop continuations even when a later base selection prefers another effort of the same model. If that exact owner is excluded during an active tool loop or nonportable continuation, routing rejects the request rather than restoring the owner or falling back to a different effort. Portable turns may select a new eligible candidate; observe and bypass modes do not enforce the protection decision.
- Candidate model names must resolve through
routing.modelCardsandproviders.modelsin a complete config. - Learned selectors need artifacts produced for the same embedding dimension and candidate labels used at runtime.
- Latency and load observations are local to a Router process; they are not a cluster-wide scheduler.
- Looper algorithms share request content with their configured workers. Apply privacy and provider-boundary decisions before choosing them.
- Looper-generated planner, worker, verifier, judge, and synthesis prompts are checked against each target Model's known context window before dispatch. Missing context metadata remains eligible for compatibility.
- Validate a complete config with
vllm-sr config validate --config config.yaml.