Skip to main content
Version: Latest (unreleased)

Multi Factor

Overview​

multi_factor chooses one candidate from quality, latency, cost, and load. Hard eligibility rules run first; the surviving candidates are then compared with a weighted or lexicographic objective.

FactorSourceDirection
QualityVersioned Overall, capability, or operator indexHigher is better
LatencyObserved TTFT or TPOT at the selected percentileLower is better
CostInput/output pricing applied to this request's token budgetLower is better
LoadCurrent in-flight requests in this Router processLower is better

Quality is resolved for the candidate's exact reasoning effort. A score from a different effort is never borrowed. Coverage is not part of the objective; when both the final objective value and intelligence score tie, higher coverage is the deterministic tie-breaker.

Set latency_metric: ttft to favor a fast first token, or tpot to favor fast streaming after generation starts. With either setting, a missing measurement stays unknown; the other metric is never substituted. Lexicographic selection compares latency only when every candidate still in the objective has a measurement. With no measurements or partial coverage, it skips latency and continues to the next priority. A measured model therefore cannot exclude an unmeasured model simply because it received the first request after startup. Omitting the setting preserves the existing TPOT-then-TTFT behavior.

What Problem Does It Solve?​

A model pool often contains a stronger model, a cheaper model, and a faster model. This selector keeps eligibility explicit and makes that trade-off one auditable decision policy.

When to Use​

Use multi_factor when a decision has at least two candidate models and must optimize a quality/latency/cost/load trade-off. Use static when order alone is the policy, or latency_aware when latency is the only selection signal.

Configuration​

Balanced​

The default weighted strategy min-max normalizes each available factor over the surviving pool and computes:

S(m)=wQQ^(m)+wT(1−T^(m))+wC(1−C^(m))+wL(1−L^(m))S(m)=w_Q\hat Q(m)+w_T(1-\hat T(m))+w_C(1-\hat C(m))+w_L(1-\hat L(m))
algorithm:
type: multi_factor
multi_factor:
objective:
strategy: weighted
quality:
index: vllm-sr/coding@1.0.0
on_missing: exclude
min_coverage: 1.0
weights:
quality: 0.4
latency: 0.2
cost: 0.2
load: 0.2

Weights are normalized to sum to one. If every configured weight is zero, the selector recovers to equal weights.

Accuracy-first​

lexicographic applies priorities in order. Each stage keeps candidates within the declared relative tolerance of the best observed value, then passes that band to the next stage.

Latency requires complete measurement coverage in the current band. This does not relax quality eligibility or cost restrictions; missing quality and cost evidence continue to follow their existing policies.

Only candidates surviving every priority remain eligible for later adaptation, session protection, and dispatch. These steps cannot restore a model excluded by an earlier quality or cost band. Recorded scores may still include excluded models to explain the selection.

Preview returns selection_trace for lexicographic selection. Insights saves the same structure under route_diagnostics.selection_trace: reached stages, measurement coverage, available values, skipped stages, elimination reasons, and the final objective survivors. Unknown values stay absent, while measured zero remains zero. These are the base selector's candidates before Router Learning; the actual selected model is reported separately. Live measurements can change after a Preview, so its trace explains that observation rather than guaranteeing a later route.

algorithm:
type: multi_factor
multi_factor:
objective:
strategy: lexicographic
priorities:
- {factor: quality, tolerance: 0.03}
- {factor: cost, tolerance: 0.05}
- {factor: latency, tolerance: 0.05}
quality:
index: vllm-sr/intelligence@1.0.0
on_missing: exclude
min_coverage: 1.0

This keeps models within 3% of the best quality score, then chooses the cheaper and faster candidate inside that band.

Cost-first with a quality floor​

algorithm:
type: multi_factor
multi_factor:
quality:
index: vllm-sr/intelligence@1.0.0
on_missing: exclude
min_coverage: 1.0
min_score: 65
objective:
strategy: lexicographic
priorities:
- {factor: cost, tolerance: 0.05}
- {factor: quality, tolerance: 0.03}
- {factor: latency, tolerance: 0.05}

The quality floor excludes models below 65 before optimization. The objective then keeps the models within 5% of the cheapest estimate and chooses the best quality inside that cost band.

weights cannot be combined with a lexicographic objective. A priority factor may appear only once.

Quality eligibility and missing data​

quality:
index: acme/clinical-quality@1.0.0
on_missing: exclude
min_coverage: 0.5
min_score: 60
  • index selects a versioned built-in or operator index.
  • min_coverage applies after the index's own missing-data policy. It can make a route stricter than the index definition.
  • min_score is a hard floor on the index's declared scale and requires on_missing: exclude.
  • exclude removes a candidate without qualifying exact-effort evidence.
  • disable_quality keeps the pool, but one missing candidate disables quality for the entire comparison. It never changes weights for only one model.
  • Selection diagnostics report the chosen intelligence score and coverage, or mark intelligence evidence unavailable; missing evidence is not shown as zero.

See Open Intelligence Index for the built-in hierarchy and Custom evaluations for operator-defined evidence.

Cost and SLOs​

Input cost is calculated from the request's pre-routing token estimate. Output cost uses max_output_tokens when the caller supplies it. If neither is known, the selector falls back to the configured per-million-token rates. This same request mix is used for the cost factor, max_cost_per_1m, and the cheapest fallback.

multi_factor:
slo:
max_tpot_ms: 200
max_ttft_ms: 800
max_cost_per_1m: 5.0
max_inflight: 50
latency_percentile: 95
on_no_candidates: cheapest

SLO ceilings are enforced only when the corresponding observation is available. If every candidate is excluded, on_no_candidates chooses cheapest, first, or fail. fail is a strict policy: the Router returns HTTP 503 and never substitutes the first configured candidate. Eval dry-runs report the selection as unavailable for the same request.

Parameters​

ParameterDefaultMeaning
objective.strategyweightedweighted or lexicographic
objective.priorities[].factor—quality, latency, cost, or load
objective.priorities[].tolerance0Relative band from the best value, from 0 to 1
quality.indexCatalog defaultVersioned quality index
quality.on_missingexclude when explicitexclude or pool-wide disable_quality
quality.min_coverageIndex policyMinimum admitted evidence coverage, from 0 to 1
quality.min_scoreOffHard quality floor on the index scale
weights.*0.25Balanced quality, latency, cost, and load weights
latency_percentile95Observed latency percentile, from 1 to 100
latency_metricTPOT, then TTFTCompare ttft or tpot consistently across candidates
on_no_candidatescheapestcheapest, first, or fail

Latency and load observations are local to each Router process. Replicas may therefore make different choices. Use shared telemetry or an infrastructure router when fleet-global capacity is required.

See the complete fragment: config/fragments/algorithm/selection/multi-factor.yaml.