Hybrid
Overview
hybrid combines Elo ratings, Router-DC description similarity, AutoMix's
one-model value estimate, and cost into one weighted candidate score.
Paper: Hybrid LLM: Cost-Efficient Quality-Aware Query Routing
Key Advantages
- Blends multiple selectors instead of committing to only one.
- Makes weighting explicit and easy to audit.
- Makes it possible to introduce one component gradually by changing its weight.
- Cost-aware scoring to balance quality and operational expense.
Algorithm Principle
Hybrid first min-max normalizes the available Elo, Router-DC, and AutoMix
scores when normalize_scores is enabled. It combines those components using
their relative weights, renormalized across the components that returned data.
It then applies a multiplicative bonus to cheaper models when cost adjustment
is enabled. Cost is therefore a second-stage adjustment, not another linear
term in the component average.
Cache affinity is the last, bounded adjustment. For a request that continues a
session, it favors the model that served the previous turn so the backend's
prompt cache stays useful, and it fades to zero as the gap between the top two
base scores grows. A prompt_cache_key also marks a request as a continuation.
When the Router has no previous model for the session, it favors the model that
last served the same key in the recipe. Each Router replica remembers that model
for an hour after the key's last response.
Select Flow
Component Selectors
The Hybrid selector internally instantiates three sub-selectors:
| Component | Source | What it provides |
|---|---|---|
EloSelector | Its own in-memory ratings | Relative model rating |
RouterDCSelector | Model descriptions | Semantic query-model similarity |
AutoMixSelector | One-shot request path | Cost-quality value estimate |
Each component shares the same SelectionContext and runs independently.
Composition uses typed scores attached to the complete candidate reference, not
model-name or display-label lookups. Two references to the same model at low
and high reasoning effort therefore remain separate through normalization,
comparison, Router Learning protection, and provider request encoding.
Elo, description similarity, and cache observations can remain model-level signals; their values are explicitly applied to each candidate variant. AutoMix contributes its raw expected value, before its separate cost-aware starting-model adjustment. Neither those model-level priors nor learned feedback are presented as effort-specific benchmark measurements.
What Problem Does It Solve?
No single ranking signal is reliable for every workload: pure cost, pure similarity, or pure feedback each misses part of the routing picture. hybrid combines multiple selectors into one auditable score so routes can balance semantic fit, historical quality, and operational cost.
When to Use
- One route should combine several ranking signals.
- You want a weighted transition between older and newer selectors.
- No single selector captures all relevant information.
- The final choice should reflect both quality and operational cost.
Known Limitations
- Higher computational cost than any single selector (runs 3 sub-selectors per request).
- Weight tuning requires domain knowledge — suboptimal weights can degrade performance.
Configuration
algorithm:
type: hybrid
hybrid:
experience_weight: 0.3 # Elo component weight
router_dc_weight: 0.3 # Weight for embedding similarity
automix_weight: 0.2 # Weight for AutoMix's one-model value
cost_weight: 0.2 # Weight for cost consideration
normalize_scores: true # Normalize component scores to [0,1]
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
experience_weight | float | 0.3 | Weight for the Elo selector score (0–1) |
router_dc_weight | float | 0.3 | Weight for RouterDC embedding similarity (0–1) |
automix_weight | float | 0.2 | Weight for AutoMix's one-model value estimate (0–1) |
cost_weight | float | 0.2 | Weight for cost consideration (0–1) |
quality_gap_threshold | float | 0.1 | Accepted for compatibility; it has no effect in the current online selector |
normalize_scores | bool | true | Normalize component scores before combination |
Feedback
Hybrid does not read Router Learning snapshots or
global.router.learning.adaptation. Its Elo, Router-DC, and AutoMix components
own separate in-memory state, and the current Router Learning outcome endpoint
does not feed that state automatically.
Request text is embedded for Router-DC and AutoMix components. Missing model
descriptions, pricing, or initialized component state make the corresponding component
less informative, so tune weights against the data actually available. The
complete example is
config/fragments/algorithm/selection/hybrid.yaml.