Skip to main content
Version: Latest (unreleased)

Hybrid

Overview​

hybrid combines Elo ratings, Router-DC description similarity, AutoMix's one-model value estimate, and cost into one weighted candidate score.

Paper: Hybrid LLM: Cost-Efficient Quality-Aware Query Routing

Key Advantages​

  • Blends multiple selectors instead of committing to only one.
  • Makes weighting explicit and easy to audit.
  • Makes it possible to introduce one component gradually by changing its weight.
  • Cost-aware scoring to balance quality and operational expense.

Algorithm Principle​

Hybrid first min-max normalizes the available Elo, Router-DC, and AutoMix scores when normalize_scores is enabled. It combines those components using their relative weights, renormalized across the components that returned data. It then applies a multiplicative bonus to cheaper models when cost adjustment is enabled. Cost is therefore a second-stage adjustment, not another linear term in the component average.

Cache affinity is the last, bounded adjustment. For a request that continues a session, it favors the model that served the previous turn so the backend's prompt cache stays useful, and it fades to zero as the gap between the top two base scores grows. A prompt_cache_key also marks a request as a continuation. When the Router has no previous model for the session, it favors the model that last served the same key in the recipe. Each Router replica remembers that model for an hour after the key's last response.

Select Flow​

Component Selectors​

The Hybrid selector internally instantiates three sub-selectors:

ComponentSourceWhat it provides
EloSelectorIts own in-memory ratingsRelative model rating
RouterDCSelectorModel descriptionsSemantic query-model similarity
AutoMixSelectorOne-shot request pathCost-quality value estimate

Each component shares the same SelectionContext and runs independently. Composition uses typed scores attached to the complete candidate reference, not model-name or display-label lookups. Two references to the same model at low and high reasoning effort therefore remain separate through normalization, comparison, Router Learning protection, and provider request encoding.

Elo, description similarity, and cache observations can remain model-level signals; their values are explicitly applied to each candidate variant. AutoMix contributes its raw expected value, before its separate cost-aware starting-model adjustment. Neither those model-level priors nor learned feedback are presented as effort-specific benchmark measurements.

What Problem Does It Solve?​

No single ranking signal is reliable for every workload: pure cost, pure similarity, or pure feedback each misses part of the routing picture. hybrid combines multiple selectors into one auditable score so routes can balance semantic fit, historical quality, and operational cost.

When to Use​

  • One route should combine several ranking signals.
  • You want a weighted transition between older and newer selectors.
  • No single selector captures all relevant information.
  • The final choice should reflect both quality and operational cost.

Known Limitations​

  • Higher computational cost than any single selector (runs 3 sub-selectors per request).
  • Weight tuning requires domain knowledge — suboptimal weights can degrade performance.

Configuration​

algorithm:
type: hybrid
hybrid:
experience_weight: 0.3 # Elo component weight
router_dc_weight: 0.3 # Weight for embedding similarity
automix_weight: 0.2 # Weight for AutoMix's one-model value
cost_weight: 0.2 # Weight for cost consideration
normalize_scores: true # Normalize component scores to [0,1]

Parameters​

ParameterTypeDefaultDescription
experience_weightfloat0.3Weight for the Elo selector score (0–1)
router_dc_weightfloat0.3Weight for RouterDC embedding similarity (0–1)
automix_weightfloat0.2Weight for AutoMix's one-model value estimate (0–1)
cost_weightfloat0.2Weight for cost consideration (0–1)
quality_gap_thresholdfloat0.1Accepted for compatibility; it has no effect in the current online selector
normalize_scoresbooltrueNormalize component scores before combination

Feedback​

Hybrid does not read Router Learning snapshots or global.router.learning.adaptation. Its Elo, Router-DC, and AutoMix components own separate in-memory state, and the current Router Learning outcome endpoint does not feed that state automatically.

Request text is embedded for Router-DC and AutoMix components. Missing model descriptions, pricing, or initialized component state make the corresponding component less informative, so tune weights against the data actually available. The complete example is config/fragments/algorithm/selection/hybrid.yaml.