Embedding Based Routing
This guide shows you how to route requests using semantic similarity with embedding models. Embedding routing matches user queries to predefined categories based on meaning rather than exact keywords, making it ideal for handling diverse phrasings and rapidly evolving categories.
Key Advantages
- Scalable: Handles unlimited categories without retraining models
- Fast: 10-50ms inference with efficient embedding models (Qwen3, Gemma)
- Flexible: Add/remove categories by updating keyword lists, no model retraining
- Semantic: Captures meaning beyond exact keyword matching
What Problem Does It Solve?
Keyword matching fails when users phrase questions differently. Embedding routing solves:
- Paraphrase handling: "How to install?" matches "installation guide" without exact words
- Intent detection: Routes based on semantic meaning, not surface patterns
- Fuzzy matching: Handles typos, abbreviations, informal language
- Dynamic categories: Add new categories without retraining classification models
- Multilingual support: Embeddings capture cross-lingual semantics
When to Use
- Customer support with diverse query phrasings
- Product inquiries where users ask the same thing many different ways
- Technical support needing semantic understanding of error descriptions
- Rapidly evolving categories where you need to add/update categories frequently
- Moderate latency tolerance (10-50ms acceptable for better semantic accuracy)
Configuration
Add embedding rules to your config.yaml:
# Define embedding signals
signals:
embeddings:
- name: "technical_support"
threshold: 0.75
candidates:
- "how to configure the system"
- "installation guide"
- "troubleshooting steps"
- "error message explanation"
aggregation_method: "max"
- name: "product_inquiry"
threshold: 0.70
candidates:
- "product features and specifications"
- "pricing information"
- "availability and stock"
aggregation_method: "avg"
- name: "account_management"
threshold: 0.72
candidates:
- "password reset"
- "account settings"
- "subscription management"
aggregation_method: "max"
# Define decisions using embedding signals
decisions:
- name: technical_support
description: "Route technical support queries"
priority: 100
rules:
operator: "OR"
conditions:
- type: "embedding"
name: "technical_support"
modelRefs:
- model: "openai/gpt-oss-120b"
use_reasoning: true
plugins:
- type: "system_prompt"
configuration:
system_prompt: "You are a technical support specialist with deep knowledge of system configuration and troubleshooting."
- name: product_inquiry
description: "Route product inquiry queries"
priority: 100
rules:
operator: "OR"
conditions:
- type: "embedding"
name: "product_inquiry"
modelRefs:
- model: "openai/gpt-oss-120b"
use_reasoning: false
plugins:
- type: "system_prompt"
configuration:
system_prompt: "You are a product specialist with comprehensive knowledge of features, pricing, and availability."
- type: "semantic-cache"
configuration:
enabled: true
similarity_threshold: 0.85
- name: account_management
description: "Route account management queries"
priority: 100
rules:
operator: "OR"
conditions:
- type: "embedding"
name: "account_management"
modelRefs:
- model: "openai/gpt-oss-120b"
use_reasoning: false
plugins:
- type: "system_prompt"
configuration:
system_prompt: "You are an account management specialist. Handle user account queries with care and security."
Embedding Models
- qwen3: High quality, 1024-dim, 32K context
- gemma: Balanced, 768-dim, 8K context, Matryoshka support (128/256/512/768)
- auto: Automatically selects based on quality/latency priorities
Aggregation Methods
- max: Uses highest similarity score
- avg: Uses average similarity across keywords
- any: Matches if any keyword exceeds threshold
Example Requests
# Technical support query
curl -X POST http://localhost:8801/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "MoM",
"messages": [{"role": "user", "content": "How do I troubleshoot connection errors?"}]
}'
# Product inquiry
curl -X POST http://localhost:8801/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "MoM",
"messages": [{"role": "user", "content": "What are the pricing options?"}]
}'