Embeddings
An embedding turns text into a vector so the router can compare meaning. One embedding model serves every feature that needs one:
- embedding signals and knowledge bases: route on similarity to example texts;
- the semantic cache: answer a repeated question from the cache;
- memory and RAG: find stored facts and documents;
- tool selection: offer the tools that fit the request.
The default model is Vela 1.0 Embedding. It reads up to 32,768 tokens and can return smaller vectors from earlier layers when you need speed: layers 3, 6, 11 and 22, and 64, 128, 256, 512 or 768 dimensions.
Turn it on
Enabling a feature that needs embeddings is enough; the router runs Vela Embedding on the CPU. To set the vector size used for routing:
global:
model_catalog:
embeddings:
semantic:
embedding_config:
model_type: mmbert
target_layer: 22
target_dimension: 768
mmbert names the Vela Embedding slot. Use qwen3 for Qwen3-Embedding-0.6B
(1,024 dimensions) and multimodal for Vela Omni (see
Images and audio).
Choose where it runs
The binding is embedding and reads vectors (embedding.v1). Every service
that needs embeddings (caches, memory, vector stores, tools) uses the global
binding, so declare it once:
global:
model_catalog:
deployments:
vela-embedding:
provider: model_runtime
artifact: vllm-sr/Vela-1.0-Encoder-307M-Embedding
device: cpu
input:
max_tokens: 8192
overflow: truncate
bindings:
embedding:
deployment: vela-embedding
contract: embedding.v1
Different features may read different views of the same model, for example the semantic cache a 256-dimension vector of layer 6 and routing a 768-dimension vector of layer 22. They share one deployment; each view is its own vector space and is stored separately.
Use an external embedding service
Any OpenAI-compatible /embeddings endpoint, for example vllm serve with an
embedding model, can replace the local model for text. It replaces the
embedding binding above, so leave that binding out:
global:
model_catalog:
embeddings:
semantic:
embedding_config:
backend: openai_compatible
model_type: remote
target_dimension: 1024
endpoint:
base_url: https://embedding.example.com/v1
model: BAAI/bge-m3
api_key_env: EMBEDDING_API_KEY
dimensions: 1024
The service receives the text being embedded. Features that need images, audio, layer views or windows of long text need a local model.
Check it
vllm-sr serve vllm-sr/Vela-1.0-Encoder-307M-Embedding --device cpu --port 8100
curl -s localhost:8100/v1/embeddings -H 'content-type: application/json' \
-d '{"input": ["How do I reset my password?", "I forgot my login password."], "dimensions": 256}'
The response is the OpenAI embeddings format. Add
"options": {"return_meta": true} to also get meta.representation: the
model, layer and dimension that identify the vector space.
Change the model without mixing vectors
Vectors of different models, layers or sizes are never compared with each other. After you change any of them, re-index vector stores and re-add the memories you need, and expect the semantic cache to fill again. See Re-embed when the embedding model changes.