Skip to main content
Version: v0.4

Train Vela router models

Adapt Vela to the languages, topics, and routing policies in your application. The family includes a shared encoder, task classifiers, an embedding model, and a reranker. Start with a published task model when you need inference; train a model when your evaluation shows a gap.

For deployment, follow Use Vela models. This section covers creating and evaluating your own task models.

Choose a training workflow​

Your goalWorkflowWhat you train
Improve semantic search or request similarityEmbeddingVectors for queries and documents
Improve the order of retrieved candidatesRerankingA relevance score for each query-document pair
Adapt domain, feedback, prompt-attack, fact-check or output-modality detectionClassifiersA request label
Detect personal informationPIIEntity spans
Apply a content-risk policySafety and HazardSafe/unsafe and risk categories

Vela's model catalog lists all eleven releases. Vela 1.0 accepts text; multimodal embeddings are a separate family. Provider model evaluation and learned model selection help choose the downstream LLM.

Choose your starting checkpoint​

Use Vela Encoder to train a new task. It is the common base for Vela's classifiers, embeddings, and reranker. Download an explicit Hub revision and keep its tokenizer and configuration together with the weights.

To improve an existing task, continue its complete Vela checkpoint, including its trained head. Use the same labels and preprocessing unless you deliberately train a new output contract. Each workflow explains fresh initialization and continuation.

Prepare examples from your application​

Start with real requests and the behavior you expect. Include languages and input lengths your application will receive, ordinary requests that should not trigger a signal, and examples close to the decision boundary.

Keep related conversations, documents, and translations in the same split. Use training data to update weights, development data to choose checkpoints and thresholds, and a separate test set for the final comparison. The task guides link to data formats and existing dataset builders.

Train and compare​

Run a small training job first to check the data and output labels. Then compare your trained model with the published task model using the same inputs and inference settings.

TaskMeasure
ClassifierPer-class precision, recall, F1, and false positives
PIIEntity-level precision, recall, and F1
EmbeddingRetrieval and semantic-similarity quality
RerankerRanking quality on fixed candidate lists
Safety and HazardMissed risks, safe-request false positives, and category coverage

Check language and length slices separately. For long-context applications, include short requests and actual long documents with relevant content near the beginning, middle, and end. Measure latency at the input lengths you plan to serve.

Use the trained model in the router​

Export a complete checkpoint with its tokenizer and labels. Choose a supported engine and input limit in Run models locally, then validate your configuration:

vllm-sr config validate --config config.yaml

Use route preview to check the signal, selected decision, and latency on representative requests before deploying the model to traffic.