Model runtime
Every model the router uses runs in the built-in model runtime: the classifiers behind signals such as domain, PII and jailbreak, the embedding models behind the semantic cache, memory and RAG, the reranker, the hallucination detector, and decision models that answer routing questions.
You usually do not have to do anything for this to work. When a feature needs a model, the router downloads it, checks every file, starts the runtime and sends it the request text. It starts serving once the models its routes need have loaded. If a runtime is slow or crashes later, requests keep flowing: the feature reports "unknown" and your routes fall back the way you configured.
Three ways to use it
| You want to | Do this | Read |
|---|---|---|
| Use the router's built-in features | Nothing extra. The router starts and supervises the runtime for you. | Run it with the router |
| Put the models on a GPU or share them between routers | Start a runtime yourself and point the router at it with endpoint. | Run it with the router |
| Call the models from your own code | Run vllm-sr serve <model> and send HTTP requests. | Quickstart |
What it can serve
| Task | Built-in models | Guide |
|---|---|---|
| Pick a domain, detect a need for fact checking, read user feedback, detect the requested output modality | Vela 1.0 Domain, FactCheck, Feedback, Modality | Classify requests |
| Find personal information | Vela 1.0 PII | Detect PII |
| Stop prompt attacks and unsafe content | Vela 1.0 Guard, Safety, Shield, Hazard | Prompt attacks and unsafe content |
| Check an answer against its sources | Vela 1.0 Halu | Hallucination checks |
| Semantic cache, memory, RAG, tool selection, embedding signals | Vela 1.0 Embedding, Qwen3-Embedding-0.6B | Embeddings |
| Rerank retrieved documents | Vela 1.0 Reranker | Rerank documents |
| Route on images and audio | Vela 1.0 Omni Nano and Mini | Images and audio |
| Ask your own routing questions in plain language | Decision 2.0, Decision 1.0, Vela 2.0 (private preview) | Decision models |
Choose a model helps you pick a size and hardware.
What you can rely on
- Pinned and verified. Built-in models are pinned to an exact Hugging Face revision. Every file is checked against its recorded SHA-256 before it is loaded, and code shipped inside a model repository is never run.
- Same answers as the released models. The default
exactprofile gives the answers the model publishers measured. Faster settings are opt-in and say that they may change results. See Profiles. - Requests never wait on a broken model. A model that is too slow, still restarting or crashed makes its feature "unknown" for that request. The router restarts a crashed runtime and keeps routing meanwhile.
- Few calls per request. The model work a request's signals send to one runtime process goes as a single bundled call, and CPU models in separate processes answer in parallel, so adding signals does not add round trips.
- Pluggable. New model families, engines and hardware back ends are ordinary Python packages. See Add your own model family.
Hardware
CPU and AMD GPUs (MI300X, MI325X) are validated. NVIDIA GPUs work but are not
yet validated; Intel GPUs (xpu) and Apple GPUs (mps) are available and not
yet validated. Every router image runs models on the CPU. The AMD and NVIDIA
images (vllm-sr serve --platform amd or --platform nvidia) also run them on
the GPU.
Coming from an older release?
The candle, ONNX Runtime and OpenVINO back ends are gone. Run
vllm-sr config migrate to update your configuration; see
Migrate from the native bindings.