Skip to main content
Version: Latest

Model runtime

Every model the router uses runs in the built-in model runtime: the classifiers behind signals such as domain, PII and jailbreak, the embedding models behind the semantic cache, memory and RAG, the reranker, the hallucination detector, and decision models that answer routing questions.

You usually do not have to do anything for this to work. When a feature needs a model, the router downloads it, checks every file, starts the runtime and sends it the request text. It starts serving once the models its routes need have loaded. If a runtime is slow or crashes later, requests keep flowing: the feature reports "unknown" and your routes fall back the way you configured.

Three ways to use it​

You want toDo thisRead
Use the router's built-in featuresNothing extra. The router starts and supervises the runtime for you.Run it with the router
Put the models on a GPU or share them between routersStart a runtime yourself and point the router at it with endpoint.Run it with the router
Call the models from your own codeRun vllm-sr serve <model> and send HTTP requests.Quickstart

What it can serve​

TaskBuilt-in modelsGuide
Pick a domain, detect a need for fact checking, read user feedback, detect the requested output modalityVela 1.0 Domain, FactCheck, Feedback, ModalityClassify requests
Find personal informationVela 1.0 PIIDetect PII
Stop prompt attacks and unsafe contentVela 1.0 Guard, Safety, Shield, HazardPrompt attacks and unsafe content
Check an answer against its sourcesVela 1.0 HaluHallucination checks
Semantic cache, memory, RAG, tool selection, embedding signalsVela 1.0 Embedding, Qwen3-Embedding-0.6BEmbeddings
Rerank retrieved documentsVela 1.0 RerankerRerank documents
Route on images and audioVela 1.0 Omni Nano and MiniImages and audio
Ask your own routing questions in plain languageDecision 2.0, Decision 1.0, Vela 2.0 (private preview)Decision models

Choose a model helps you pick a size and hardware.

What you can rely on​

  • Pinned and verified. Built-in models are pinned to an exact Hugging Face revision. Every file is checked against its recorded SHA-256 before it is loaded, and code shipped inside a model repository is never run.
  • Same answers as the released models. The default exact profile gives the answers the model publishers measured. Faster settings are opt-in and say that they may change results. See Profiles.
  • Requests never wait on a broken model. A model that is too slow, still restarting or crashed makes its feature "unknown" for that request. The router restarts a crashed runtime and keeps routing meanwhile.
  • Few calls per request. The model work a request's signals send to one runtime process goes as a single bundled call, and CPU models in separate processes answer in parallel, so adding signals does not add round trips.
  • Pluggable. New model families, engines and hardware back ends are ordinary Python packages. See Add your own model family.

Hardware​

CPU and AMD GPUs (MI300X, MI325X) are validated. NVIDIA GPUs work but are not yet validated; Intel GPUs (xpu) and Apple GPUs (mps) are available and not yet validated. Every router image runs models on the CPU. The AMD and NVIDIA images (vllm-sr serve --platform amd or --platform nvidia) also run them on the GPU.

Coming from an older release?​

The candle, ONNX Runtime and OpenVINO back ends are gone. Run vllm-sr config migrate to update your configuration; see Migrate from the native bindings.