Run it with the router
The router can run its models in two ways:
- Managed (the default). The router starts the runtime as a child process, restarts it if it crashes, and stops it when the router stops.
- Attached. You run the runtime yourself, for example on a GPU machine or
as a Kubernetes service, and the router connects to it with
endpoint.
Both use the same configuration, the same models and the same answers.
Built-in features need no configuration
When you turn on a built-in feature, such as a domain signal or the semantic
cache, the router already knows which Vela model it needs. It runs that model
in a managed runtime on the CPU. Set use_cpu: false on the feature's module
to let the runtime pick a GPU instead (device: auto).
Only models that an active feature uses are started. Declaring a deployment that nothing uses does not load it.
Describe a deployment
Write a deployment when you want to choose the model, the device or the
process yourself. A deployment is a named model under
global.model_catalog.deployments:
global:
model_catalog:
deployments:
vela-domain:
provider: model_runtime
artifact: vllm-sr/Vela-1.0-Encoder-307M-Domain
device: cpu
input:
max_tokens: 512
overflow: reject
decision-kai:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Kai-0.6B
device: auto
process: decisions
| Field | Meaning |
|---|---|
provider | Always model_runtime for models the runtime serves. |
artifact | A Hugging Face repository, or an absolute path to a local copy. |
revision | The exact 40-character commit to load. Built-in models are already pinned; other repositories need one. |
device | auto (default), cpu, cuda:N, rocm:N, xpu:N, mps, or an accelerator a plugin adds. |
profile | exact (default) or an opt-in faster profile. See Profiles. |
input | For task models: the longest input in tokens (max_tokens) and what to do with longer input (overflow: reject, truncate or window). |
process | Runs deployments with the same name in one runtime process. |
endpoint | Attach to a runtime you run instead of starting one. |
served_name | The model's name on an attached runtime that serves several models (default: the deployment name). |
Then connect a feature to the deployment with a binding:
global:
model_catalog:
bindings:
domain_classifier:
deployment: vela-domain
contract: label_distribution.v1
Each feature reads one kind of answer, its contract: label probabilities
(label_distribution.v1), independent label scores (label_scores.v1), text
spans (token_spans.v1), vectors (embedding.v1) or relevance scores
(relevance_scores.v1). The task guides list the binding for each feature.
At startup the router checks every binding against the model's own
description (its heads, labels, dimensions and input limit) and refuses a
mismatch before it serves traffic.
A binding under global.model_catalog.bindings applies everywhere. A recipe
can override it under its own routing.model_bindings.
Group models into processes
By default the models of one GPU share one runtime process, so all models on
rocm:0 answer a request in one call and use memory efficiently. They take
turns on the GPU, one device call at a time; put a model on a GPU of its own
when it must not wait for the others. CPU models
are spread over several processes, one per model up to one per two cores the
router may use, so a request's models run in parallel; each process runs an
equal share of the cores as threads. VLLM_SR_RUNTIME_CPU_PROCESSES caps the
number of CPU processes, and 1 keeps every CPU model in one process. Keep the
default where you can: in one process, an ONNX Runtime model such as Vela Omni
and a PyTorch model share the CPU's threads, and one of them answers more
slowly under load.
Deployments on device: auto (the default) are grouped by the device auto
picks on the router's host. On a host without a GPU that is the CPU, so each
of them gets a CPU process of its own, as with device: cpu. On a GPU host
they share the process of the first GPU, such as rocm:0, with the
deployments you put on that GPU; the runtime may still place a model on
another device when that GPU lacks the memory, and a model it places on the
CPU there may use every core the router has. vllm-sr-runtime devices shows
the device auto picks first. The router asks once, the first time it runs a
deployment on auto; if the runtime cannot answer, the auto deployments
share one process until the router restarts, and the router logs
auto_device_unresolved.
Give a deployment its own process name when it should not share a fault
domain or memory with the others, for example a large decision model:
global:
model_catalog:
deployments:
decision-lux:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Lux-9B
device: rocm:0
process: large-decisions
If that process crashes or runs out of memory, only its deployments become unavailable while the router restarts it; the other models keep answering.
Attach to a runtime you run
Start a runtime anywhere the router can reach, then point a deployment at it:
vllm-sr serve vllm-sr/Decision-2.0-Lux-9B vllm-sr/Vela-1.0-Encoder-307M-PII --device rocm:0 --host 0.0.0.0 --port 8100
global:
model_catalog:
deployments:
gpu-lux:
provider: model_runtime
endpoint: http://gpu-runtime.internal:8100
served_name: vllm-sr/Decision-2.0-Lux-9B
gpu-pii:
provider: model_runtime
endpoint: http://gpu-runtime.internal:8100
served_name: vllm-sr/Vela-1.0-Encoder-307M-PII
endpoint takes http://host:port, https://host:port or a Unix socket,
unix:///path/to/runtime.sock. The router does not start, restart or stop an
attached runtime; it checks its health and treats it as unavailable while it
is down.
A runtime started this way listens on 127.0.0.1 unless you pass --host.
Expose it only on a private network: it has no authentication of its own.
On a large host, give such a runtime --threads, or run it in a cpuset, when
it serves an ONNX Runtime model such as Vela Omni: each graph of that model
runs its own pool of up to --threads CPU threads, and without the option up
to every CPU the process may run on (an Omni bundle has four graphs). The
runtimes the router starts always get their share of the cores as --threads.
On Kubernetes
The router image already contains the CPU runtime, so managed deployments work
in any cluster. The ROCm router image,
ghcr.io/vllm-project/semantic-router/extproc-rocm, contains the runtime with
PyTorch for ROCm: give the router pod an AMD GPU and set device: rocm:0 on a
deployment, and the router runs that model on the GPU itself.
Models you keep on the CPU in the ROCm image run on its ROCm build of PyTorch.
For most Vela task models that build fails the runtime's load-time check that
a model answers the same alone and in a batch, so under exact they answer
queued requests one at a time: the same answers, with less throughput under
load. A router whose models all run on the CPU should use the CPU image.
The ROCm image's PyTorch has no CPU LAPACK, so decoders with gated delta rule
layers (the Qwen3.5-based models, such as Decision 2.0 Lux-9B or Vela 2.0 4B)
can't load on its CPU: the runtime refuses device: cpu for them at startup,
before their weights load, and names the cause. Serve them on the GPU, or
from the CPU image.
To share GPU models between several routers, run the runtime as its own
Deployment from the same image and attach every router to its Service. The
image's vllm-sr-runtime command starts it:
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-runtime
spec:
replicas: 1
selector:
matchLabels: {app: model-runtime}
template:
metadata:
labels: {app: model-runtime}
spec:
containers:
- name: runtime
image: ghcr.io/vllm-project/semantic-router/extproc-rocm:latest
command: ["vllm-sr-runtime"]
args: ["serve", "vllm-sr/Decision-2.0-Lux-9B", "--device", "rocm:0", "--host", "0.0.0.0", "--port", "8100"]
resources:
limits: {amd.com/gpu: 1}
ports:
- {name: http, containerPort: 8100}
readinessProbe:
httpGet: {path: /health, port: http}
livenessProbe:
httpGet: {path: /health/live, port: http}
---
apiVersion: v1
kind: Service
metadata:
name: model-runtime
spec:
selector: {app: model-runtime}
ports:
- {name: http, port: 8100, targetPort: http}
Set the deployment's endpoint to http://model-runtime.<namespace>.svc.cluster.local:8100.
Use /health as the readiness probe: it succeeds only after every model has
loaded and passed its self-check.
When a model is not ready
At startup the router waits until the task models its routes use have loaded: domain, PII, guard, safety, fact-check, feedback and hallucination models, and your own classifiers. If one of them fails to load, the router does not start, and its log names the model and the reason (see Troubleshooting). This holds for an attached runtime too, so start it before the router. Decision models do not hold up the start: until one answers, its signals are unknown.
Once the router is serving, requests never wait for a model that cannot answer:
| Situation | What a feature sees |
|---|---|
| The model is still downloading or loading | Unknown |
| The answer arrives after the feature's timeout | Unknown |
| The runtime is overloaded | Unknown |
| The runtime process crashed | Unknown until the router has restarted it (back-off from 1 s to 60 s) |
| Every model of a process failed to load | Unknown; the router restarts that process with the same back-off until the model loads |
The input is longer than the deployment allows and overflow: reject | An error for that input |
An unknown signal does not match. A decision's rules.on_unknown chooses what
an unknown signal means for that route (no_match by default, or
fail_request), and guard and PII modules have on_error: allow | block. The
decision selection algorithm falls back to the first model in modelRefs.
Check what is running
- The router exports
vsr_model_runtime_ready{deployment="..."}(1 when the deployment answers),vsr_model_runtime_restarts_total,vsr_model_runtime_requests_totalby outcome andvsr_model_runtime_unknown_answers_totalby reason on its metrics port. - The router's management API lists every deployment with its process,
state, restarts and the model it serves (labels, device, profile):
curl -s localhost:8080/api/v1/inventory/model-runtime. - An attached runtime answers
GET /v1/modelsandGET /healthdirectly.
The reference lists the environment variables that change the runtime command, the socket directory and the model cache of managed runtimes.