Skip to main content
Version: Latest

Built-in Model Runtime

Overview​

The model runtime serves decision models for the Router. It is a separate process with one HTTP contract: POST /v1/decisions (a strict superset of System One, also available as POST /v1/systemone), GET /v1/models, GET /health and GET /metrics. The Router talks to it over a Unix socket when it manages the runtime, or over HTTP when it attaches to an engine started elsewhere.

Phase 1 serves the six Decision 2.0 models out of the box:

ModelRevisionBackbone
vllm-sr/Decision-2.0-Kai-0.6B881bee41Qwen3 dense
vllm-sr/Decision-2.0-Eos-0.8Bad0aa724Qwen3.5 Gated DeltaNet hybrid
vllm-sr/Decision-2.0-Sol-2B4b75b521Qwen3.5 Gated DeltaNet hybrid
vllm-sr/Decision-2.0-Nox-4Bce1bdc9dQwen3.5 Gated DeltaNet hybrid
vllm-sr/Decision-2.0-Lux-9B214ffa43Qwen3.5 Gated DeltaNet hybrid
vllm-sr/Decision-2.0-Vega-27B9b067a95Qwen3.8-27B with a LoRA adapter

Every package is pinned by revision and verified file by file against its manifest before load. Model code is built into the runtime; code shipped inside a model package is never executed.

What Problem Does It Solve?​

Decision models are decoders with custom readouts, so they do not fit the Router's in-process encoder bindings. The runtime gives them one supervised, pluggable home: model families, execution engines and hardware accelerators are plugins, and faster numerics are opt-in profiles measured against the default.

When to Use​

Use it whenever a decision signal or the decision selection algorithm is configured. Run it on its own with vllm-sr serve <hf-model> to serve a decision model to other clients or to share one engine between Routers.

Configuration​

global:
model_catalog:
deployments:
decision-kai:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Kai-0.6B # Hub repository or an absolute package path
revision: 881bee413681d80ebeac86afcda8b4138dae516e
device: auto # auto, cpu, cuda[:N] or rocm[:N]
profile: exact # exact, shared_context, batching or max_speed
decision-shared:
provider: model_runtime
endpoint: http://decision-runtime:8100 # attach; unix:///path also works

Without endpoint, the Router starts vllm-sr-runtime serve for each deployment that a signal or selector uses, on a private Unix socket, restarts it with back-off when it exits, and stops it on shutdown. Until the runtime reports ready, decision signals are unknown and decision selectors fall back. VLLM_SR_RUNTIME_COMMAND overrides the command, VLLM_SR_RUNTIME_CACHE_DIR sets the model cache and VLLM_SR_RUNTIME_DIR the socket directory.

The vllm-sr image ships the runtime with CPU PyTorch, so managed CPU deployments work out of the box, and downloaded models persist under /app/models/model-runtime. For a GPU, start the runtime with a GPU build of PyTorch (vllm-sr serve <model> --device rocm:0, or the runtime image built from src/model-runtime/Dockerfile with a GPU wheel) and attach to it with endpoint.

Run an engine on its own:

pip install -e src/model-runtime
vllm-sr serve vllm-sr/Decision-2.0-Kai-0.6B --device cpu --port 8100
vllm-sr serve vllm-sr/Decision-2.0-Lux-9B --device rocm:0 --profile shared_context
ProfileNumericsPurpose
exactidentical to the released packagesthe default
shared_contextapproximaterun the shared state of a multi-question request once
batchingapproximateshare forwards between concurrent requests
max_speedapproximatefaster kernels that may change rounding

CPU and ROCm (MI300X, MI325X) are validated. CUDA is implemented and unit tested but not yet validated on NVIDIA hardware. The runtime exposes Prometheus metrics at /metrics, and the Router exports vsr_model_runtime_* metrics per deployment.