Skip to main content
Version: Latest (unreleased)

Quickstart

In about ten minutes you install the vllm-sr CLI, serve a model on your CPU, ask it a question, and then let the router use the same model for routing.

You need Linux, macOS or WSL2 with Docker or Podman, Python 3.10 or newer, and a few gigabytes of free disk for the router image and the model download. No GPU is needed. Engine mode (vllm-sr serve ARTIFACT --engine) is newer than the 0.4.0 release; see the release channel note.

1. Install​

Install the development channel, which includes Engine mode, into the installer's isolated virtual environment. --no-launch leaves startup for the next step:

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev --no-launch

The model runtime, the vllm-srun Python package, ships only inside the router images. vllm-sr serve ARTIFACT --engine starts the instance frontend from the vllm-sr image, which the CLI pulls on first use, so nothing else is installed on your machine. On a GPU host, --platform rocm or --platform cuda selects the vllm-sr-rocm or vllm-sr-cuda image and passes the GPUs through; on macOS the runtime runs on the CPU, because containers there get no GPU. The images carry the release's own PyTorch build, so their answers are the ones the models were released with (Choose a model).

2. Serve a model​

Start Decision 2.0 Kai, the smallest decision model. A decision model answers questions you write in plain language.

vllm-sr serve vllm-sr/Decision-2.0-Kai-0.6B --engine --platform cpu

On an AMD GPU, the same model runs on the first GPU with:

vllm-sr serve vllm-sr/Decision-2.0-Kai-0.6B --engine --platform rocm --device-ids 0

The vllm-sr-rocm image is a 6.5 GB download on first use. rocm:N picks another GPU of the host in canonical YAML; --device-ids N selects a host GPU at startup.

Decision 3.0 models also answer questions about images. On an AMD Instinct MI325X, start the smallest one with vllm-sr serve vllm-sr/d3-lite --engine --platform rocm, and send images as described in Router and Engine modes.

The CLI starts the persistent frontend, Dashboard and a managed model worker. The first start downloads the model into the instance model cache. Later starts reuse it. Startup waits for readiness; use vllm-sr status to inspect the stack and vllm-sr stop to stop it.

The initial Engine configuration publishes the selected model on the default listener. An existing config keeps its explicit listeners[].systemone.models allowlist and API keys; restarting in another mode or selecting a model never broadens it. Every start without --engine uses Router mode; MODEL alone does not select Engine mode. Bare vllm-sr serve preserves the configured default judgment deployment, or initializes Vela 2.0 0.3B in a new configuration. In a second terminal, inspect the public native model list:

curl -s localhost:8899/v1/systemone/models

3. Send a request​

Ask two questions about one request at once: which kind of work it is (a choice), and whether it needs multi-step reasoning (a yes/no question, called noul).

curl -s localhost:8899/v1/systemone -H 'content-type: application/json' -d '{
"model": "vllm-sr/Decision-2.0-Kai-0.6B",
"state": "Write a Python function that merges two sorted lists.",
"questions": {
"kind": {"type": "choice", "instructions": "What kind of work is this?",
"criteria": {"code": "Writing or fixing code", "math": "Mathematics", "chat": "Anything else"}},
"reasoning": {"type": "noul", "instructions": "Does answering this need multi-step reasoning?"}
}
}'

The answer for kind names the chosen option and the probability of every option. The answer for reasoning is the probability that the answer is yes:

Response
{
"model": "vllm-sr/Decision-2.0-Kai-0.6B",
"answers": {
"kind": {"type": "choice", "choice": "code", "probabilities": {"code": 0.504, "math": 0.133, "chat": 0.362}, "confidence": 0.106},
"reasoning": {"type": "noul", "noul": 0.519}
},
"usage": {"input_tokens": 168, "output_tokens": 0}
}

Add "options": {"return_meta": true} to the request to also get meta: the revision, profile and device that answered and how long it took. On 16 CPU cores this request takes about 0.2 seconds in the recorded worker benchmark. POST /v1/decisions is an alias of /v1/systemone; both use the same explicit public model ID. Native discovery uses /v1/systemone/models, while /v1/models remains Chat discovery. The worker's classify, embeddings, rerank and bundle APIs are separate; see the task guides.

4. Use it from the router​

The router runs models for you. Name the model as a deployment with provider: model_runtime, then ask it questions in a decision signal. Save this as config.yaml, replacing host.docker.internal:8000 with an OpenAI-compatible backend that answers your users:

version: v0.3
listeners:
- name: http
address: 0.0.0.0
port: 8899
systemone:
models: [vllm-sr/Decision-2.0-Kai-0.6B]
providers:
defaults:
model: answer-model
models:
- name: answer-model
backend_refs:
- name: answer
endpoint: host.docker.internal:8000
protocol: http
routing:
modelCards:
- name: answer-model
signals:
decision:
- name: needs_reasoning
deployment: primary
question:
type: noul
instructions: Does answering this request need multi-step reasoning?
predicate:
gte: 0.7
decisions:
- name: think-first
priority: 100
rules:
operator: AND
on_unknown: no_match
conditions:
- type: decision
name: needs_reasoning
modelRefs:
- model: answer-model
use_reasoning: true
- name: default-route
priority: 1
rules:
operator: AND
conditions: []
modelRefs:
- model: answer-model
use_reasoning: false
global:
model_catalog:
deployments:
primary:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Kai-0.6B
device: cpu

Validate the file, then restart in Router mode using this configuration:

vllm-sr config validate --config config.yaml
vllm-sr serve --config config.yaml --replace-active-config

--replace-active-config applies the file you just wrote over the saved Engine configuration. Later restarts can omit it to retain Dashboard edits. The primary deployment answers both the routing question and direct System One requests; the listener explicitly publishes its native model name.

Send a request through the router and look at which route it took:

curl -s -D - -o /dev/null localhost:8899/v1/chat/completions \
-H 'content-type: application/json' -H 'x-vsr-debug: true' \
-d '{"model": "vllm-sr/auto", "messages": [{"role": "user", "content": "Plan a three-step proof that there are infinitely many primes."}]}' \
| grep -i '^x-vsr-'

x-vsr-selected-decision names the route, and x-vsr-matched-decision-model lists the decision signals that matched. If the runtime restarts later, the signal is unknown until the model is back and on_unknown: no_match sends requests to default-route meanwhile.

Router and Engine mode share one managed instance. To attach an independently operated worker instead, use an explicit endpoint and served_name; the deployment guide describes its separate worker API and lifecycle.

Next steps​