Skip to main content
Version: Latest

Quickstart

In about ten minutes you install the model runtime, serve a model on your CPU, ask it a question, and then let the router use the same model for routing.

You need Linux or macOS, Python 3.10 or newer, and about 3 GB of free disk for the model download. No GPU is needed.

1. Install​

The runtime is a Python package that lives next to the vllm-sr CLI in the repository. Install both into a virtual environment, with the CPU build of PyTorch:

git clone https://github.com/vllm-project/semantic-router.git
cd semantic-router
python3 -m venv .venv
. .venv/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install ./src/vllm-sr ./src/model-runtime

Check that the runtime sees its built-in models:

vllm-sr-runtime models

2. Serve a model​

Start Decision 2.0 Kai, the smallest decision model. A decision model answers questions you write in plain language.

vllm-sr serve vllm-sr/Decision-2.0-Kai-0.6B --device cpu --port 8100

The first start downloads the model (about 1.5 GB) into your Hugging Face cache and checks every file against its pinned hash. The model is ready when its health check passes. In a second terminal:

curl -s localhost:8100/health

It answers {"status": "ready", ...} once the model has loaded and passed its self-check. Until then it answers HTTP 503 with the current stage.

3. Send a request​

Ask two questions about one request at once: which kind of work it is (a choice), and whether it needs multi-step reasoning (a yes/no question, called noul).

curl -s localhost:8100/v1/decisions -H 'content-type: application/json' -d '{
"state": "Write a Python function that merges two sorted lists.",
"questions": {
"kind": {"type": "choice", "instructions": "What kind of work is this?",
"criteria": {"code": "Writing or fixing code", "math": "Mathematics", "chat": "Anything else"}},
"reasoning": {"type": "noul", "instructions": "Does answering this need multi-step reasoning?"}
}
}'

The answer for kind names the chosen option and the probability of every option. The answer for reasoning is the probability that the answer is yes:

Response
{
"model": "Decision-2.0-Kai-0.6B",
"answers": {
"kind": {"type": "choice", "choice": "code", "probabilities": {"code": 0.504, "math": 0.133, "chat": 0.362}, "confidence": 0.106},
"reasoning": {"type": "noul", "noul": 0.519}
},
"usage": {"input_tokens": 168, "output_tokens": 0}
}

Add "options": {"return_meta": true} to the request to also get meta: the revision, profile and device that answered and how long it took. On 16 CPU cores this request takes about 0.2 seconds. GET /v1/models shows what is loaded, where it runs and whether it passed its self-check.

The same command serves classifiers. Stop the server with Ctrl-C and serve the Vela Domain classifier instead:

vllm-sr serve vllm-sr/Vela-1.0-Encoder-307M-Domain --device cpu --port 8100
curl -s localhost:8100/v1/classify -H 'content-type: application/json' \
-d '{"input": ["What is the derivative of x squared?"]}'

The result lists the most likely domain (label) and the probability of each of the 14 domains.

4. Use it from the router​

The router runs models for you. Name the model as a deployment with provider: model_runtime, then ask it questions in a decision signal. Save this as config.yaml, replacing vllm:8000 with an OpenAI-compatible backend that answers your users:

version: v0.3
listeners:
- name: http
address: 0.0.0.0
port: 8899
providers:
defaults:
model: answer-model
models:
- name: answer-model
backend_refs:
- name: answer
endpoint: vllm:8000
protocol: http
routing:
modelCards:
- name: answer-model
signals:
decision:
- name: needs_reasoning
deployment: decision-kai
question:
type: noul
instructions: Does answering this request need multi-step reasoning?
predicate:
gte: 0.7
decisions:
- name: think-first
priority: 100
rules:
operator: AND
on_unknown: no_match
conditions:
- type: decision
name: needs_reasoning
modelRefs:
- model: answer-model
use_reasoning: true
- name: default-route
priority: 1
rules:
operator: AND
conditions: []
modelRefs:
- model: answer-model
use_reasoning: false
global:
model_catalog:
deployments:
decision-kai:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Kai-0.6B
device: cpu

Validate the file and start the router:

vllm-sr config validate --config config.yaml
vllm-sr serve --config config.yaml

The router starts its own runtime for decision-kai. The first time, it downloads its own copy of the model into the models/ directory next to your config.yaml and keeps it there for later starts. Send a request through the router and look at which route it took:

curl -s -D - -o /dev/null localhost:8899/v1/chat/completions \
-H 'content-type: application/json' -H 'x-vsr-debug: true' \
-d '{"model": "auto", "messages": [{"role": "user", "content": "Plan a three-step proof that there are infinitely many primes."}]}' \
| grep -i '^x-vsr-'

x-vsr-selected-decision names the route, and x-vsr-matched-decision-model lists the decision signals that matched. While the model is still loading, the signal is unknown and on_unknown: no_match sends requests to default-route.

To reuse the server you started in step 2 instead of a second copy of the model, replace artifact and device with its address, for example endpoint: http://host.docker.internal:8100, and start that server with --host 0.0.0.0 so the router container can reach it.

Next steps​