Skip to main content
Version: Latest

Decision models

A decision model answers questions you write in plain language about a request, without training a classifier for each one. The router uses it in two places:

Kinds of questions​

TypeAsksThe answer
choiceWhich of these options fits?The chosen option and the probability of each option
noulIs this true?The probability of yes
scoreHow much, on an ordered scale?The expected level and the probability of each level
setWhich of these labels apply? (Vela 2.0)Every label above its threshold
spanWhere in the text is ...? (Vela 2.0)Labelled spans of the text

All questions of one request that go to the same model travel in one call and are answered together.

Ask a question​

global:
model_catalog:
deployments:
decision-kai:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Kai-0.6B
device: auto
routing:
signals:
decision:
- name: needs_reasoning
deployment: decision-kai
question:
type: noul
instructions: Does answering this request need multi-step reasoning?
predicate:
gte: 0.7
timeout_ms: 1000
- name: request_kind
deployment: decision-kai
question:
type: choice
instructions: What kind of request is this?
choices:
- key: code
description: Writing, reviewing or debugging code
- key: math
description: Mathematics or quantitative reasoning
- key: chat
description: Anything else
decisions:
- name: hard-code
priority: 200
rules:
operator: AND
on_unknown: no_match
conditions:
- type: decision
name: request_kind
label: code
- type: decision
name: needs_reasoning
modelRefs:
- model: large-coder

Write instructions the way you would ask a colleague, and describe each option in a few words. Short, concrete options give the most reliable answers.

Let it choose the model​

routing:
decisions:
- name: code-route
priority: 100
rules:
operator: AND
conditions:
- type: decision
name: request_kind
label: code
modelRefs:
- model: qwen3-8b
- model: qwen3-32b
algorithm:
type: decision
decision:
deployment: decision-kai
instructions: Which model should answer this request?
candidates:
qwen3-8b: Fast general model for routine code
qwen3-32b: Strong reasoning model for hard code
timeout_ms: 1000

The model's probability for each candidate becomes its selection score. If the model is not ready or answers late, the first model in modelRefs answers.

Which decision model​

Decision 2.0 is the default family; Kai-0.6B runs on a CPU and the larger sizes are more accurate on a GPU. Decision 1.0 models answer the same questions. Vela 2.0 also answers set and span questions and has ready-made questions for PII, hallucination and toxicity; it is a private preview and needs a Hugging Face token with access. See Choose a model.

Check it​

Ask the model the same question directly with /v1/decisions; see the Quickstart. Through the router, x-vsr-matched-decision-model lists the decision signals that matched and x-vsr-selected-model the model the selector chose.