Decision models
A decision model answers questions you write in plain language about a request, without training a classifier for each one. The Router can use it for registered judgment tasks, custom questions, and backend selection:
- task bindings connect domain, complexity, preference, classification, safety, PII judgments, hallucination judgments, and other supported consumers to a deployment; capability checks use the model's actual question types;
- a
decisionsignal asks a question and routes on the answer; - the
decisionselection algorithm asks which of a route's models should answer.
Kinds of questions
| Type | Asks | The answer |
|---|---|---|
choice | Which of these options fits? | The chosen option and the probability of each option |
noul | Is this true? | The probability of yes |
score | How much, on an ordered scale? | The expected level and the probability of each level |
set | Which of these labels apply? (Vela 2.0) | Every label's probability, and the labels above its threshold |
span | Where in the text is ...? (Vela 2.0) | Labelled spans of the text, each with its probability |
These are native question types; their availability depends on the model card.
Router task adapters can also compose a set from one noul question per label.
That adapter does not add native set support to the public System One API.
Exact span results require a span-capable model.
Compatible questions for the same model and input within one routing stage can travel together, including PII questions when they use that deployment. Window scans, batch limits, and later routing stages can require additional calls or forwards. In particular, selecting the Chat backend after signal evaluation is a separate task; sharing a deployment avoids another model copy, not that task.
Ask a question
global:
model_catalog:
deployments:
decision-kai:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Kai-0.6B
device: auto
routing:
signals:
decision:
- name: needs_reasoning
deployment: decision-kai
question:
type: noul
instructions: Does answering this request need multi-step reasoning?
predicate:
gte: 0.7
timeout_ms: 1000
- name: request_kind
deployment: decision-kai
question:
type: choice
instructions: What kind of request is this?
choices:
- key: code
description: Writing, reviewing or debugging code
- key: math
description: Mathematics or quantitative reasoning
- key: chat
description: Anything else
decisions:
- name: hard-code
priority: 200
rules:
operator: AND
on_unknown: no_match
conditions:
- type: decision
name: request_kind
label: code
- type: decision
name: needs_reasoning
modelRefs:
- model: large-coder
Write instructions the way you would ask a colleague, and describe each option in a few words. Short, concrete options give the most reliable answers.
Let it choose the model
routing:
decisions:
- name: code-route
priority: 100
rules:
operator: AND
conditions:
- type: decision
name: request_kind
label: code
modelRefs:
- model: qwen3-8b
- model: qwen3-32b
algorithm:
type: decision
decision:
deployment: decision-kai
instructions: Which model should answer this request?
candidates:
qwen3-8b: Fast general model for routine code
qwen3-32b: Strong reasoning model for hard code
timeout_ms: 1000
The model's probability for each candidate becomes its selection score. If the
model is not ready or answers late, the first model in modelRefs answers.
Without a deployment, the Router's
decision model chooses: the
Vela 2.0 model that answered the request's signals also picks its model, with
no second copy loaded.
Route on labels and spans
Vela 2.0 also answers set and span questions. Both take labels, and a
condition names the label it routes on:
global:
model_catalog:
deployments:
vela2:
provider: model_runtime
artifact: vllm-sr/Vela-2.0-0.3B
device: cpu
routing:
signals:
decision:
- name: support_topics
deployment: vela2
question:
type: set
instructions: Which topics does the request mention?
labels:
- key: billing
description: payments, invoices or refunds
- key: shipping
description: deliveries, tracking or returns
- name: places
deployment: vela2
question:
type: span
instructions: Which spans name a city?
labels:
- key: city
description: a city name
decisions:
- name: billing-in-a-city
priority: 150
rules:
operator: AND
conditions:
- type: decision
name: support_topics
label: billing
- type: decision
name: places
label: city
modelRefs:
- model: support-model
A set label matches when the model selects it, and a span label when the
model finds a span of it; a predicate on the rule matches on the label's
probability instead. Every label's probability is a signal value
(decision:support_topics:billing). On models with a broad span head (0.8B,
4B, 9B), head: router | broad picks the head that answers a span question;
threshold replaces the model's own threshold. See the
signal reference.
The Router compiles each question against the selected model's capabilities.
A set can use a native set head or composed Noul judgments, including on
Decision 1.0 and 2.0. A span needs native span support; a model that lacks it
cannot provide exact positions. Managed models are checked during preparation.
Offline attached models remain unavailable until their current card can be
checked. The public System One API accepts the model's native question types,
so use Vela 2.0 for direct Set and Span requests.
Vela 2.0 also answers PII and unsupported claims with its router span head:
bind the pii and
hallucination signals
to the same deployment, and the PII question travels with the request's other
questions to it.
Which decision model
The Router defaults to Vela 2.0 0.3B. For custom questions, Decision 2.0
Kai-0.6B also runs on a CPU; its larger sizes target more demanding questions
on a GPU. Decision 1.0 models answer the same
questions. Vela 2.0 also answers set and span questions, has ready-made
questions for PII and unsupported claims, and its 0.3B answers the router's
built-in signals by default, batching
compatible questions within a routing stage. See
Choose a model.
Check it
Ask the model the same question directly with /v1/decisions; see the
Quickstart. Through the router,
x-vsr-matched-decision-model lists the decision signals that matched and
x-vsr-selected-model the model the selector chose.