跳到主要内容
版本:最新版

Classify requests

Classifiers give every request a label: its subject, whether it needs fact checking, how the user reacts to the last answer, or whether they ask for an image. Routes match on those labels.

SignalModelLabels
domainVela 1.0 Domain14 domains: biology, business, chemistry, computer science, economics, engineering, health, history, law, math, other, philosophy, physics, psychology
fact_checkVela 1.0 FactCheckFACT_CHECK_NEEDED, NO_FACT_CHECK_NEEDED
user_feedbackVela 1.0 Feedbacksatisfied, need clarification, wrong answer, want different, no feedback
modalityVela 1.0 ModalityAR (text), DIFFUSION (image), BOTH
classifieryour own modelyour labels

Turn it on​

Add the signal and use it in a route. The router runs the right model for you on the CPU; nothing else is needed:

routing:
signals:
domains:
- name: math
description: Mathematics and quantitative reasoning.
mmlu_categories: [math]
- name: computer science
description: Programming and computer science.
mmlu_categories: [computer science]
decisions:
- name: math-route
priority: 100
rules:
operator: AND
conditions:
- type: domain
name: math
modelRefs:
- model: math-model

Choose where it runs​

To pick the device, the input limit or another model, describe a deployment and bind the feature to it. The binding names are domain_classifier, fact_check_classifier, feedback_detector and modality_detector; all of them read label probabilities (label_distribution.v1):

global:
model_catalog:
deployments:
vela-domain:
provider: model_runtime
artifact: vllm-sr/Vela-1.0-Encoder-307M-Domain
device: cpu
input:
max_tokens: 2048
overflow: truncate
bindings:
domain_classifier:
deployment: vela-domain
contract: label_distribution.v1

overflow: truncate classifies the first 2,048 tokens of a long request and reports that it did; reject (the default) leaves the signal unknown for input over the limit instead.

Check it​

Ask the model directly. Start it on its own, or use any runtime that serves it:

vllm-sr serve vllm-sr/Vela-1.0-Encoder-307M-Domain --device cpu --port 8100
curl -s localhost:8100/v1/classify -H 'content-type: application/json' \
-d '{"input": ["What is the derivative of x squared?", "Fix this segfault in my C code."]}'

Each result has the most likely label and the probabilities of all labels in the order of labels. Through the router, send a request with the x-vsr-debug: true header: x-vsr-matched-domains lists the domains that matched and x-vsr-selected-decision names the route.

Use your own classifier​

A Hugging Face ModernBERT or mmBERT sequence classifier works as a classifier signal with your own labels. Pin its revision, bind it to the rule and route on its labels:

routing:
signals:
classifiers:
- name: ticket_topic
type: local
labels: [billing, shipping, other]
decisions:
- name: billing-desk
priority: 100
rules:
operator: AND
conditions:
- type: classifier
name: ticket_topic
label: billing
predicate:
gte: 0.5
modelRefs:
- model: support-model
model_bindings:
classifier.ticket_topic:
deployment: ticket-topics
contract: label_distribution.v1
global:
model_catalog:
deployments:
ticket-topics:
provider: model_runtime
artifact: your-org/ticket-topic-classifier
revision: 0123456789abcdef0123456789abcdef01234567
device: cpu
input:
max_tokens: 512
overflow: reject

List labels in the model's own label order (its id2label). The router reads the model's labels from the runtime and refuses to start if they do not match the rule. A model with independent labels and a published operating point uses label_scores.v1 instead; see classifier signals.