Skip to main content
Version: Latest (unreleased)

Choose a model, size and hardware

Start from the task. Every built-in model below is pinned to an exact Hugging Face revision, so the same name always loads the same files.

By task​

You want toDefault modelVela 1.0 specialistNotes
Route by subject (math, law, code, ...)Vela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-Domain14 domains
Spot requests that need fact checkingVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-FactCheckIt flags the need; it does not check facts
Read how a user reacts to the last answerVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-FeedbackSatisfied, needs clarification, wrong answer, wants something different, no feedback
Tell text requests from image requestsVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-ModalityReads the written request only
Find personal informationVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-PII17 entity types, with exact character spans
Stop prompt injection and jailbreaksVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-Guard
Flag unsafe contentVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-Safety or -ShieldShield is an alternative safety model
Check an answer against its sourcesVela 2.0 0.3Bvllm-sr/Vela-1.0-Encoder-307M-HaluMarks unsupported spans of the answer
Name the kind of riskExplicit bindingvllm-sr/Vela-1.0-Encoder-307M-Hazard12 independent categories; bind the hazard deployment with its published operating point
Embeddings for cache, memory, RAG and toolsvllm-sr/Vela-1.0-Encoder-307M-EmbeddingSmaller sizes and fewer layers trade quality for speed
Larger or instructed text embeddingsQwen/Qwen3-Embedding-0.6B0.6B, 1,024 dimensions
Rerank retrieved documentsvllm-sr/Vela-1.0-Encoder-307M-Reranker
Embed text, images and audio togethervllm-sr/Vela-1.0-Omni-Nano or -Mini164M / 1.36B; Mini is more accurate and accepts longer text
Ask your own questions in plain languageA decision model (next section)0.6B to 27B

With no model configured, the built-in signals the table gives Vela 2.0 0.3B share one deployment, which batches compatible questions within each routing stage (below). Hazard, embeddings, reranking and Omni run their own models. The Vela 1.0 specialists remain built in, and naming them restores them. Each is a 307M encoder that runs well on a CPU: on 16 cores the median Vela Domain request takes about 12 ms (measurements). Most read up to 32,768 tokens; the 0.3B reads 8,192. Each model card and GET /v1/models list the limits.

These input limits do not guarantee latency. Complete long-input embeddings and scans can exceed the signal deadline on a CPU with few cores. Measure your input lengths and concurrency, then choose dedicated CPU capacity or a GPU and tune worker threads for that workload.

Decision models​

Decision models answer questions you write yourself, such as "does this need step-by-step reasoning?" or "which of these models should answer?". Pick the smallest one that is accurate enough for your questions.

ModelSizeRuns well onGood for
vllm-sr/Decision-2.0-Kai-0.6B0.6BCPU (about 0.2 s for two questions on 16 cores) or any GPUFast, simple routing questions; the default choice to start with
vllm-sr/Decision-2.0-Eos-0.8B0.8BCPU or any GPUSlightly harder questions at similar cost
vllm-sr/Decision-2.0-Sol-2B2BGPU; CPU for low trafficQuestions that need more judgment
vllm-sr/Decision-2.0-Nox-4B4BGPUNuanced questions and many options
vllm-sr/Decision-2.0-Lux-9B9BGPU (24 GB or more)The most accurate at moderate cost
vllm-sr/Decision-2.0-Vega-27B27BOne GPU with 64 GB or moreThe most accurate overall

Decision 3.0 models also read images: a request may carry images that every question sees. They are built in as vllm-sr/d3 (27B), vllm-sr/d3-flash (9B), vllm-sr/d3-mini (4B), vllm-sr/d3-nano (2B) and vllm-sr/d3-lite (0.8B), and run on a GPU. Start one with, for example, vllm-sr serve vllm-sr/d3-lite --engine --platform rocm on an AMD Instinct MI325X. They answer choice, noul and score questions, and their answers match the packages' own runtime bit for bit (parity record).

Decision 1.0 models (vllm-sr/Decision-1.0-Kai-0.6B, -Lex-0.6B, -Route-0.6B, -Eos-0.8B, -Sol-2B, -Nox-4B, -Lux-9B) are also built in and answer the same kinds of questions. Vela 2.0 (vllm-sr/Vela-2.0-0.3B, -0.8B, -4B, -9B) adds questions that pick several labels (set) or mark spans of text (span), and the router routes on both. Its router span head also answers the pii and hallucination signals, which is how the default 0.3B deployment replaces the separate PII and Halu models. On a CPU, run the 0.3B. On a GPU, the larger sizes read inputs of up to 16,384 tokens (the 0.3B reads 8,192): the 0.8B costs the least of them, and the 4B and 9B are the most accurate.

vllm-srun models prints every built-in model with its pinned revision.

The built-in signals run on Vela 2.0 0.3B​

The domain, prompt guard, safety, fact check, user feedback, modality, PII and hallucination signals default to vllm-sr/Vela-2.0-0.3B (collection). All of them share one deployment, primary, which batches compatible questions from the same routing stage. One API call can carry several questions; window scans and batch limits may still require multiple model forward passes. Later routing stages can make additional calls.

  • Questions: each signal asks the question the model was trained on for it, with the labels of its Vela 1.0 model, so rules and policies read the answer as before. PII and hallucination use the model's span head, so their spans keep exact character offsets.
  • CPU profile: on a CPU the deployment runs max_speed, a packed copy of the model's weights. It gives the same answers to within about 0.00001 and is about 1.6 times faster than exact.
  • Input: ordinary routing judgments may truncate to the model's input limit. Prompt guard, safety, PII and hallucination require complete input; supported window scans can extend coverage up to the scan budget. Incomplete coverage produces an error or unknown result. See Long inputs for limits and routing policies. Vela 1.0 Guard and PII scan up to 32K in windows.
  • Thresholds: the module defaults are calibrated to the 0.3B's scores (below).

The maintainers chose this default although it misses two goals they had set for it (#4639): level or better accuracy on every signal, and level or better latency on a CPU. Measured through the Router on the router signal suite (A/B record):

  • Ahead: prompt guard (held-out AUC +0.026; on the E2E attack fixtures it blocks all six attacks, Vela 1.0 Guard five) and safety (+0.052 held-out, and ahead on every set). The signals share one model.
  • Level: PII and hallucination on held-out and fresh files.
  • Behind, most: modality (held-out AUC −0.180; the 0.3B misses most requests that ask for a new image) and user feedback (accuracy −0.038 held-out, −0.178 fresh).
  • Behind: domain (accuracy −0.037 held-out, −0.088 fresh) and fact check (held-out AUC −0.101).
  • CPU time: in this measurement, each request carries the questions, their options and the 17 PII labels (at least 560 tokens) through the 307M-parameter model, where each Vela 1.0 model reads only the request. On 12 CPU cores, for the five request signals of the latency record, the median request takes about 4.9 times as long:
Router on 12 CPU coresp50p95Requests per secondAt concurrency 16
Vela 1.0 specialists (restored)16 ms58 ms38.951.8
Vela 2.0 0.3B (the default)79 ms100 ms11.912.8

#4668 works on the CPU latency. On a GPU (use_cpu: false), the 0.3B answers the same questions in about 7 ms at the median on one AMD Instinct MI325X (measurements).

Thresholds​

Each default threshold keeps the Vela 1.0 specialist's operating point on the suite's dev split: its false-positive rate, or for a confidence floor its share of requests below the floor. The defaults are prompt guard 0.75, domain 0.28, PII 0.01, fact check 0.93 and user feedback 0.37.

  • PII: the 0.3B's span head applies its own per-label thresholds before it returns a span, so 0.01 accepts every span it returns.
  • Other models: a module that runs any other model and sets no threshold keeps its earlier default.
  • Your own rule thresholds (routing.signals.jailbreak[].threshold and the like) are yours, and they were likely chosen for Vela 1.0. The record maps each Vela 1.0 value to the 0.3B: prompt guard 0.3–0.9 → 0.74–0.77, PII → 0.01, safety 0.5 → 0.46, fact check 0.95 → 0.93, modality confidence_threshold 0.7 → 0.51.

Restore the Vela 1.0 specialists​

One block brings them back. Module thresholds you do not set return to the specialists' defaults with them:

global:
model_catalog:
system:
safety: models/Vela-1.0-Encoder-307M-Safety
prompt_guard: models/Vela-1.0-Encoder-307M-Guard
domain_classifier: models/Vela-1.0-Encoder-307M-Domain
pii_classifier: models/Vela-1.0-Encoder-307M-PII
fact_check_classifier: models/Vela-1.0-Encoder-307M-FactCheck
hallucination_detector: models/Vela-1.0-Encoder-307M-Halu
feedback_detector: models/Vela-1.0-Encoder-307M-Feedback

A modality classifier names models/Vela-1.0-Encoder-307M-Modality as its classifier.model_path.

To bring back one signal only, set its line alone. User feedback, for example:

global:
model_catalog:
system:
feedback_detector: models/Vela-1.0-Encoder-307M-Feedback
SignalLine under global.model_catalog
Domainsystem.domain_classifier: models/Vela-1.0-Encoder-307M-Domain
Prompt guardsystem.prompt_guard: models/Vela-1.0-Encoder-307M-Guard
Safetysystem.safety: models/Vela-1.0-Encoder-307M-Safety
Fact checksystem.fact_check_classifier: models/Vela-1.0-Encoder-307M-FactCheck
User feedbacksystem.feedback_detector: models/Vela-1.0-Encoder-307M-Feedback
PIIsystem.pii_classifier: models/Vela-1.0-Encoder-307M-PII
Hallucinationsystem.hallucination_detector: models/Vela-1.0-Encoder-307M-Halu
Modalitymodules.modality_detector.classifier.model_path: models/Vela-1.0-Encoder-307M-Modality

Rule thresholds that a configuration sets itself stay where they are. The built-in recipes' rules are calibrated to the 0.3B, so a signal moved back takes its Vela 1.0 rule thresholds with it. In mom-v1 those are prompt guard 0.5, safety 0.5 and PII 0.7; the record lists every recipe's.

Specialist overrides remain explicit task bindings; selecting a default decision model does not remove them.

Choose a size​

The default decision binding names a deployment that answers the Router's judgment tasks and decision questions without an override. Declare the resource once, then select its exact key. Without an override, the built-in default is Vela 2.0 0.3B. Passing a model artifact updates the active default deployment; --platform selects the runtime platform while explicit deployment placement remains authoritative. For example:

vllm-sr serve vllm-sr/Vela-2.0-4B --platform rocm
global:
model_catalog:
deployments:
primary:
provider: model_runtime
artifact: vllm-sr/Vela-2.0-4B
device: rocm:0
system:
decision_model:
deployment: primary

The CLI saves this deployment choice in the active configuration as a new version, which vllm-sr config versions lists and vllm-sr config rollback undoes; later starts keep it, and vllm-sr status shows it. The Helm chart's decisionModel value and the operator's spec.config.decision_model set the same binding. Deployment keys are exact and case-sensitive. Model identity, device and profile belong to the deployment, not the binding.

Measured through the Router on the router signal suite, against the Vela 1.0 specialists, and for the latency record's five request signals (record):

Decision modelHardwareHeld-out accuracy against Vela 1.0p50 on a GPUp50 on 12 CPU cores
Vela-2.0-0.3B (default)CPU or GPUAhead on prompt guard and safety, behind on domain, modality and feedback6.6 ms79 ms
Vela-2.0-0.8BCPU or GPUAhead on domain, prompt guard, safety, modality and hallucination; behind on PII40.1 msabout 3 s
Vela-2.0-4BGPU, about 17 GBAhead on every signal but fact check55.2 msNot measured
Vela-2.0-9BGPU, about 32 GBAhead on every signal76.5 msNot measured
Vela-1.0CPU or GPUThe specialists themselvesn/a16 ms
  • GPU: one AMD Instinct MI325X, sequential requests. At concurrency 16 a GPU serves about 154 (0.3B), 25 (0.8B), 18 (4B) and 13 (9B) requests per second.
  • Use a GPU for the 4B and 9B in this workload. Their CPU latency was not measured in this record. Placement follows the deployment and runtime requirements, not a CLI ban based on the model name. Explicit GPU devices must match the selected platform and be available on the host.
  • The 0.8B on a CPU is a decoder: a request takes seconds, as the table shows. Serve it on a GPU, or keep the 0.3B on a CPU.
  • Every size is behind Vela 1.0 on user feedback's fresh file (CrossWOZ) and on PII in distribution. Per-signal numbers with intervals are in the record.
  • Decision 1.0 and Decision 2.0 may be the default judgment deployment. Available tasks follow the model's native capabilities; an unsupported task is unavailable regardless of the model family.
  • Specialists remain explicit task overrides and can run alongside the default decision deployment.

Each size has its own module thresholds, which a module that sets none takes when you switch:

Decision modelPrompt guardDomainPIIFact checkUser feedback
Vela-2.0-0.3B0.750.280.010.930.37
Vela-2.0-0.8B0.710.380.070.9940.34
Vela-2.0-4B0.630.450.050.99840.33
Vela-2.0-9B0.420.460.140.9980.35
Vela-1.00.50.50.90.950.7

Rule thresholds a configuration sets, such as the built-in recipes', stay; the record maps each to every size (for example mom-v1's prompt_attack 0.75 is 0.71 on the 0.8B, 0.63 on the 4B and 0.42 on the 9B). A system.<module> line or a binding keeps that one signal on its own model.

Hardware​

HardwareStatusUse
CPUValidatedEvery router image runs models on CPU out of the box.
AMD Instinct MI300X, MI325XValidatedSet device: rocm:0. vllm-sr serve --platform rocm and the vllm-sr-rocm image ship PyTorch for ROCm.
NVIDIA GPUsWorks, not yet validatedSet device: cuda:0. vllm-sr serve --platform cuda ships PyTorch for CUDA.
Intel GPUsAvailable, not yet validateddevice: xpu:0, with the runtime installed next to an XPU build of PyTorch.
Apple siliconCPU only in this releaseOn macOS the docker target runs the CPU image, because Docker's Linux VM gets no GPU. Host GPU support is tracked in #4636.

On AMD GPUs the router images ship the stack the runtime is validated on: PyTorch 2.12 for ROCm 7.2, FLA 0.5.2, and causal-conv1d 1.7.0 built for ROCm. Its causal-conv1d also carries code for MI200 and MI350 GPUs, so the models that use it run there too, though only MI300X and MI325X are validated. The built-in models' GPU reference answers are checked on that stack, and every model compares itself with them when it loads. With another PyTorch, ROCm or kernel build, a model can fail that check or report unverified. If a model's reference answers had to be recorded again on this stack, its family's record says so and gives how often it agrees with the released answers.

device: auto (the default) picks the first validated GPU with enough free memory and otherwise the CPU. A GPU you name explicitly must exist, or the model fails to load with a clear reason instead of quietly running on the CPU.

How much memory​

Memory depends on the model family and profile, not just the device. FP32 weights use about 4 bytes per parameter; a lower-precision model may use about 2, with additional memory needed for inputs, activations, and concurrency. A 307M FP32 task model therefore needs roughly 1.3 GB for its weights alone. Check the model record and measure peak memory for your workload. Each replica has its own process and weight copy; place replicas on suitable devices (see Run it with the router).

Your own models​

Hugging Face ModernBERT and mmBERT classifiers, token classifiers and embedding models load the same way as the built-in Vela models: give a Hub repository with revision, or an absolute path to a local copy. Models of another architecture need a family plugin; see Add your own model family.