跳到主要内容
版本:最新版(未发布)

Profiles

A profile says how a model runs: exactly as published, or faster with slightly different numbers. You choose it per deployment.

ProfileIn plain wordsWhen to use it
exact (default)Gives the same answers the model's publishers measured, bit for bit for Decision 2.0. Task models run in full 32-bit precision on every device.Always, unless you measured that you need more speed.
shared_contextWhen you ask a decision model several questions about one request, it reads the request once and answers every question from that reading.Many decision questions per request on a GPU.
batchingRequests that arrive within a few milliseconds of each other are run together.High traffic on a GPU, where combining requests raises throughput.
max_speedUses faster kernels and 16-bit numbers where the hardware supports them, on top of the above.GPU deployments where latency matters more than the last decimal place.

Everything except exact is approximate: answers can differ slightly from the published ones, and a borderline answer can flip. On NVIDIA GPUs from Ampere on, max_speed runs Decision 2.0 with fused Triton kernels, 11% to 48% faster with no decision changed in its record; Vela 2.0 keeps the exact kernels there. Each approximate profile has a measured accuracy record in the runtime's records.

Set a profile​

For a router-managed model, set profile on the deployment:

global:
model_catalog:
deployments:
decision-lux:
provider: model_runtime
artifact: vllm-sr/Decision-2.0-Lux-9B
device: rocm:0
profile: shared_context

For a runtime you start yourself, pass --runtime-profile (--profile to vllm-srun serve):

vllm-sr serve vllm-sr/Decision-2.0-Lux-9B --engine --platform rocm --device-ids 0 --runtime-profile shared_context

A runtime started with an approximate profile still answers requests that ask for exact ("options": {"profile": "exact"}), so you can compare the two on your own traffic before you switch.

Performance without changing answers​

The runtime is fast in the default profile too. It never trades correctness for these:

  • compatible work for the same model and routing stage can share a bundled call; one bundle may still require several forwards;
  • task families can fuse compatible work, while decision-model shared-context execution is an explicit profile choice;
  • independent workers can execute in parallel when hardware capacity allows; placing multiple workers on one GPU does not add compute capacity;
  • cacheable repeated work can reuse recent results while they remain in the worker's bounded result cache (--result-cache-entries); replicas have independent worker caches;
  • long inputs are split into windows only when you ask for windowing.