CLI Commands
Use vllm-sr to configure and run the router, manage recipes, send requests, and evaluate routing. See Installation to install the CLI and Local Docker deployment to start a router.
This reference is generated from the registered CLI commands. Command descriptions, arguments, options, and declared defaults match the source. An option without a declared default is omitted until supplied; commands may resolve it from configuration or the environment as described in their help. Repeatable options collect values. Run vllm-sr COMMAND --help for help from your installed version.
Command index
| Command | Description |
|---|---|
vllm-sr | vLLM Semantic Router CLI - Signal-driven routing across LLM providers, with a built-in model runtime. |
vllm-sr benchmark | Prepare, run, inspect, and compare sr-bench 1.0 evaluations. |
vllm-sr benchmark cancel | Cancel remaining work while retaining all existing evidence. |
vllm-sr benchmark candidate-plan | Reuse a terminal baseline's frozen protocol without repeating its requests. |
vllm-sr benchmark catalog | Show the nine benchmark adapters and evaluation profiles. |
vllm-sr benchmark compare | Compare matched cases against the best observed single model. |
vllm-sr benchmark comparison-options | List eligible baselines, or comparable live runs for BASELINE; no model calls. |
vllm-sr benchmark dataset | Prepare reproducible fixed benchmark case sets. |
vllm-sr benchmark dataset combine | Create a reusable multi-benchmark dataset from prepared manifests. |
vllm-sr benchmark dataset exclusions | Freeze finite named history without inspecting outcomes or running models. |
vllm-sr benchmark dataset options | List downloadable sources, profiles, access notes, and dependencies. |
vllm-sr benchmark dataset preparations | List shared download jobs or inspect one preparation by ID. |
vllm-sr benchmark dataset prepare | Download and freeze a dataset through the shared service by default. |
vllm-sr benchmark dataset show | Inspect a frozen dataset or list datasets in the shared store. |
vllm-sr benchmark experiment | Group durable baseline, routing checks, candidates and validation runs. |
vllm-sr benchmark experiment attach | Link existing evidence without changing or rerunning it. |
vllm-sr benchmark experiment create | Create an experiment without submitting model work. |
vllm-sr benchmark experiment delete | Delete a finished experiment's grouping and links; keep every run and result. |
vllm-sr benchmark experiment list | Read one page of experiments. |
vllm-sr benchmark experiment show | Read the experiment and one page of its linked runs. |
vllm-sr benchmark export | Export a dev response matrix for training; holdout export is rejected. |
vllm-sr benchmark plan | Validate and freeze all cases, targets, profiles, and limits without inference. |
vllm-sr benchmark preview | Inspect routing decisions without producing quality scores. |
vllm-sr benchmark reconcile-usage | Append an offline accounting correction from saved streams; no inference. |
vllm-sr benchmark recover | Create a separate attempt from a reviewed recovery plan; never auto-retry. |
vllm-sr benchmark recover-plan | Inspect eligible continuation cells without making model requests. |
vllm-sr benchmark regrade | Regrade saved MCQ/grid final outputs without mutating original evidence. |
vllm-sr benchmark replay | Estimate eligible static routes from saved answers without inference. |
vllm-sr benchmark replay-options | List eligible baselines, or compatible previews for BASELINE; no model calls. |
vllm-sr benchmark report | Show quality, four-bucket usage, cost, latency, time, and limitations. |
vllm-sr benchmark run | Execute one frozen live evaluation through the shared service. |
vllm-sr benchmark runs | |
vllm-sr benchmark serve | Own the durable journal and workers independently of a browser. |
vllm-sr benchmark setup | Inspect prerequisites or explicitly install optional benchmark harnesses. |
vllm-sr benchmark show | Read a run, bounded evidence page, or one complete saved call. |
vllm-sr benchmark target | Manage the operator-owned target registry used by the Dashboard. |
vllm-sr benchmark target list | |
vllm-sr benchmark target register | Replace the local registry from a JSON list of credential references. |
vllm-sr completion | Generate or install shell completion for vllm-sr. |
vllm-sr completion install | Install shell completions into your shell configuration. |
vllm-sr completion show | Print the completion script for a shell. |
vllm-sr config | Print generated configuration or run config subcommands. |
vllm-sr config apply | Plan, compare-and-swap, persist, and hot-reload a configuration. |
vllm-sr config envoy | Print the generated Envoy configuration. |
vllm-sr config get | Read the active canonical configuration from a Router. |
vllm-sr config init | Create a minimal canonical configuration template. |
vllm-sr config migrate | Migrate a legacy or mixed config file to canonical v0.3 YAML. |
vllm-sr config plan | Validate and plan an exact remote mutation without changing the Router. |
vllm-sr config rollback | Compare-and-swap the active configuration to a recorded version. |
vllm-sr config router | Print the canonical router configuration. |
vllm-sr config schema | Discover the canonical config contract progressively. |
vllm-sr config validate | Validate configuration file. |
vllm-sr config versions | List the configuration history, newest first. |
vllm-sr dashboard | Open the dashboard in your default web browser. |
vllm-sr instance | Inspect the serving state of an existing local instance. |
vllm-sr instance models | Print actual native model cards for readiness checks, without inference. |
vllm-sr instance status | Print desired/observed mode and durable operation state. |
vllm-sr logs | Show logs from vLLM Semantic Router service. |
vllm-sr optimize | Analyze routing evidence and produce candidate recipe changes. |
vllm-sr optimize recipe-learning | Analyze replay and outcomes to produce recipe-learning artifacts. |
vllm-sr recipe | Validate, plan, apply, inspect, or package routing recipes. |
vllm-sr recipe apply | Validate then compare-and-swap one recipe into the active config. |
vllm-sr recipe builtin | Discover, export, or bind installed Recipes without a running Router. |
vllm-sr recipe builtin export | Export an exact five-file BUNDLE to a new directory. |
vllm-sr recipe builtin init | Bind one installed recipe to explicit models and validate the new config. |
vllm-sr recipe builtin list | List installed bundles, recipe names, and required candidate pools as JSON. |
vllm-sr recipe delete | Compare-and-swap deletion of an unreferenced named recipe. |
vllm-sr recipe get | Read one managed recipe. |
vllm-sr recipe list | List recipes and the collection ETag. |
vllm-sr recipe pack | Create a deterministic ZIP from an exact five-file RECIPE_DIR. |
vllm-sr recipe plan | Validate a recipe and bind the plan to the current config ETag. |
vllm-sr recipe validate | Validate a recipe against the running Router without changing config. |
vllm-sr request | Send requests through the stack's listener. |
vllm-sr request chat | Send a one-shot chat completion through the Envoy-routed HTTP API. |
vllm-sr route | Preview routing decisions or probe the routed inference path. |
vllm-sr route preview | Preview signals and model selection without generating an answer. |
vllm-sr route probe | Probe a real route and assert complete assistant delivery for expected 2xx. |
vllm-sr serve | Start vLLM Semantic Router. |
vllm-sr status | Show status of vLLM Semantic Router services. |
vllm-sr stop | Stop vLLM Semantic Router. |
vllm-sr storage | Inspect Router storage or manage local storage credentials. |
vllm-sr storage rotate | Replace this stack's storage credentials with freshly generated ones. |
vllm-sr storage vector-stores | List vector stores known to the Router management API. |
vllm-sr
Usage: vllm-sr [OPTIONS] [COMMAND] [ARGS]...
vLLM Semantic Router CLI - Signal-driven routing across LLM providers, with a built-in model runtime.
| Parameter | Description |
|---|---|
--version | Show version and exit. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark
Usage: vllm-sr benchmark [OPTIONS] COMMAND [ARGS]...
Prepare, run, inspect, and compare sr-bench 1.0 evaluations.
| Parameter | Description |
|---|---|
--url TEXT | Shared sr-bench service URL; discovers the current local stack by default. Environment: SR_BENCH_URL. |
--store PATH | Durable service store; discovers the current local stack by default. Environment: SR_BENCH_STORE. |
--no-autostart | Require an already running service. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark cancel
Usage: vllm-sr benchmark cancel [OPTIONS] RUN_ID
Cancel remaining work while retaining all existing evidence.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark candidate-plan
Usage: vllm-sr benchmark candidate-plan [OPTIONS] BASELINE_RUN_ID
Reuse a terminal baseline's frozen protocol without repeating its requests.
Failed, cancelled and interrupted full-plan baselines may supply the same questions and settings. This does not qualify their measurements for Compare.
| Parameter | Description |
|---|---|
BASELINE_RUN_ID | Required argument. Type: text. |
--target TEXT | Registered MoM target; may be repeated. [required] May be repeated. |
--mode CHOICE | [default: live] Choices: live, preview. |
--name TEXT | — |
--experiment TEXT | — |
--hypothesis TEXT | Default: . |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark catalog
Usage: vllm-sr benchmark catalog [OPTIONS]
Show the nine benchmark adapters and evaluation profiles.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark compare
Usage: vllm-sr benchmark compare [OPTIONS] BASELINE_RUN_ID CANDIDATE_RUN_ID
Compare matched cases against the best observed single model.
| Parameter | Description |
|---|---|
BASELINE_RUN_ID | Required argument. Type: text. |
CANDIDATE_RUN_ID | Required argument. Type: text. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark comparison-options
Usage: vllm-sr benchmark comparison-options [OPTIONS] [BASELINE]
List eligible baselines, or comparable live runs for BASELINE; no model calls.
| Parameter | Description |
|---|---|
[BASELINE] | Optional argument. Type: text. |
--after TEXT | Opaque cursor from the previous eligible options page. |
--limit INTEGER RANGE | [default: 10; 1<=x<=25] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset
Usage: vllm-sr benchmark dataset [OPTIONS] COMMAND [ARGS]...
Prepare reproducible fixed benchmark case sets.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset combine
Usage: vllm-sr benchmark dataset combine [OPTIONS] MANIFESTS...
Create a reusable multi-benchmark dataset from prepared manifests.
| Parameter | Description |
|---|---|
MANIFESTS... | Required argument. Type: path. Accepts multiple values. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset exclusions
Usage: vllm-sr benchmark dataset exclusions [OPTIONS]
Freeze finite named history without inspecting outcomes or running models.
| Parameter | Description |
|---|---|
--dataset TEXT | Named prepared dataset whose entire membership is reserved. May be repeated. |
--run TEXT | Named frozen run whose entire planned membership is reserved. May be repeated. |
--output PATH | [required] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset options
Usage: vllm-sr benchmark dataset options [OPTIONS]
List downloadable sources, profiles, access notes, and dependencies.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset preparations
Usage: vllm-sr benchmark dataset preparations [OPTIONS] [PREPARATION_ID]
List shared download jobs or inspect one preparation by ID.
| Parameter | Description |
|---|---|
[PREPARATION_ID] | Optional argument. Type: text. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset prepare
Usage: vllm-sr benchmark dataset prepare [OPTIONS]
Download and freeze a dataset through the shared service by default.
| Parameter | Description |
|---|---|
--benchmark TEXT | Benchmark ID; repeat to prepare a shared collection. [required] May be repeated. |
--profile CHOICE | Choices: smoke, quick, standard. Default: quick. |
--source-path PATH | — |
--local | Prepare on this host; required for source files and history options. Default: false. |
--no-wait | Return the shared service preparation job immediately. Default: false. |
--revision TEXT | — |
--seed INTEGER | Default: 20260918. |
--source-partition TEXT | Frozen upstream partition for native task identity; never an evaluation split. |
--exclusion-snapshot PATH | — |
--evaluation-role CHOICE | Explicit family role; retest is never selected automatically. Choices: holdout, retest. |
--limit INTEGER | Custom case cap; cannot be represented as an upstream full benchmark. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark dataset show
Usage: vllm-sr benchmark dataset show [OPTIONS] [PATH]
Inspect a frozen dataset or list datasets in the shared store.
| Parameter | Description |
|---|---|
[PATH] | Optional argument. Type: path. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark experiment
Usage: vllm-sr benchmark experiment [OPTIONS] COMMAND [ARGS]...
Group durable baseline, routing checks, candidates and validation runs.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark experiment attach
Usage: vllm-sr benchmark experiment attach [OPTIONS] EXPERIMENT_ID
Link existing evidence without changing or rerunning it.
| Parameter | Description |
|---|---|
EXPERIMENT_ID | Required argument. Type: text. |
--run TEXT | [required] |
--role CHOICE | [required] Choices: baseline, candidate, estimate, initial, preview, recovery, smoke, validation. |
--hypothesis TEXT | Default: . |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark experiment create
Usage: vllm-sr benchmark experiment create [OPTIONS] NAME
Create an experiment without submitting model work.
| Parameter | Description |
|---|---|
NAME | Required argument. Type: text. |
--idempotency-key TEXT | — |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark experiment delete
Usage: vllm-sr benchmark experiment delete [OPTIONS] EXPERIMENT_ID
Delete a finished experiment's grouping and links; keep every run and result.
| Parameter | Description |
|---|---|
EXPERIMENT_ID | Required argument. Type: text. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark experiment list
Usage: vllm-sr benchmark experiment list [OPTIONS]
Read one page of experiments.
| Parameter | Description |
|---|---|
--after INTEGER RANGE | [x>=0] Default: 0. |
--limit INTEGER RANGE | [1<=x<=50] Default: 20. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark experiment show
Usage: vllm-sr benchmark experiment show [OPTIONS] EXPERIMENT_ID
Read the experiment and one page of its linked runs.
| Parameter | Description |
|---|---|
EXPERIMENT_ID | Required argument. Type: text. |
--after INTEGER RANGE | [x>=0] Default: 0. |
--limit INTEGER RANGE | [1<=x<=50] Default: 20. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark export
Usage: vllm-sr benchmark export [OPTIONS] RUN_ID
Export a dev response matrix for training; holdout export is rejected.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--output PATH | [required] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark plan
Usage: vllm-sr benchmark plan [OPTIONS]
Validate and freeze all cases, targets, profiles, and limits without inference.
| Parameter | Description |
|---|---|
--manifest PATH | [required] |
--output PATH | — |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark preview
Usage: vllm-sr benchmark preview [OPTIONS]
Inspect routing decisions without producing quality scores.
| Parameter | Description |
|---|---|
--idempotency-key TEXT | Bind repeated submissions to the same frozen plan, without reissuing calls. |
--detach | Return immediately with a durable run ID. Default: false. |
--manifest PATH | [required] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark reconcile-usage
Usage: vllm-sr benchmark reconcile-usage [OPTIONS] RUN_ID
Append an offline accounting correction from saved streams; no inference.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark recover
Usage: vllm-sr benchmark recover [OPTIONS] RUN_ID
Create a separate attempt from a reviewed recovery plan; never auto-retry.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--plan PATH | [required] |
--idempotency-key TEXT | [required] |
--acknowledge-new-attempt | Authorize new paid attempts for the exact reviewed failed cells. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark recover-plan
Usage: vllm-sr benchmark recover-plan [OPTIONS] RUN_ID
Inspect eligible continuation cells without making model requests.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--mode CHOICE | Choices: undispatched, failed. Default: undispatched. |
--output PATH | — |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark regrade
Usage: vllm-sr benchmark regrade [OPTIONS] RUN_ID
Regrade saved MCQ/grid final outputs without mutating original evidence.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--output PATH | [required] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark replay
Usage: vllm-sr benchmark replay [OPTIONS]
Estimate eligible static routes from saved answers without inference.
| Parameter | Description |
|---|---|
--baseline TEXT | Completed single-model answer matrix run ID. [required] |
--preview TEXT | Completed deterministic routing preview run ID. [required] |
--idempotency-key TEXT | — |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark replay-options
Usage: vllm-sr benchmark replay-options [OPTIONS] [BASELINE]
List eligible baselines, or compatible previews for BASELINE; no model calls.
| Parameter | Description |
|---|---|
[BASELINE] | Optional argument. Type: text. |
--after TEXT | Opaque cursor from the previous eligible options page. |
--limit INTEGER RANGE | [default: 10; 1<=x<=25] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark report
Usage: vllm-sr benchmark report [OPTIONS] RUN_ID
Show quality, four-bucket usage, cost, latency, time, and limitations.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--output PATH | — |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark run
Usage: vllm-sr benchmark run [OPTIONS]
Execute one frozen live evaluation through the shared service.
| Parameter | Description |
|---|---|
--idempotency-key TEXT | Bind repeated submissions to the same frozen plan, without reissuing calls. |
--detach | Return immediately with a durable run ID. Default: false. |
--manifest PATH | [required] |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark runs
Usage: vllm-sr benchmark runs [OPTIONS]
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark serve
Usage: vllm-sr benchmark serve [OPTIONS]
Own the durable journal and workers independently of a browser.
| Parameter | Description |
|---|---|
--host TEXT | Default: 127.0.0.1. |
--port INTEGER | Default: 8090. |
--store-identity TEXT | Canonical host store path SHA256 for an isolated runtime container. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark setup
Usage: vllm-sr benchmark setup [OPTIONS]
Inspect prerequisites or explicitly install optional benchmark harnesses.
| Parameter | Description |
|---|---|
--benchmark TEXT | Default: all. |
--install | Install pinned optional harnesses and task sources; makes no model requests. Default: false. |
--build-sandbox | Build a local offline grading image and record its content digest. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark show
Usage: vllm-sr benchmark show [OPTIONS] RUN_ID
Read a run, bounded evidence page, or one complete saved call.
| Parameter | Description |
|---|---|
RUN_ID | Required argument. Type: text. |
--results | Default: false. |
--calls | Default: false. |
--active | Read only in-progress calls; requires --calls. Default: false. |
--events | Default: false. |
--after INTEGER RANGE | Evidence cursor from the previous page. [x>=0] Default: 0. |
--limit INTEGER RANGE | Calls/results per page. [1<=x<=500] Default: 100. |
--call-id TEXT | Read one full saved call including prompt and final response. |
--help | Show this message and exit. Default: false. |
vllm-sr benchmark target
Usage: vllm-sr benchmark target [OPTIONS] COMMAND [ARGS]...
Manage the operator-owned target registry used by the Dashboard.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark target list
Usage: vllm-sr benchmark target list [OPTIONS]
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr benchmark target register
Usage: vllm-sr benchmark target register [OPTIONS]
Replace the local registry from a JSON list of credential references.
| Parameter | Description |
|---|---|
--file PATH | [required] |
--help | Show this message and exit. Default: false. |
vllm-sr completion
Usage: vllm-sr completion [OPTIONS] [COMMAND] [ARGS]...
Generate or install shell completion for vllm-sr.
Examples:
vllm-sr completion show bash # Print bash completion script
vllm-sr completion show zsh # Print zsh completion script
vllm-sr completion install # Auto-install for current shell
vllm-sr completion install bash # Install bash completions
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr completion install
Usage: vllm-sr completion install [OPTIONS] [bash|zsh|fish]
Install shell completions into your shell configuration.
Automatically appends the completion setup to your shell's rc file (~/.bashrc, ~/.zshrc) or writes to the fish completions directory. Safe to run multiple times — skips if already configured.
Examples:
vllm-sr completion install # Auto-detect shell
vllm-sr completion install bash # Install for bash
vllm-sr completion install zsh # Install for zsh
vllm-sr completion install fish # Install for fish
| Parameter | Description |
|---|---|
[SHELL] | Optional argument. Type: choice. Choices: bash, zsh, fish. |
--help | Show this message and exit. Default: false. |
vllm-sr completion show
Usage: vllm-sr completion show [OPTIONS] [bash|zsh|fish]
Print the completion script for a shell.
Outputs the completion script to stdout. If SHELL is omitted, the command attempts to detect the current shell automatically.
Examples:
vllm-sr completion show bash
vllm-sr completion show zsh
vllm-sr completion show fish
eval "$(vllm-sr completion show zsh)"
| Parameter | Description |
|---|---|
[SHELL] | Optional argument. Type: choice. Choices: bash, zsh, fish. |
--help | Show this message and exit. Default: false. |
vllm-sr config
Usage: vllm-sr config [OPTIONS] [COMMAND] [ARGS]...
Print generated configuration or run config subcommands.
Examples:
vllm-sr config init --output config.yaml
vllm-sr config validate --config config.yaml
vllm-sr config apply --config config.yaml
vllm-sr config router
vllm-sr config migrate --config old.yaml
vllm-sr config envoy # with --gateway extproc
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr config apply
Usage: vllm-sr config apply [OPTIONS]
Plan, compare-and-swap, persist, and hot-reload a configuration.
A change the running Router can't take without a restart is saved for the
next vllm-sr serve of a local stack, as the Dashboard saves one.
| Parameter | Description |
|---|---|
--config FILE | [default: config.yaml] |
--mode CHOICE | [default: replace] Choices: replace, merge. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 120] |
--token-env TEXT | Environment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr config envoy
Usage: vllm-sr config envoy [OPTIONS]
Print the generated Envoy configuration.
| Parameter | Description |
|---|---|
--config TEXT | Path to config file (default: config.yaml) Default: config.yaml. |
--help | Show this message and exit. Default: false. |
vllm-sr config get
Usage: vllm-sr config get [OPTIONS]
Read the active canonical configuration from a Router.
| Parameter | Description |
|---|---|
--format CHOICE | [default: yaml] Choices: json, yaml. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | Environment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr config init
Usage: vllm-sr config init [OPTIONS]
Create a minimal canonical configuration template.
| Parameter | Description |
|---|---|
--output TEXT | Path for the new canonical configuration template. [default: config.yaml] |
--force | Overwrite the output file if it already exists. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr config migrate
Usage: vllm-sr config migrate [OPTIONS]
Migrate a legacy or mixed config file to canonical v0.3 YAML.
| Parameter | Description |
|---|---|
--config TEXT | Path to source config file (default: config.yaml) Default: config.yaml. |
--output TEXT | Path for migrated canonical config (default: <config>.migrated.yaml) |
--force | Overwrite the output file if it already exists. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr config plan
Usage: vllm-sr config plan [OPTIONS]
Validate and plan an exact remote mutation without changing the Router.
| Parameter | Description |
|---|---|
--config FILE | [default: config.yaml] |
--mode CHOICE | [default: replace] Choices: replace, merge. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | Environment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr config rollback
Usage: vllm-sr config rollback [OPTIONS] VERSION
Compare-and-swap the active configuration to a recorded version.
VERSION is a configuration version number from vllm-sr config versions,
or the timestamp of a backup. The restored document activates as a new
version.
| Parameter | Description |
|---|---|
VERSION | Required argument. Type: text. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 120] |
--token-env TEXT | Environment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr config router
Usage: vllm-sr config router [OPTIONS]
Print the canonical router configuration.
| Parameter | Description |
|---|---|
--config TEXT | Path to config file (default: config.yaml) Default: config.yaml. |
--help | Show this message and exit. Default: false. |
vllm-sr config schema
Usage: vllm-sr config schema [OPTIONS]
Discover the canonical config contract progressively.
| Parameter | Description |
|---|---|
--endpoint TEXT | Read the contract from a running Router management origin instead of the schema bundled with this CLI. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--full | Print the complete JSON Schema instead of the compact index. Default: false. |
--section PATH | Print one config path and only its referenced definitions. |
--surface KIND:NAME | Print one signal, algorithm, plugin, or projection contract. |
--expanded | Include the selected section's self-contained JSON Schema. Default: false. |
--help | Show this message and exit. Default: false. |
vllm-sr config validate
Usage: vllm-sr config validate [OPTIONS]
Validate configuration file.
The CLI's own checks run first. The Router's validation then decides, as
it does when vllm-sr serve or vllm-sr config apply loads the file: the
Router in its local image (never pulled), or the running Router --endpoint
names.
Examples:
vllm-sr config validate
vllm-sr config validate --config my-config.yaml
vllm-sr config validate --endpoint http://localhost:8080
| Parameter | Description |
|---|---|
--config TEXT | Path to config file (default: config.yaml) Default: config.yaml. |
--endpoint TEXT | Validate with this running Router instead of the local Router image. |
--image TEXT | Router image whose own validation to run (default: the stack's, if present). |
--gateway CHOICE | The gateway mode the configuration is served in (default: standalone). Choices: standalone, extproc. |
--offline | Run only the CLI's own checks, without the Router's validation. Default: false. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr config versions
Usage: vllm-sr config versions [OPTIONS]
List the configuration history, newest first.
| Parameter | Description |
|---|---|
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | Environment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr dashboard
Usage: vllm-sr dashboard [OPTIONS]
Open the dashboard in your default web browser.
Examples:
vllm-sr dashboard # Docker dashboard
vllm-sr dashboard --target kubernetes # Show K8s address and port forward
vllm-sr dashboard --no-open
| Parameter | Description |
|---|---|
--no-open | Don't open browser, just show URL Default: false. |
--target TEXT | Deployment target: docker, kubernetes (default: docker) |
--namespace TEXT | Kubernetes namespace (kubernetes target only) |
--context TEXT | kubectl / Helm context (kubernetes target only) |
--container-runtime CHOICE | Container runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman. |
--help | Show this message and exit. Default: false. |
vllm-sr instance
Usage: vllm-sr instance [OPTIONS] COMMAND [ARGS]...
Inspect the serving state of an existing local instance.
| Parameter | Description |
|---|---|
--config FILE | Default: config.yaml. |
--help | Show this message and exit. Default: false. |
vllm-sr instance models
Usage: vllm-sr instance models [OPTIONS]
Print actual native model cards for readiness checks, without inference.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr instance status
Usage: vllm-sr instance status [OPTIONS]
Print desired/observed mode and durable operation state.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr logs
Usage: vllm-sr logs [OPTIONS] {envoy|router|dashboard}
Show logs from vLLM Semantic Router service.
Examples:
vllm-sr logs envoy
vllm-sr logs router
vllm-sr logs dashboard
vllm-sr logs envoy --follow
vllm-sr logs router -f
vllm-sr logs router --target kubernetes # Kubernetes logs
vllm-sr logs router --target kubernetes -f # Follow K8s logs
| Parameter | Description |
|---|---|
SERVICE | Required argument. Type: choice. Choices: envoy, router, dashboard. |
-f, --follow | Follow log output Default: false. |
--target TEXT | Deployment target: docker, kubernetes (default: docker) |
--namespace TEXT | Kubernetes namespace (kubernetes target only) |
--context TEXT | kubectl / Helm context (kubernetes target only) |
--container-runtime CHOICE | Container runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman. |
--help | Show this message and exit. Default: false. |
vllm-sr optimize
Usage: vllm-sr optimize [OPTIONS] COMMAND [ARGS]...
Analyze routing evidence and produce candidate recipe changes.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr optimize recipe-learning
Usage: vllm-sr optimize recipe-learning [OPTIONS]
Analyze replay and outcomes to produce recipe-learning artifacts.
| Parameter | Description |
|---|---|
--replay-file FILE | Router replay JSON file. Accepts a router_replay.list payload or a record array. |
--endpoint TEXT | Router management base URL (origin or /api/v1). Defaults to http://localhost:8080, with VLLM_SR_PORT_OFFSET added to the port, when --replay-file is omitted. Uses VSR_MGMT_TOKEN for bearer auth when set. |
--cases-file FILE | Optional eval cases JSON with replay_id/request_id plus expected_decision or expected_model. |
--recipe-file FILE | Optional current recipe YAML used to materialize complete candidate recipe variants. |
--limit INTEGER | Replay records to fetch from the endpoint. [default: 100] |
--output-dir DIRECTORY | Directory for metrics/findings/patch/seed-pack artifacts. |
--json | Print the full recipe-learning artifact. Default: false. |
--report-only | Compute metrics and findings without generating recipe patches. Default: false. |
--timeout INTEGER | HTTP request timeout in seconds. [default: 15] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe
Usage: vllm-sr recipe [OPTIONS] COMMAND [ARGS]...
Validate, plan, apply, inspect, or package routing recipes.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr recipe apply
Usage: vllm-sr recipe apply [OPTIONS] RECIPE_FILE
Validate then compare-and-swap one recipe into the active config.
| Parameter | Description |
|---|---|
RECIPE_FILE | Required argument. Type: file. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe builtin
Usage: vllm-sr recipe builtin [OPTIONS] COMMAND [ARGS]...
Discover, export, or bind installed Recipes without a running Router.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr recipe builtin export
Usage: vllm-sr recipe builtin export [OPTIONS] BUNDLE
Export an exact five-file BUNDLE to a new directory.
| Parameter | Description |
|---|---|
BUNDLE | Required argument. Type: text. |
--output-dir DIRECTORY | [required] |
--catalog-version TEXT | [default: latest] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe builtin init
Usage: vllm-sr recipe builtin init [OPTIONS] NAME
Bind one installed recipe to explicit models and validate the new config.
No model, endpoint, capability, or candidate pool is invented. Every recipe decision must meet its declared minimum_candidates before this succeeds.
| Parameter | Description |
|---|---|
NAME | Required argument. Type: text. |
--bundle TEXT | Bundle from recipe builtin list. [required] |
--config FILE | Existing canonical provider config to preserve. [required] |
--bindings FILE | YAML mapping each decision that calls a backend to modelRefs; omit immediate responses. [required] |
--model-name TEXT | Explicit public entrypoint name. [required] |
--exclude-decision TEXT | Explicitly omit a named lane for an unavailable capability; creates a recipe derivative. Repeatable. May be repeated. |
--output FILE | New config path; existing files are never replaced. Keep beside --config when it references local KB assets. [required] |
--catalog-version TEXT | [default: latest] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe builtin list
Usage: vllm-sr recipe builtin list [OPTIONS]
List installed bundles, recipe names, and required candidate pools as JSON.
| Parameter | Description |
|---|---|
--catalog-version TEXT | [default: latest] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe delete
Usage: vllm-sr recipe delete [OPTIONS] NAME
Compare-and-swap deletion of an unreferenced named recipe.
| Parameter | Description |
|---|---|
NAME | Required argument. Type: text. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe get
Usage: vllm-sr recipe get [OPTIONS] NAME
Read one managed recipe.
| Parameter | Description |
|---|---|
NAME | Required argument. Type: text. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe list
Usage: vllm-sr recipe list [OPTIONS]
List recipes and the collection ETag.
| Parameter | Description |
|---|---|
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe pack
Usage: vllm-sr recipe pack [OPTIONS] RECIPE_DIR
Create a deterministic ZIP from an exact five-file RECIPE_DIR.
| Parameter | Description |
|---|---|
RECIPE_DIR | Required argument. Type: directory. |
--output PATH | Archive path or output directory (default: the Recipe parent directory). |
--help | Show this message and exit. Default: false. |
vllm-sr recipe plan
Usage: vllm-sr recipe plan [OPTIONS] RECIPE_FILE
Validate a recipe and bind the plan to the current config ETag.
| Parameter | Description |
|---|---|
RECIPE_FILE | Required argument. Type: file. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr recipe validate
Usage: vllm-sr recipe validate [OPTIONS] RECIPE_FILE
Validate a recipe against the running Router without changing config.
| Parameter | Description |
|---|---|
RECIPE_FILE | Required argument. Type: file. |
--endpoint TEXT | Router management base URL; defaults to the local Router API port. |
--timeout FLOAT | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |
vllm-sr request
Usage: vllm-sr request [OPTIONS] COMMAND [ARGS]...
Send requests through the stack's listener.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr request chat
Usage: vllm-sr request chat [OPTIONS] [MESSAGE]...
Send a one-shot chat completion through the Envoy-routed HTTP API.
Uses the first listener port in config.yaml plus the stack port offset. The default model is the namespaced automatic-routing alias vllm-sr/auto.
Examples:
vllm-sr request chat "hello"
vllm-sr request chat --model vllm-sr/auto --prompt "Explain mixture of models"
vllm-sr request chat --json "hello"
| Parameter | Description |
|---|---|
[MESSAGE]... | Optional argument. Type: text. Accepts multiple values. |
--prompt TEXT | User message (alternative to a positional prompt). |
--model TEXT | Model name sent to the router. [default: vllm-sr/auto] |
--system TEXT | Optional system message prepended to the conversation. |
--config TEXT | Config file used to resolve listener host port (Docker default only). [default: config.yaml] |
--base-url TEXT | Explicit routed listener origin or OpenAI /v1 base URL for a remote or port-forwarded stack. |
--json | Print the raw JSON response instead of assistant text. Default: false. |
--timeout FLOAT | HTTP timeout in seconds for the completion request. [default: 120.0] |
--temperature FLOAT | Optional sampling temperature passed through to the API. |
--target TEXT | Deployment target (docker or k8s). |
--help | Show this message and exit. Default: false. |
vllm-sr route
Usage: vllm-sr route [OPTIONS] COMMAND [ARGS]...
Preview routing decisions or probe the routed inference path.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr route preview
Usage: vllm-sr route preview [OPTIONS]
Preview signals and model selection without generating an answer.
| Parameter | Description |
|---|---|
--prompt TEXT | Plain text prompt to evaluate. |
--messages TEXT | OpenAI-style messages JSON array string. |
--model TEXT | Routing model or entrypoint whose recipe should be evaluated. |
--request-file FILE | JSON Router Preview request; supported Chat messages and prompt fields only. |
--session-id TEXT | Read-only Learning session identity. |
--conversation-id TEXT | Read-only Learning conversation identity. |
--sampling-seed INTEGER RANGE | Preview-only exploration seed; does not fix a later live random draw. [-9223372036854775808<=x<=9223372036854775807] |
--endpoint TEXT | Router management origin or /api/v1 root; defaults to the local management port. |
--token-env TEXT | Environment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN] |
--trace / --no-trace | Include per-decision routing trace trees. Default: false. |
--json | Print the full JSON response payload. Default: false. |
--timeout FLOAT | HTTP request timeout in seconds. [default: 15.0] |
--help | Show this message and exit. Default: false. |
vllm-sr route probe
Usage: vllm-sr route probe [OPTIONS]
Probe a real route and assert complete assistant delivery for expected 2xx.
| Parameter | Description |
|---|---|
--prompt TEXT | Plain-text user prompt. |
--messages TEXT | OpenAI-style messages JSON array string. |
--model TEXT | [default: vllm-sr/auto] |
--config TEXT | [default: config.yaml] |
--base-url TEXT | Explicit Envoy listener origin or OpenAI /v1 base URL; otherwise derive it from --config. |
--api-key-env TEXT | Environment variable containing the bearer token; omitted when unset. [default: OPENAI_API_KEY] |
--temperature FLOAT | — |
--max-completion-tokens INTEGER RANGE | Completion token budget, including reasoning; omitted uses the backend default. [x>=1] |
--timeout FLOAT | [default: 120.0] |
--target TEXT | Deployment target used for URL resolution. |
--debug / --no-debug | [default: debug] |
--expect-status INTEGER | [default: OK] |
--expect-recipe TEXT | — |
--expect-decision TEXT | — |
--expect-algorithm TEXT | — |
--expect-selected-model TEXT | Assert the Router's x-vsr-selected-model receipt header. |
--expect-response-model TEXT | Assert the upstream OpenAI response body's top-level model field. |
--help | Show this message and exit. Default: false. |
vllm-sr serve
Usage: vllm-sr serve [OPTIONS] [MODEL]
Start vLLM Semantic Router.
Serve uses --config or config.yaml and preserves the Dashboard-first setup flow. Connect physical models and publish Mixture-of-Model entrypoints in the Dashboard, then keep the same stack running with this single command.
Virtual models are routing policies. Semantic Router starts the Router, which serves the OpenAI-compatible API itself, the Dashboard and supporting services; it does not download or launch the physical LLM engines referenced by provider backends. Connect user-owned single or multiple model endpoints through one canonical config or the Dashboard.
Ports are configured in the selected config under the listeners section.
Local startup waits up to 1800 seconds for Router readiness, or Dashboard readiness during first-run setup. During setup the command keeps waiting, and once you activate a config in the Dashboard it starts the Router from it. Use --startup-timeout SECONDS for a different positive budget when model loading or GPU compilation needs more time. The wait begins after containers start. It does not change inference deadlines. Timeout exits the CLI with an error and leaves containers available for inspection.
GATEWAY MODES:
standalone - The Router serves the OpenAI-compatible API itself (default)
extproc - An Envoy-based gateway in front of the Router: the Envoy container
on docker, as before standalone became the default; your gateway
on kubernetes
DEPLOYMENT TARGETS:
docker - Local Docker deployment (default)
kubernetes - Kubernetes deployment via Helm (k8s is the old name, for this
release only)
MODEL SELECTION ALGORITHMS:
static - Use first configured model (default, no learning)
router_dc - Query-model matching via embedding similarity
automix - Cost-quality optimization using POMDP
hybrid - Combine multiple methods with configurable weights
workflows - Router Flow static/dynamic micro-agent orchestration
latency_aware - TPOT/TTFT percentile-aware selection
knn - KNN selector using shared ML model-selection settings
kmeans - KMeans selector using shared ML model-selection settings
svm - SVM selector using shared ML model-selection settings
mlp - MLP selector using shared ML model-selection settings
multi_factor - Quality, latency, cost, and load scoring
Cross-request learning lives under global.router.learning.adaptation and global.router.learning.protection instead of --algorithm.
Examples:
# Dashboard-first setup or an existing ./config.yaml
vllm-sr serve
# Envoy in front of the Router, as before standalone became the default
vllm-sr serve --gateway extproc
# User-owned single or multi-model topology
vllm-sr serve --config my-models.yaml
# Explicitly replace Dashboard-edited runtime state from reviewed source YAML
vllm-sr serve --config my-models.yaml --replace-active-config
# Deploy a user-owned config to Kubernetes
vllm-sr serve --target kubernetes --config my-models.yaml --namespace my-ns
# Runtime policy and image overrides
vllm-sr serve --algorithm latency_aware
vllm-sr serve --image-pull-policy always
vllm-sr serve --readonly
vllm-sr serve --minimal
vllm-sr serve --log-level debug
# AMD ROCm image, device passthrough, and router internal GPU defaults
vllm-sr serve --platform rocm
vllm-sr serve --platform rocm --startup-timeout 7200
VLLM_SR_AMD_ROUTER_VISIBLE_DEVICES=7 vllm-sr serve --platform rocm
INSTANCE MODES:
vllm-sr serve
vllm-sr serve vllm-sr/Decision-2.0-Kai-0.6B --engine
vllm-sr serve vllm-sr/Vela-2.0-4B --platform rocm -dp 2 --device-ids 0
Without --engine the instance starts in Router mode, including on restart. --engine (-e) disables recipe routing; the frontend, Dashboard and native System One APIs remain available. Saved routing configuration is retained.
MODEL overrides the configured default judgment deployment's artifact. Omitting MODEL preserves that deployment (a new configuration uses Vela 2.0 0.3B). Only explicit model/placement options override saved settings. Backend LLMs, named deployments, listeners and API grants belong in --config.
| Parameter | Description |
|---|---|
[MODEL] | Optional argument. Type: text. |
-e, --engine | Start without recipe routing; otherwise start Router mode. Default: false. |
--config TEXT | Path to the Router configuration. [default: config.yaml] |
--replace-active-config | Replace this local Docker stack's active runtime config from --config, discarding Dashboard edits. Default: false. |
--image TEXT | Docker image to use (default: ghcr.io/vllm-project/semantic-router/vllm-sr:latest) |
--router-image TEXT | Docker image for the router container (Docker target only; defaults to --image or VLLM_SR_IMAGE) |
--envoy-image TEXT | Docker image for the Envoy container (docker target with --gateway extproc; defaults to --image or VLLM_SR_IMAGE) |
--dashboard-image TEXT | Docker image for the dashboard container (Docker target only; defaults to --image or VLLM_SR_IMAGE) |
--image-pull-policy CHOICE | Image pull policy: always, ifnotpresent, never (default: always) Choices: always, ifnotpresent, never. Default: always. |
--startup-timeout SECONDS | Local Docker startup readiness budget in seconds, including model loading and compilation (default: 1800). [x>=1] |
--readonly | Run dashboard in read-only mode (disable config editing, allow playground only) Default: false. |
--minimal | Start in minimal mode: no Dashboard or observability stack (Jaeger, Prometheus, Grafana) Default: false. |
--log-level CHOICE | Log level of the Router, or of the runtime in engine mode (debug, info, warn, error, dpanic, panic, fatal) Choices: debug, info, warn, warning, error, dpanic, panic, fatal. |
--platform CHOICE | Execution backend: auto (default) discovers the deployment target; cpu, cuda or rocm select it explicitly. Choices: auto, cpu, cuda, rocm. |
--algorithm CHOICE | Request-time base algorithm override for payload-safe algorithms: static, router_dc, automix, hybrid, workflows, latency_aware, knn, kmeans, svm, mlp, multi_factor. Algorithms that require an authored payload remain available in config.yaml. Cross-request learning uses global.router.learning.adaptation/protection. Choices: static, router_dc, automix, hybrid, workflows, latency_aware, knn, kmeans, svm, mlp, multi_factor. |
--target TEXT | Deployment target: docker, kubernetes (default: docker) |
--gateway CHOICE | Where client traffic enters: standalone (default; the Router serves the OpenAI-compatible API on the config's listeners, with no Envoy) or extproc (an Envoy-based gateway in front of the Router: the Envoy container on the docker target, your gateway on kubernetes). Choices: standalone, extproc. |
--namespace TEXT | Kubernetes namespace (kubernetes target only) |
--context TEXT | kubectl / Helm context (kubernetes target only) |
--profile TEXT | Deployment profile: dev, prod (kubernetes target only). Selects values-<profile>.yaml defaults. |
--chart-dir TEXT | Path to Helm chart directory (kubernetes target only; default: ./deploy/helm/semantic-router, else the published chart for this version) |
--container-runtime CHOICE | Container runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman. |
--recipe-env NAME | Explicitly bind one host environment variable for the active Recipe. Repeat for multiple names; NAME=value is rejected. May be repeated. |
--revision TEXT | Optional model branch, tag or commit; resolved once to an immutable startup revision. |
-dp, --data-parallel-size INTEGER RANGE | Number of model replicas. Preserve configured placement; new GPU deployments use distinct available GPUs. [1<=x<=64] |
--device-ids IDS | Docker host GPU indices, e.g. 0 or 0,1. One index shares a GPU across replicas; otherwise use one per replica. Existing visibility masks are respected, not changed. |
--runtime-profile PROFILE | Model runtime numerics profile (default exact; vllm-srun plugins lists the installed ones). |
--help | Show this message and exit. Default: false. |
vllm-sr status
Usage: vllm-sr status [OPTIONS] [envoy|router|dashboard|all]
Show status of vLLM Semantic Router services.
Examples:
vllm-sr status # Show all services (Docker)
vllm-sr status all # Show all services
vllm-sr status router # Show router status
vllm-sr status dashboard # Show dashboard status
vllm-sr status --target kubernetes # Show Kubernetes status
| Parameter | Description |
|---|---|
[SERVICE] | Optional argument. Type: choice. Choices: envoy, router, dashboard, all. Default: all. |
--target TEXT | Deployment target: docker, kubernetes (default: docker) |
--namespace TEXT | Kubernetes namespace (kubernetes target only) |
--context TEXT | kubectl / Helm context (kubernetes target only) |
--container-runtime CHOICE | Container runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman. |
--help | Show this message and exit. Default: false. |
vllm-sr stop
Usage: vllm-sr stop [OPTIONS]
Stop vLLM Semantic Router.
Examples:
vllm-sr stop # Stop Docker stack
vllm-sr stop --target kubernetes # Uninstall Helm release
| Parameter | Description |
|---|---|
--target TEXT | Deployment target: docker, kubernetes (default: docker) |
--namespace TEXT | Kubernetes namespace (kubernetes target only) |
--context TEXT | kubectl / Helm context (kubernetes target only) |
--container-runtime CHOICE | Container runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman. |
--help | Show this message and exit. Default: false. |
vllm-sr storage
Usage: vllm-sr storage [OPTIONS] COMMAND [ARGS]...
Inspect Router storage or manage local storage credentials.
| Parameter | Description |
|---|---|
--help | Show this message and exit. Default: false. |
vllm-sr storage rotate
Usage: vllm-sr storage rotate [OPTIONS]
Replace this stack's storage credentials with freshly generated ones.
Rotation generates new values, applies them to Postgres with ALTER ROLE
and to Redis by rebuilding it against the same named volume, then asks you
to re-run serve so Router picks the new values up.
Router has to be restarted, and this command deliberately does not do it
for you. Router receives the credentials as environment values captured
when its container was created, so restart would bring back the old ones
and only a re-create picks up the new ones -- and re-creating it here would
have to guess the images, profile, and Recipe bindings you originally
served with. Re-run your own serve command instead.
Until you do, the stack is degraded: ALTER ROLE takes effect at once, so
connections Router already holds keep working while every new one fails.
The order is forced -- restarting Router first would start it on a
credential Postgres has not accepted yet -- so run the two steps back to
back and plan the rotation for a moment when a brief restart is acceptable.
The scope is one stack, resolved from VLLM_SR_STACK_NAME exactly like
serve, stop, and status. Rotate other stacks one at a time. There is
deliberately no cross-stack mode: a failure partway through would leave
some stacks revoked and others not, with no value left to roll back to.
Examples:
vllm-sr storage rotate
VLLM_SR_STACK_NAME=staging vllm-sr storage rotate
| Parameter | Description |
|---|---|
--config TEXT | Config file whose directory holds the runtime state (default: config.yaml, matching vllm-sr serve). |
--runtime CHOICE | Container runtime for the local Docker target: docker, podman Choices: docker, podman. |
--help | Show this message and exit. Default: false. |
vllm-sr storage vector-stores
Usage: vllm-sr storage vector-stores [OPTIONS]
List vector stores known to the Router management API.
| Parameter | Description |
|---|---|
--endpoint TEXT | Router management origin or /api/v1 root. |
--timeout INTEGER | [default: 15] |
--token-env TEXT | [default: VSR_MGMT_TOKEN] |
--help | Show this message and exit. Default: false. |