Skip to main content
Version: Latest (unreleased)

CLI Commands

Use vllm-sr to configure and run the router, manage recipes, send requests, and evaluate routing. See Installation to install the CLI and Local Docker deployment to start a router.

This reference is generated from the registered CLI commands. Command descriptions, arguments, options, and declared defaults match the source. An option without a declared default is omitted until supplied; commands may resolve it from configuration or the environment as described in their help. Repeatable options collect values. Run vllm-sr COMMAND --help for help from your installed version.

Command index​

CommandDescription
vllm-srvLLM Semantic Router CLI - Signal-driven routing across LLM providers, with a built-in model runtime.
vllm-sr benchmarkPrepare, run, inspect, and compare sr-bench 1.0 evaluations.
vllm-sr benchmark cancelCancel remaining work while retaining all existing evidence.
vllm-sr benchmark candidate-planReuse a terminal baseline's frozen protocol without repeating its requests.
vllm-sr benchmark catalogShow the nine benchmark adapters and evaluation profiles.
vllm-sr benchmark compareCompare matched cases against the best observed single model.
vllm-sr benchmark comparison-optionsList eligible baselines, or comparable live runs for BASELINE; no model calls.
vllm-sr benchmark datasetPrepare reproducible fixed benchmark case sets.
vllm-sr benchmark dataset combineCreate a reusable multi-benchmark dataset from prepared manifests.
vllm-sr benchmark dataset exclusionsFreeze finite named history without inspecting outcomes or running models.
vllm-sr benchmark dataset optionsList downloadable sources, profiles, access notes, and dependencies.
vllm-sr benchmark dataset preparationsList shared download jobs or inspect one preparation by ID.
vllm-sr benchmark dataset prepareDownload and freeze a dataset through the shared service by default.
vllm-sr benchmark dataset showInspect a frozen dataset or list datasets in the shared store.
vllm-sr benchmark experimentGroup durable baseline, routing checks, candidates and validation runs.
vllm-sr benchmark experiment attachLink existing evidence without changing or rerunning it.
vllm-sr benchmark experiment createCreate an experiment without submitting model work.
vllm-sr benchmark experiment deleteDelete a finished experiment's grouping and links; keep every run and result.
vllm-sr benchmark experiment listRead one page of experiments.
vllm-sr benchmark experiment showRead the experiment and one page of its linked runs.
vllm-sr benchmark exportExport a dev response matrix for training; holdout export is rejected.
vllm-sr benchmark planValidate and freeze all cases, targets, profiles, and limits without inference.
vllm-sr benchmark previewInspect routing decisions without producing quality scores.
vllm-sr benchmark reconcile-usageAppend an offline accounting correction from saved streams; no inference.
vllm-sr benchmark recoverCreate a separate attempt from a reviewed recovery plan; never auto-retry.
vllm-sr benchmark recover-planInspect eligible continuation cells without making model requests.
vllm-sr benchmark regradeRegrade saved MCQ/grid final outputs without mutating original evidence.
vllm-sr benchmark replayEstimate eligible static routes from saved answers without inference.
vllm-sr benchmark replay-optionsList eligible baselines, or compatible previews for BASELINE; no model calls.
vllm-sr benchmark reportShow quality, four-bucket usage, cost, latency, time, and limitations.
vllm-sr benchmark runExecute one frozen live evaluation through the shared service.
vllm-sr benchmark runs
vllm-sr benchmark serveOwn the durable journal and workers independently of a browser.
vllm-sr benchmark setupInspect prerequisites or explicitly install optional benchmark harnesses.
vllm-sr benchmark showRead a run, bounded evidence page, or one complete saved call.
vllm-sr benchmark targetManage the operator-owned target registry used by the Dashboard.
vllm-sr benchmark target list
vllm-sr benchmark target registerReplace the local registry from a JSON list of credential references.
vllm-sr completionGenerate or install shell completion for vllm-sr.
vllm-sr completion installInstall shell completions into your shell configuration.
vllm-sr completion showPrint the completion script for a shell.
vllm-sr configPrint generated configuration or run config subcommands.
vllm-sr config applyPlan, compare-and-swap, persist, and hot-reload a configuration.
vllm-sr config envoyPrint the generated Envoy configuration.
vllm-sr config getRead the active canonical configuration from a Router.
vllm-sr config initCreate a minimal canonical configuration template.
vllm-sr config migrateMigrate a legacy or mixed config file to canonical v0.3 YAML.
vllm-sr config planValidate and plan an exact remote mutation without changing the Router.
vllm-sr config rollbackCompare-and-swap the active configuration to a recorded version.
vllm-sr config routerPrint the canonical router configuration.
vllm-sr config schemaDiscover the canonical config contract progressively.
vllm-sr config validateValidate configuration file.
vllm-sr config versionsList the configuration history, newest first.
vllm-sr dashboardOpen the dashboard in your default web browser.
vllm-sr instanceInspect the serving state of an existing local instance.
vllm-sr instance modelsPrint actual native model cards for readiness checks, without inference.
vllm-sr instance statusPrint desired/observed mode and durable operation state.
vllm-sr logsShow logs from vLLM Semantic Router service.
vllm-sr optimizeAnalyze routing evidence and produce candidate recipe changes.
vllm-sr optimize recipe-learningAnalyze replay and outcomes to produce recipe-learning artifacts.
vllm-sr recipeValidate, plan, apply, inspect, or package routing recipes.
vllm-sr recipe applyValidate then compare-and-swap one recipe into the active config.
vllm-sr recipe builtinDiscover, export, or bind installed Recipes without a running Router.
vllm-sr recipe builtin exportExport an exact five-file BUNDLE to a new directory.
vllm-sr recipe builtin initBind one installed recipe to explicit models and validate the new config.
vllm-sr recipe builtin listList installed bundles, recipe names, and required candidate pools as JSON.
vllm-sr recipe deleteCompare-and-swap deletion of an unreferenced named recipe.
vllm-sr recipe getRead one managed recipe.
vllm-sr recipe listList recipes and the collection ETag.
vllm-sr recipe packCreate a deterministic ZIP from an exact five-file RECIPE_DIR.
vllm-sr recipe planValidate a recipe and bind the plan to the current config ETag.
vllm-sr recipe validateValidate a recipe against the running Router without changing config.
vllm-sr requestSend requests through the stack's listener.
vllm-sr request chatSend a one-shot chat completion through the Envoy-routed HTTP API.
vllm-sr routePreview routing decisions or probe the routed inference path.
vllm-sr route previewPreview signals and model selection without generating an answer.
vllm-sr route probeProbe a real route and assert complete assistant delivery for expected 2xx.
vllm-sr serveStart vLLM Semantic Router.
vllm-sr statusShow status of vLLM Semantic Router services.
vllm-sr stopStop vLLM Semantic Router.
vllm-sr storageInspect Router storage or manage local storage credentials.
vllm-sr storage rotateReplace this stack's storage credentials with freshly generated ones.
vllm-sr storage vector-storesList vector stores known to the Router management API.

vllm-sr​

Usage: vllm-sr [OPTIONS] [COMMAND] [ARGS]...

vLLM Semantic Router CLI - Signal-driven routing across LLM providers, with a built-in model runtime.

ParameterDescription
--versionShow version and exit. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr benchmark​

Usage: vllm-sr benchmark [OPTIONS] COMMAND [ARGS]...

Prepare, run, inspect, and compare sr-bench 1.0 evaluations.

ParameterDescription
--url TEXTShared sr-bench service URL; discovers the current local stack by default. Environment: SR_BENCH_URL.
--store PATHDurable service store; discovers the current local stack by default. Environment: SR_BENCH_STORE.
--no-autostartRequire an already running service. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr benchmark cancel​

Usage: vllm-sr benchmark cancel [OPTIONS] RUN_ID

Cancel remaining work while retaining all existing evidence.

ParameterDescription
RUN_IDRequired argument. Type: text.
--helpShow this message and exit. Default: false.

vllm-sr benchmark candidate-plan​

Usage: vllm-sr benchmark candidate-plan [OPTIONS] BASELINE_RUN_ID

Reuse a terminal baseline's frozen protocol without repeating its requests.

Failed, cancelled and interrupted full-plan baselines may supply the same questions and settings. This does not qualify their measurements for Compare.

ParameterDescription
BASELINE_RUN_IDRequired argument. Type: text.
--target TEXTRegistered MoM target; may be repeated. [required] May be repeated.
--mode CHOICE[default: live] Choices: live, preview.
--name TEXT—
--experiment TEXT—
--hypothesis TEXTDefault: .
--helpShow this message and exit. Default: false.

vllm-sr benchmark catalog​

Usage: vllm-sr benchmark catalog [OPTIONS]

Show the nine benchmark adapters and evaluation profiles.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark compare​

Usage: vllm-sr benchmark compare [OPTIONS] BASELINE_RUN_ID CANDIDATE_RUN_ID

Compare matched cases against the best observed single model.

ParameterDescription
BASELINE_RUN_IDRequired argument. Type: text.
CANDIDATE_RUN_IDRequired argument. Type: text.
--helpShow this message and exit. Default: false.

vllm-sr benchmark comparison-options​

Usage: vllm-sr benchmark comparison-options [OPTIONS] [BASELINE]

List eligible baselines, or comparable live runs for BASELINE; no model calls.

ParameterDescription
[BASELINE]Optional argument. Type: text.
--after TEXTOpaque cursor from the previous eligible options page.
--limit INTEGER RANGE[default: 10; 1<=x<=25]
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset​

Usage: vllm-sr benchmark dataset [OPTIONS] COMMAND [ARGS]...

Prepare reproducible fixed benchmark case sets.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset combine​

Usage: vllm-sr benchmark dataset combine [OPTIONS] MANIFESTS...

Create a reusable multi-benchmark dataset from prepared manifests.

ParameterDescription
MANIFESTS...Required argument. Type: path. Accepts multiple values.
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset exclusions​

Usage: vllm-sr benchmark dataset exclusions [OPTIONS]

Freeze finite named history without inspecting outcomes or running models.

ParameterDescription
--dataset TEXTNamed prepared dataset whose entire membership is reserved. May be repeated.
--run TEXTNamed frozen run whose entire planned membership is reserved. May be repeated.
--output PATH[required]
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset options​

Usage: vllm-sr benchmark dataset options [OPTIONS]

List downloadable sources, profiles, access notes, and dependencies.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset preparations​

Usage: vllm-sr benchmark dataset preparations [OPTIONS] [PREPARATION_ID]

List shared download jobs or inspect one preparation by ID.

ParameterDescription
[PREPARATION_ID]Optional argument. Type: text.
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset prepare​

Usage: vllm-sr benchmark dataset prepare [OPTIONS]

Download and freeze a dataset through the shared service by default.

ParameterDescription
--benchmark TEXTBenchmark ID; repeat to prepare a shared collection. [required] May be repeated.
--profile CHOICEChoices: smoke, quick, standard. Default: quick.
--source-path PATH—
--localPrepare on this host; required for source files and history options. Default: false.
--no-waitReturn the shared service preparation job immediately. Default: false.
--revision TEXT—
--seed INTEGERDefault: 20260918.
--source-partition TEXTFrozen upstream partition for native task identity; never an evaluation split.
--exclusion-snapshot PATH—
--evaluation-role CHOICEExplicit family role; retest is never selected automatically. Choices: holdout, retest.
--limit INTEGERCustom case cap; cannot be represented as an upstream full benchmark.
--helpShow this message and exit. Default: false.

vllm-sr benchmark dataset show​

Usage: vllm-sr benchmark dataset show [OPTIONS] [PATH]

Inspect a frozen dataset or list datasets in the shared store.

ParameterDescription
[PATH]Optional argument. Type: path.
--helpShow this message and exit. Default: false.

vllm-sr benchmark experiment​

Usage: vllm-sr benchmark experiment [OPTIONS] COMMAND [ARGS]...

Group durable baseline, routing checks, candidates and validation runs.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark experiment attach​

Usage: vllm-sr benchmark experiment attach [OPTIONS] EXPERIMENT_ID

Link existing evidence without changing or rerunning it.

ParameterDescription
EXPERIMENT_IDRequired argument. Type: text.
--run TEXT[required]
--role CHOICE[required] Choices: baseline, candidate, estimate, initial, preview, recovery, smoke, validation.
--hypothesis TEXTDefault: .
--helpShow this message and exit. Default: false.

vllm-sr benchmark experiment create​

Usage: vllm-sr benchmark experiment create [OPTIONS] NAME

Create an experiment without submitting model work.

ParameterDescription
NAMERequired argument. Type: text.
--idempotency-key TEXT—
--helpShow this message and exit. Default: false.

vllm-sr benchmark experiment delete​

Usage: vllm-sr benchmark experiment delete [OPTIONS] EXPERIMENT_ID

Delete a finished experiment's grouping and links; keep every run and result.

ParameterDescription
EXPERIMENT_IDRequired argument. Type: text.
--helpShow this message and exit. Default: false.

vllm-sr benchmark experiment list​

Usage: vllm-sr benchmark experiment list [OPTIONS]

Read one page of experiments.

ParameterDescription
--after INTEGER RANGE[x>=0] Default: 0.
--limit INTEGER RANGE[1<=x<=50] Default: 20.
--helpShow this message and exit. Default: false.

vllm-sr benchmark experiment show​

Usage: vllm-sr benchmark experiment show [OPTIONS] EXPERIMENT_ID

Read the experiment and one page of its linked runs.

ParameterDescription
EXPERIMENT_IDRequired argument. Type: text.
--after INTEGER RANGE[x>=0] Default: 0.
--limit INTEGER RANGE[1<=x<=50] Default: 20.
--helpShow this message and exit. Default: false.

vllm-sr benchmark export​

Usage: vllm-sr benchmark export [OPTIONS] RUN_ID

Export a dev response matrix for training; holdout export is rejected.

ParameterDescription
RUN_IDRequired argument. Type: text.
--output PATH[required]
--helpShow this message and exit. Default: false.

vllm-sr benchmark plan​

Usage: vllm-sr benchmark plan [OPTIONS]

Validate and freeze all cases, targets, profiles, and limits without inference.

ParameterDescription
--manifest PATH[required]
--output PATH—
--helpShow this message and exit. Default: false.

vllm-sr benchmark preview​

Usage: vllm-sr benchmark preview [OPTIONS]

Inspect routing decisions without producing quality scores.

ParameterDescription
--idempotency-key TEXTBind repeated submissions to the same frozen plan, without reissuing calls.
--detachReturn immediately with a durable run ID. Default: false.
--manifest PATH[required]
--helpShow this message and exit. Default: false.

vllm-sr benchmark reconcile-usage​

Usage: vllm-sr benchmark reconcile-usage [OPTIONS] RUN_ID

Append an offline accounting correction from saved streams; no inference.

ParameterDescription
RUN_IDRequired argument. Type: text.
--helpShow this message and exit. Default: false.

vllm-sr benchmark recover​

Usage: vllm-sr benchmark recover [OPTIONS] RUN_ID

Create a separate attempt from a reviewed recovery plan; never auto-retry.

ParameterDescription
RUN_IDRequired argument. Type: text.
--plan PATH[required]
--idempotency-key TEXT[required]
--acknowledge-new-attemptAuthorize new paid attempts for the exact reviewed failed cells. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr benchmark recover-plan​

Usage: vllm-sr benchmark recover-plan [OPTIONS] RUN_ID

Inspect eligible continuation cells without making model requests.

ParameterDescription
RUN_IDRequired argument. Type: text.
--mode CHOICEChoices: undispatched, failed. Default: undispatched.
--output PATH—
--helpShow this message and exit. Default: false.

vllm-sr benchmark regrade​

Usage: vllm-sr benchmark regrade [OPTIONS] RUN_ID

Regrade saved MCQ/grid final outputs without mutating original evidence.

ParameterDescription
RUN_IDRequired argument. Type: text.
--output PATH[required]
--helpShow this message and exit. Default: false.

vllm-sr benchmark replay​

Usage: vllm-sr benchmark replay [OPTIONS]

Estimate eligible static routes from saved answers without inference.

ParameterDescription
--baseline TEXTCompleted single-model answer matrix run ID. [required]
--preview TEXTCompleted deterministic routing preview run ID. [required]
--idempotency-key TEXT—
--helpShow this message and exit. Default: false.

vllm-sr benchmark replay-options​

Usage: vllm-sr benchmark replay-options [OPTIONS] [BASELINE]

List eligible baselines, or compatible previews for BASELINE; no model calls.

ParameterDescription
[BASELINE]Optional argument. Type: text.
--after TEXTOpaque cursor from the previous eligible options page.
--limit INTEGER RANGE[default: 10; 1<=x<=25]
--helpShow this message and exit. Default: false.

vllm-sr benchmark report​

Usage: vllm-sr benchmark report [OPTIONS] RUN_ID

Show quality, four-bucket usage, cost, latency, time, and limitations.

ParameterDescription
RUN_IDRequired argument. Type: text.
--output PATH—
--helpShow this message and exit. Default: false.

vllm-sr benchmark run​

Usage: vllm-sr benchmark run [OPTIONS]

Execute one frozen live evaluation through the shared service.

ParameterDescription
--idempotency-key TEXTBind repeated submissions to the same frozen plan, without reissuing calls.
--detachReturn immediately with a durable run ID. Default: false.
--manifest PATH[required]
--helpShow this message and exit. Default: false.

vllm-sr benchmark runs​

Usage: vllm-sr benchmark runs [OPTIONS]
ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark serve​

Usage: vllm-sr benchmark serve [OPTIONS]

Own the durable journal and workers independently of a browser.

ParameterDescription
--host TEXTDefault: 127.0.0.1.
--port INTEGERDefault: 8090.
--store-identity TEXTCanonical host store path SHA256 for an isolated runtime container.
--helpShow this message and exit. Default: false.

vllm-sr benchmark setup​

Usage: vllm-sr benchmark setup [OPTIONS]

Inspect prerequisites or explicitly install optional benchmark harnesses.

ParameterDescription
--benchmark TEXTDefault: all.
--installInstall pinned optional harnesses and task sources; makes no model requests. Default: false.
--build-sandboxBuild a local offline grading image and record its content digest. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr benchmark show​

Usage: vllm-sr benchmark show [OPTIONS] RUN_ID

Read a run, bounded evidence page, or one complete saved call.

ParameterDescription
RUN_IDRequired argument. Type: text.
--resultsDefault: false.
--callsDefault: false.
--activeRead only in-progress calls; requires --calls. Default: false.
--eventsDefault: false.
--after INTEGER RANGEEvidence cursor from the previous page. [x>=0] Default: 0.
--limit INTEGER RANGECalls/results per page. [1<=x<=500] Default: 100.
--call-id TEXTRead one full saved call including prompt and final response.
--helpShow this message and exit. Default: false.

vllm-sr benchmark target​

Usage: vllm-sr benchmark target [OPTIONS] COMMAND [ARGS]...

Manage the operator-owned target registry used by the Dashboard.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark target list​

Usage: vllm-sr benchmark target list [OPTIONS]
ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr benchmark target register​

Usage: vllm-sr benchmark target register [OPTIONS]

Replace the local registry from a JSON list of credential references.

ParameterDescription
--file PATH[required]
--helpShow this message and exit. Default: false.

vllm-sr completion​

Usage: vllm-sr completion [OPTIONS] [COMMAND] [ARGS]...

Generate or install shell completion for vllm-sr.

Examples:
vllm-sr completion show bash # Print bash completion script
vllm-sr completion show zsh # Print zsh completion script
vllm-sr completion install # Auto-install for current shell
vllm-sr completion install bash # Install bash completions
ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr completion install​

Usage: vllm-sr completion install [OPTIONS] [bash|zsh|fish]

Install shell completions into your shell configuration.

Automatically appends the completion setup to your shell's rc file (~/.bashrc, ~/.zshrc) or writes to the fish completions directory. Safe to run multiple times — skips if already configured.

Examples:
vllm-sr completion install # Auto-detect shell
vllm-sr completion install bash # Install for bash
vllm-sr completion install zsh # Install for zsh
vllm-sr completion install fish # Install for fish
ParameterDescription
[SHELL]Optional argument. Type: choice. Choices: bash, zsh, fish.
--helpShow this message and exit. Default: false.

vllm-sr completion show​

Usage: vllm-sr completion show [OPTIONS] [bash|zsh|fish]

Print the completion script for a shell.

Outputs the completion script to stdout. If SHELL is omitted, the command attempts to detect the current shell automatically.

Examples:
vllm-sr completion show bash
vllm-sr completion show zsh
vllm-sr completion show fish
eval "$(vllm-sr completion show zsh)"
ParameterDescription
[SHELL]Optional argument. Type: choice. Choices: bash, zsh, fish.
--helpShow this message and exit. Default: false.

vllm-sr config​

Usage: vllm-sr config [OPTIONS] [COMMAND] [ARGS]...

Print generated configuration or run config subcommands.

Examples:

vllm-sr config init --output config.yaml
vllm-sr config validate --config config.yaml
vllm-sr config apply --config config.yaml
vllm-sr config router
vllm-sr config migrate --config old.yaml
vllm-sr config envoy # with --gateway extproc
ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr config apply​

Usage: vllm-sr config apply [OPTIONS]

Plan, compare-and-swap, persist, and hot-reload a configuration.

A change the running Router can't take without a restart is saved for the next vllm-sr serve of a local stack, as the Dashboard saves one.

ParameterDescription
--config FILE[default: config.yaml]
--mode CHOICE[default: replace] Choices: replace, merge.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 120]
--token-env TEXTEnvironment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr config envoy​

Usage: vllm-sr config envoy [OPTIONS]

Print the generated Envoy configuration.

ParameterDescription
--config TEXTPath to config file (default: config.yaml) Default: config.yaml.
--helpShow this message and exit. Default: false.

vllm-sr config get​

Usage: vllm-sr config get [OPTIONS]

Read the active canonical configuration from a Router.

ParameterDescription
--format CHOICE[default: yaml] Choices: json, yaml.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXTEnvironment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr config init​

Usage: vllm-sr config init [OPTIONS]

Create a minimal canonical configuration template.

ParameterDescription
--output TEXTPath for the new canonical configuration template. [default: config.yaml]
--forceOverwrite the output file if it already exists. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr config migrate​

Usage: vllm-sr config migrate [OPTIONS]

Migrate a legacy or mixed config file to canonical v0.3 YAML.

ParameterDescription
--config TEXTPath to source config file (default: config.yaml) Default: config.yaml.
--output TEXTPath for migrated canonical config (default: <config>.migrated.yaml)
--forceOverwrite the output file if it already exists. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr config plan​

Usage: vllm-sr config plan [OPTIONS]

Validate and plan an exact remote mutation without changing the Router.

ParameterDescription
--config FILE[default: config.yaml]
--mode CHOICE[default: replace] Choices: replace, merge.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXTEnvironment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr config rollback​

Usage: vllm-sr config rollback [OPTIONS] VERSION

Compare-and-swap the active configuration to a recorded version.

VERSION is a configuration version number from vllm-sr config versions, or the timestamp of a backup. The restored document activates as a new version.

ParameterDescription
VERSIONRequired argument. Type: text.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 120]
--token-env TEXTEnvironment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr config router​

Usage: vllm-sr config router [OPTIONS]

Print the canonical router configuration.

ParameterDescription
--config TEXTPath to config file (default: config.yaml) Default: config.yaml.
--helpShow this message and exit. Default: false.

vllm-sr config schema​

Usage: vllm-sr config schema [OPTIONS]

Discover the canonical config contract progressively.

ParameterDescription
--endpoint TEXTRead the contract from a running Router management origin instead of the schema bundled with this CLI.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--fullPrint the complete JSON Schema instead of the compact index. Default: false.
--section PATHPrint one config path and only its referenced definitions.
--surface KIND:NAMEPrint one signal, algorithm, plugin, or projection contract.
--expandedInclude the selected section's self-contained JSON Schema. Default: false.
--helpShow this message and exit. Default: false.

vllm-sr config validate​

Usage: vllm-sr config validate [OPTIONS]

Validate configuration file.

The CLI's own checks run first. The Router's validation then decides, as it does when vllm-sr serve or vllm-sr config apply loads the file: the Router in its local image (never pulled), or the running Router --endpoint names.

Examples:

vllm-sr config validate
vllm-sr config validate --config my-config.yaml
vllm-sr config validate --endpoint http://localhost:8080
ParameterDescription
--config TEXTPath to config file (default: config.yaml) Default: config.yaml.
--endpoint TEXTValidate with this running Router instead of the local Router image.
--image TEXTRouter image whose own validation to run (default: the stack's, if present).
--gateway CHOICEThe gateway mode the configuration is served in (default: standalone). Choices: standalone, extproc.
--offlineRun only the CLI's own checks, without the Router's validation. Default: false.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr config versions​

Usage: vllm-sr config versions [OPTIONS]

List the configuration history, newest first.

ParameterDescription
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXTEnvironment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr dashboard​

Usage: vllm-sr dashboard [OPTIONS]

Open the dashboard in your default web browser.

Examples:

vllm-sr dashboard # Docker dashboard
vllm-sr dashboard --target kubernetes # Show K8s address and port forward
vllm-sr dashboard --no-open
ParameterDescription
--no-openDon't open browser, just show URL Default: false.
--target TEXTDeployment target: docker, kubernetes (default: docker)
--namespace TEXTKubernetes namespace (kubernetes target only)
--context TEXTkubectl / Helm context (kubernetes target only)
--container-runtime CHOICEContainer runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman.
--helpShow this message and exit. Default: false.

vllm-sr instance​

Usage: vllm-sr instance [OPTIONS] COMMAND [ARGS]...

Inspect the serving state of an existing local instance.

ParameterDescription
--config FILEDefault: config.yaml.
--helpShow this message and exit. Default: false.

vllm-sr instance models​

Usage: vllm-sr instance models [OPTIONS]

Print actual native model cards for readiness checks, without inference.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr instance status​

Usage: vllm-sr instance status [OPTIONS]

Print desired/observed mode and durable operation state.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr logs​

Usage: vllm-sr logs [OPTIONS] {envoy|router|dashboard}

Show logs from vLLM Semantic Router service.

Examples:

vllm-sr logs envoy
vllm-sr logs router
vllm-sr logs dashboard
vllm-sr logs envoy --follow
vllm-sr logs router -f
vllm-sr logs router --target kubernetes # Kubernetes logs
vllm-sr logs router --target kubernetes -f # Follow K8s logs
ParameterDescription
SERVICERequired argument. Type: choice. Choices: envoy, router, dashboard.
-f, --followFollow log output Default: false.
--target TEXTDeployment target: docker, kubernetes (default: docker)
--namespace TEXTKubernetes namespace (kubernetes target only)
--context TEXTkubectl / Helm context (kubernetes target only)
--container-runtime CHOICEContainer runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman.
--helpShow this message and exit. Default: false.

vllm-sr optimize​

Usage: vllm-sr optimize [OPTIONS] COMMAND [ARGS]...

Analyze routing evidence and produce candidate recipe changes.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr optimize recipe-learning​

Usage: vllm-sr optimize recipe-learning [OPTIONS]

Analyze replay and outcomes to produce recipe-learning artifacts.

ParameterDescription
--replay-file FILERouter replay JSON file. Accepts a router_replay.list payload or a record array.
--endpoint TEXTRouter management base URL (origin or /api/v1). Defaults to http://localhost:8080, with VLLM_SR_PORT_OFFSET added to the port, when --replay-file is omitted. Uses VSR_MGMT_TOKEN for bearer auth when set.
--cases-file FILEOptional eval cases JSON with replay_id/request_id plus expected_decision or expected_model.
--recipe-file FILEOptional current recipe YAML used to materialize complete candidate recipe variants.
--limit INTEGERReplay records to fetch from the endpoint. [default: 100]
--output-dir DIRECTORYDirectory for metrics/findings/patch/seed-pack artifacts.
--jsonPrint the full recipe-learning artifact. Default: false.
--report-onlyCompute metrics and findings without generating recipe patches. Default: false.
--timeout INTEGERHTTP request timeout in seconds. [default: 15]
--helpShow this message and exit. Default: false.

vllm-sr recipe​

Usage: vllm-sr recipe [OPTIONS] COMMAND [ARGS]...

Validate, plan, apply, inspect, or package routing recipes.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr recipe apply​

Usage: vllm-sr recipe apply [OPTIONS] RECIPE_FILE

Validate then compare-and-swap one recipe into the active config.

ParameterDescription
RECIPE_FILERequired argument. Type: file.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr recipe builtin​

Usage: vllm-sr recipe builtin [OPTIONS] COMMAND [ARGS]...

Discover, export, or bind installed Recipes without a running Router.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr recipe builtin export​

Usage: vllm-sr recipe builtin export [OPTIONS] BUNDLE

Export an exact five-file BUNDLE to a new directory.

ParameterDescription
BUNDLERequired argument. Type: text.
--output-dir DIRECTORY[required]
--catalog-version TEXT[default: latest]
--helpShow this message and exit. Default: false.

vllm-sr recipe builtin init​

Usage: vllm-sr recipe builtin init [OPTIONS] NAME

Bind one installed recipe to explicit models and validate the new config.

No model, endpoint, capability, or candidate pool is invented. Every recipe decision must meet its declared minimum_candidates before this succeeds.

ParameterDescription
NAMERequired argument. Type: text.
--bundle TEXTBundle from recipe builtin list. [required]
--config FILEExisting canonical provider config to preserve. [required]
--bindings FILEYAML mapping each decision that calls a backend to modelRefs; omit immediate responses. [required]
--model-name TEXTExplicit public entrypoint name. [required]
--exclude-decision TEXTExplicitly omit a named lane for an unavailable capability; creates a recipe derivative. Repeatable. May be repeated.
--output FILENew config path; existing files are never replaced. Keep beside --config when it references local KB assets. [required]
--catalog-version TEXT[default: latest]
--helpShow this message and exit. Default: false.

vllm-sr recipe builtin list​

Usage: vllm-sr recipe builtin list [OPTIONS]

List installed bundles, recipe names, and required candidate pools as JSON.

ParameterDescription
--catalog-version TEXT[default: latest]
--helpShow this message and exit. Default: false.

vllm-sr recipe delete​

Usage: vllm-sr recipe delete [OPTIONS] NAME

Compare-and-swap deletion of an unreferenced named recipe.

ParameterDescription
NAMERequired argument. Type: text.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr recipe get​

Usage: vllm-sr recipe get [OPTIONS] NAME

Read one managed recipe.

ParameterDescription
NAMERequired argument. Type: text.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr recipe list​

Usage: vllm-sr recipe list [OPTIONS]

List recipes and the collection ETag.

ParameterDescription
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr recipe pack​

Usage: vllm-sr recipe pack [OPTIONS] RECIPE_DIR

Create a deterministic ZIP from an exact five-file RECIPE_DIR.

ParameterDescription
RECIPE_DIRRequired argument. Type: directory.
--output PATHArchive path or output directory (default: the Recipe parent directory).
--helpShow this message and exit. Default: false.

vllm-sr recipe plan​

Usage: vllm-sr recipe plan [OPTIONS] RECIPE_FILE

Validate a recipe and bind the plan to the current config ETag.

ParameterDescription
RECIPE_FILERequired argument. Type: file.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr recipe validate​

Usage: vllm-sr recipe validate [OPTIONS] RECIPE_FILE

Validate a recipe against the running Router without changing config.

ParameterDescription
RECIPE_FILERequired argument. Type: file.
--endpoint TEXTRouter management base URL; defaults to the local Router API port.
--timeout FLOAT[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.

vllm-sr request​

Usage: vllm-sr request [OPTIONS] COMMAND [ARGS]...

Send requests through the stack's listener.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr request chat​

Usage: vllm-sr request chat [OPTIONS] [MESSAGE]...

Send a one-shot chat completion through the Envoy-routed HTTP API.

Uses the first listener port in config.yaml plus the stack port offset. The default model is the namespaced automatic-routing alias vllm-sr/auto.

Examples:

vllm-sr request chat "hello"
vllm-sr request chat --model vllm-sr/auto --prompt "Explain mixture of models"
vllm-sr request chat --json "hello"
ParameterDescription
[MESSAGE]...Optional argument. Type: text. Accepts multiple values.
--prompt TEXTUser message (alternative to a positional prompt).
--model TEXTModel name sent to the router. [default: vllm-sr/auto]
--system TEXTOptional system message prepended to the conversation.
--config TEXTConfig file used to resolve listener host port (Docker default only). [default: config.yaml]
--base-url TEXTExplicit routed listener origin or OpenAI /v1 base URL for a remote or port-forwarded stack.
--jsonPrint the raw JSON response instead of assistant text. Default: false.
--timeout FLOATHTTP timeout in seconds for the completion request. [default: 120.0]
--temperature FLOATOptional sampling temperature passed through to the API.
--target TEXTDeployment target (docker or k8s).
--helpShow this message and exit. Default: false.

vllm-sr route​

Usage: vllm-sr route [OPTIONS] COMMAND [ARGS]...

Preview routing decisions or probe the routed inference path.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr route preview​

Usage: vllm-sr route preview [OPTIONS]

Preview signals and model selection without generating an answer.

ParameterDescription
--prompt TEXTPlain text prompt to evaluate.
--messages TEXTOpenAI-style messages JSON array string.
--model TEXTRouting model or entrypoint whose recipe should be evaluated.
--request-file FILEJSON Router Preview request; supported Chat messages and prompt fields only.
--session-id TEXTRead-only Learning session identity.
--conversation-id TEXTRead-only Learning conversation identity.
--sampling-seed INTEGER RANGEPreview-only exploration seed; does not fix a later live random draw. [-9223372036854775808<=x<=9223372036854775807]
--endpoint TEXTRouter management origin or /api/v1 root; defaults to the local management port.
--token-env TEXTEnvironment variable containing the Router management bearer token. [default: VSR_MGMT_TOKEN]
--trace / --no-traceInclude per-decision routing trace trees. Default: false.
--jsonPrint the full JSON response payload. Default: false.
--timeout FLOATHTTP request timeout in seconds. [default: 15.0]
--helpShow this message and exit. Default: false.

vllm-sr route probe​

Usage: vllm-sr route probe [OPTIONS]

Probe a real route and assert complete assistant delivery for expected 2xx.

ParameterDescription
--prompt TEXTPlain-text user prompt.
--messages TEXTOpenAI-style messages JSON array string.
--model TEXT[default: vllm-sr/auto]
--config TEXT[default: config.yaml]
--base-url TEXTExplicit Envoy listener origin or OpenAI /v1 base URL; otherwise derive it from --config.
--api-key-env TEXTEnvironment variable containing the bearer token; omitted when unset. [default: OPENAI_API_KEY]
--temperature FLOAT—
--max-completion-tokens INTEGER RANGECompletion token budget, including reasoning; omitted uses the backend default. [x>=1]
--timeout FLOAT[default: 120.0]
--target TEXTDeployment target used for URL resolution.
--debug / --no-debug[default: debug]
--expect-status INTEGER[default: OK]
--expect-recipe TEXT—
--expect-decision TEXT—
--expect-algorithm TEXT—
--expect-selected-model TEXTAssert the Router's x-vsr-selected-model receipt header.
--expect-response-model TEXTAssert the upstream OpenAI response body's top-level model field.
--helpShow this message and exit. Default: false.

vllm-sr serve​

Usage: vllm-sr serve [OPTIONS] [MODEL]

Start vLLM Semantic Router.

Serve uses --config or config.yaml and preserves the Dashboard-first setup flow. Connect physical models and publish Mixture-of-Model entrypoints in the Dashboard, then keep the same stack running with this single command.

Virtual models are routing policies. Semantic Router starts the Router, which serves the OpenAI-compatible API itself, the Dashboard and supporting services; it does not download or launch the physical LLM engines referenced by provider backends. Connect user-owned single or multiple model endpoints through one canonical config or the Dashboard.

Ports are configured in the selected config under the listeners section.

Local startup waits up to 1800 seconds for Router readiness, or Dashboard readiness during first-run setup. During setup the command keeps waiting, and once you activate a config in the Dashboard it starts the Router from it. Use --startup-timeout SECONDS for a different positive budget when model loading or GPU compilation needs more time. The wait begins after containers start. It does not change inference deadlines. Timeout exits the CLI with an error and leaves containers available for inspection.

GATEWAY MODES:

standalone - The Router serves the OpenAI-compatible API itself (default)
extproc - An Envoy-based gateway in front of the Router: the Envoy container
on docker, as before standalone became the default; your gateway
on kubernetes

DEPLOYMENT TARGETS:

docker - Local Docker deployment (default)
kubernetes - Kubernetes deployment via Helm (k8s is the old name, for this
release only)

MODEL SELECTION ALGORITHMS:

static - Use first configured model (default, no learning)
router_dc - Query-model matching via embedding similarity
automix - Cost-quality optimization using POMDP
hybrid - Combine multiple methods with configurable weights
workflows - Router Flow static/dynamic micro-agent orchestration
latency_aware - TPOT/TTFT percentile-aware selection
knn - KNN selector using shared ML model-selection settings
kmeans - KMeans selector using shared ML model-selection settings
svm - SVM selector using shared ML model-selection settings
mlp - MLP selector using shared ML model-selection settings
multi_factor - Quality, latency, cost, and load scoring

Cross-request learning lives under global.router.learning.adaptation and global.router.learning.protection instead of --algorithm.

Examples:

# Dashboard-first setup or an existing ./config.yaml
vllm-sr serve
# Envoy in front of the Router, as before standalone became the default
vllm-sr serve --gateway extproc
# User-owned single or multi-model topology
vllm-sr serve --config my-models.yaml
# Explicitly replace Dashboard-edited runtime state from reviewed source YAML
vllm-sr serve --config my-models.yaml --replace-active-config
# Deploy a user-owned config to Kubernetes
vllm-sr serve --target kubernetes --config my-models.yaml --namespace my-ns
# Runtime policy and image overrides
vllm-sr serve --algorithm latency_aware
vllm-sr serve --image-pull-policy always
vllm-sr serve --readonly
vllm-sr serve --minimal
vllm-sr serve --log-level debug
# AMD ROCm image, device passthrough, and router internal GPU defaults
vllm-sr serve --platform rocm
vllm-sr serve --platform rocm --startup-timeout 7200
VLLM_SR_AMD_ROUTER_VISIBLE_DEVICES=7 vllm-sr serve --platform rocm
INSTANCE MODES:
vllm-sr serve
vllm-sr serve vllm-sr/Decision-2.0-Kai-0.6B --engine
vllm-sr serve vllm-sr/Vela-2.0-4B --platform rocm -dp 2 --device-ids 0

Without --engine the instance starts in Router mode, including on restart. --engine (-e) disables recipe routing; the frontend, Dashboard and native System One APIs remain available. Saved routing configuration is retained.

MODEL overrides the configured default judgment deployment's artifact. Omitting MODEL preserves that deployment (a new configuration uses Vela 2.0 0.3B). Only explicit model/placement options override saved settings. Backend LLMs, named deployments, listeners and API grants belong in --config.

ParameterDescription
[MODEL]Optional argument. Type: text.
-e, --engineStart without recipe routing; otherwise start Router mode. Default: false.
--config TEXTPath to the Router configuration. [default: config.yaml]
--replace-active-configReplace this local Docker stack's active runtime config from --config, discarding Dashboard edits. Default: false.
--image TEXTDocker image to use (default: ghcr.io/vllm-project/semantic-router/vllm-sr:latest)
--router-image TEXTDocker image for the router container (Docker target only; defaults to --image or VLLM_SR_IMAGE)
--envoy-image TEXTDocker image for the Envoy container (docker target with --gateway extproc; defaults to --image or VLLM_SR_IMAGE)
--dashboard-image TEXTDocker image for the dashboard container (Docker target only; defaults to --image or VLLM_SR_IMAGE)
--image-pull-policy CHOICEImage pull policy: always, ifnotpresent, never (default: always) Choices: always, ifnotpresent, never. Default: always.
--startup-timeout SECONDSLocal Docker startup readiness budget in seconds, including model loading and compilation (default: 1800). [x>=1]
--readonlyRun dashboard in read-only mode (disable config editing, allow playground only) Default: false.
--minimalStart in minimal mode: no Dashboard or observability stack (Jaeger, Prometheus, Grafana) Default: false.
--log-level CHOICELog level of the Router, or of the runtime in engine mode (debug, info, warn, error, dpanic, panic, fatal) Choices: debug, info, warn, warning, error, dpanic, panic, fatal.
--platform CHOICEExecution backend: auto (default) discovers the deployment target; cpu, cuda or rocm select it explicitly. Choices: auto, cpu, cuda, rocm.
--algorithm CHOICERequest-time base algorithm override for payload-safe algorithms: static, router_dc, automix, hybrid, workflows, latency_aware, knn, kmeans, svm, mlp, multi_factor. Algorithms that require an authored payload remain available in config.yaml. Cross-request learning uses global.router.learning.adaptation/protection. Choices: static, router_dc, automix, hybrid, workflows, latency_aware, knn, kmeans, svm, mlp, multi_factor.
--target TEXTDeployment target: docker, kubernetes (default: docker)
--gateway CHOICEWhere client traffic enters: standalone (default; the Router serves the OpenAI-compatible API on the config's listeners, with no Envoy) or extproc (an Envoy-based gateway in front of the Router: the Envoy container on the docker target, your gateway on kubernetes). Choices: standalone, extproc.
--namespace TEXTKubernetes namespace (kubernetes target only)
--context TEXTkubectl / Helm context (kubernetes target only)
--profile TEXTDeployment profile: dev, prod (kubernetes target only). Selects values-<profile>.yaml defaults.
--chart-dir TEXTPath to Helm chart directory (kubernetes target only; default: ./deploy/helm/semantic-router, else the published chart for this version)
--container-runtime CHOICEContainer runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman.
--recipe-env NAMEExplicitly bind one host environment variable for the active Recipe. Repeat for multiple names; NAME=value is rejected. May be repeated.
--revision TEXTOptional model branch, tag or commit; resolved once to an immutable startup revision.
-dp, --data-parallel-size INTEGER RANGENumber of model replicas. Preserve configured placement; new GPU deployments use distinct available GPUs. [1<=x<=64]
--device-ids IDSDocker host GPU indices, e.g. 0 or 0,1. One index shares a GPU across replicas; otherwise use one per replica. Existing visibility masks are respected, not changed.
--runtime-profile PROFILEModel runtime numerics profile (default exact; vllm-srun plugins lists the installed ones).
--helpShow this message and exit. Default: false.

vllm-sr status​

Usage: vllm-sr status [OPTIONS] [envoy|router|dashboard|all]

Show status of vLLM Semantic Router services.

Examples:

vllm-sr status # Show all services (Docker)
vllm-sr status all # Show all services
vllm-sr status router # Show router status
vllm-sr status dashboard # Show dashboard status
vllm-sr status --target kubernetes # Show Kubernetes status
ParameterDescription
[SERVICE]Optional argument. Type: choice. Choices: envoy, router, dashboard, all. Default: all.
--target TEXTDeployment target: docker, kubernetes (default: docker)
--namespace TEXTKubernetes namespace (kubernetes target only)
--context TEXTkubectl / Helm context (kubernetes target only)
--container-runtime CHOICEContainer runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman.
--helpShow this message and exit. Default: false.

vllm-sr stop​

Usage: vllm-sr stop [OPTIONS]

Stop vLLM Semantic Router.

Examples:

vllm-sr stop # Stop Docker stack
vllm-sr stop --target kubernetes # Uninstall Helm release
ParameterDescription
--target TEXTDeployment target: docker, kubernetes (default: docker)
--namespace TEXTKubernetes namespace (kubernetes target only)
--context TEXTkubectl / Helm context (kubernetes target only)
--container-runtime CHOICEContainer runtime: docker, podman. Equivalent to setting CONTAINER_RUNTIME=<runtime>. Choices: docker, podman.
--helpShow this message and exit. Default: false.

vllm-sr storage​

Usage: vllm-sr storage [OPTIONS] COMMAND [ARGS]...

Inspect Router storage or manage local storage credentials.

ParameterDescription
--helpShow this message and exit. Default: false.

vllm-sr storage rotate​

Usage: vllm-sr storage rotate [OPTIONS]

Replace this stack's storage credentials with freshly generated ones.

Rotation generates new values, applies them to Postgres with ALTER ROLE and to Redis by rebuilding it against the same named volume, then asks you to re-run serve so Router picks the new values up.

Router has to be restarted, and this command deliberately does not do it for you. Router receives the credentials as environment values captured when its container was created, so restart would bring back the old ones and only a re-create picks up the new ones -- and re-creating it here would have to guess the images, profile, and Recipe bindings you originally served with. Re-run your own serve command instead.

Until you do, the stack is degraded: ALTER ROLE takes effect at once, so connections Router already holds keep working while every new one fails. The order is forced -- restarting Router first would start it on a credential Postgres has not accepted yet -- so run the two steps back to back and plan the rotation for a moment when a brief restart is acceptable.

The scope is one stack, resolved from VLLM_SR_STACK_NAME exactly like serve, stop, and status. Rotate other stacks one at a time. There is deliberately no cross-stack mode: a failure partway through would leave some stacks revoked and others not, with no value left to roll back to.

Examples:

vllm-sr storage rotate
VLLM_SR_STACK_NAME=staging vllm-sr storage rotate
ParameterDescription
--config TEXTConfig file whose directory holds the runtime state (default: config.yaml, matching vllm-sr serve).
--runtime CHOICEContainer runtime for the local Docker target: docker, podman Choices: docker, podman.
--helpShow this message and exit. Default: false.

vllm-sr storage vector-stores​

Usage: vllm-sr storage vector-stores [OPTIONS]

List vector stores known to the Router management API.

ParameterDescription
--endpoint TEXTRouter management origin or /api/v1 root.
--timeout INTEGER[default: 15]
--token-env TEXT[default: VSR_MGMT_TOKEN]
--helpShow this message and exit. Default: false.