Release note: standalone mode is the default
The first release that includes
#4623 serves the
OpenAI-compatible API from the Router itself. vllm-sr serve starts no Envoy
container by default: the Router container publishes the config's listeners,
answers /health and /ready on them, and proxies to the model backends. The
routing decisions, the upstream requests and the responses are the same as
behind Envoy. See Gateway Modes.
Keep the previous stack
vllm-sr serve --gateway extproc
--gateway extproc starts the Envoy container in front of the Router exactly
as earlier releases did, with the same Envoy configuration. Choose it for Envoy
features the standalone Router does not have yet, such as token-bucket rate
limiting, mTLS, JWT or OIDC, or advanced route matching. Envoy mode also
sustains about 5–10% more requests per second at 32 or more concurrent clients;
below that, standalone mode answers faster (see the design doc's
results).
What changes for a standalone stack
- No Envoy container.
vllm-sr statusandvllm-sr logsreport the Router and the Dashboard;vllm-sr logs envoyand--envoy-imageneed--gateway extproc. The Dashboard's Playground and readiness checks use the Router's listener. - Proxy-control headers from clients are dropped.
x-envoy-internaland Envoy's retry, timeout and tracing controls (such asx-envoy-max-retriesandx-envoy-upstream-rq-timeout-ms) never reach a backend, as behind Envoy. - Identity headers from clients are dropped, both
x-authz-*and the namesglobal.services.authz.identitysets, because no authenticator stands in front of the Router. If a proxy or your application sets them, turn onlisteners[].identity.trust_headers(optionally withtrusted_peers). Without a trusting listener, a config with a decision on anauthzsignal, a rate limit rule that matchesuserorgroup, or authz providers is refused at startup, and memory, router replay and the per-user selection algorithms treat every request as anonymous. - Listener changes need a restart. A reload that changes a listener's address,
port, timeout or
tlspaths, or adds or removes a listener, is rejected asrestart_required; runvllm-sr serveagain.
New
listeners[].tls(cert_file,key_file) serves a listener over TLS in standalone mode, and a renewed certificate in those files serves new connections without a restart.--gateway extprocrefuses it rather than serve the listener in cleartext.--gateway standalone|extprocand--platform cpu|amd|nvidiawork on the kubernetes target too. The CLI writes them into the generated Helm values asgateway.mode, the image repository and a GPU request.vllm-sr serve --helplists its options by group, and an option of another group is an error that says where it applies.- The Router binary takes
-gateway standalone. Its default staysextproc, so a manifest that runs it without the flag behaves as before. - On macOS,
--platform amd|nvidiafails with a clear message: Docker's Linux VM gets no GPU there, so the docker target runs the CPU image. - Timeouts, retries and fallback, per model and per decision.
providers.models[].reliabilitygains connect, total, idle, per-try and first-byte timeouts, retriable status codes, back-off,Retry-Afterand retry budgets. A decision'sreliabilityandfallbackblocks override them for the requests it routes, in both gateway modes;first_byte_timeoutis standalone only. See Tune timeouts, retries, and endpoint health and Fall back to another model. - Configuration versions and rollback. Every accepted change activates a
numbered version, and each response names it in
x-vsr-config-version. A rejected change leaves the active version serving and reports why. The last ten activations survive a restart, andPOST /api/v1/config/rollbackactivates a recorded version as a new one. See Configuration Management.
Kubernetes runs standalone too
- The Helm chart sets
gateway.mode: standaloneby default. The Router serves the listeners in its config, the Service exposes their ports, and the probes check/readyand/healthon them. The default listener ishttp-8899.gateway.tls.secretNamemounts a TLS Secret as a volume, so a rotated certificate serves new connections without a restart. - The Operator runs the Router standalone when
spec.gatewayis unset: the Router serves port 8801 itself, the port the Operator's Envoy sidecar served before. The sidecar is gone, and the Operator deletes its ConfigMap once the rollout completes.spec.gatewaystill selects the Gateway integration, with the Router serving ext_proc. - Keep a gateway in front. If Envoy Gateway, Agent Router, Istio, KServe,
llm-d or another Envoy-based gateway calls the Router over ext_proc, upgrade
with
--set gateway.mode=extproc; the integration values files underdeploy/kubernetes/set it. An upgrade whose live config still has the old default listeners (grpc-50051,http-8080) fails to render in standalone mode, before anything changes.helm rollbackrestores the previous release.
One router image
vllm-sr, vllm-sr-rocm and vllm-sr-cuda serve vllm-sr serve, the Helm
chart and the Operator, in either gateway mode, so Kubernetes gains CUDA. With
no arguments, or with Router flags, the image runs the Router on
/app/config/config.yaml, as the former extproc image did.
The Dashboard needs no container socket
vllm-sr serve in an empty directory still opens setup in the Dashboard, and
now keeps waiting: when you activate a config there, the CLI starts the Router
(and Envoy with --gateway extproc) from it. If you stop the command first, the
next vllm-sr serve starts the Router, and vllm-sr status says that setup is
complete. The Dashboard no longer starts, stops or execs containers, and no
container socket is mounted into it.
A change the running Router hot-reloads applies as before. A change the running
containers can't take is saved and waits for the CLI: one the Router answers
restart_required for (in standalone mode, a listener's set, address, port,
timeout or TLS certificate), or, with --gateway extproc, one that changes
Envoy's generated configuration. The Dashboard answers "Restart required: run
vllm-sr serve to apply.", vllm-sr status reports it, and the next
vllm-sr serve recreates the containers from the saved config.
Recipe activation follows the same rule. A Recipe that changes the listeners,
the managed storage or the Router management API used to make the Dashboard
recreate the containers itself. Now the activation, once confirmed, is
committed and answered 202 with status: restart_required and that message,
and the next vllm-sr serve creates the containers from the Recipe's config,
starting any storage it adds. Storage the Recipe stops using keeps its data and
runs until vllm-sr stop. Deactivation works the same way.
Engine mode runs in a container
vllm-sr serve MODEL runs the model runtime in the foreground, in a container
from the router image of --platform (vllm-sr, vllm-sr-rocm or
vllm-sr-cuda), so pip install vllm-sr and Docker or Podman are all it needs.
--hostand--portsay where the host publishes the runtime (default127.0.0.1:8100).- Local package directories are mounted read-only, and downloads persist in
~/.cache/vllm-sr/models(VLLM_SR_ENGINE_CACHE_DIRmoves it). --devicetakes what the image runs:cpu,rocm[:N]with--platform amd,cuda[:N]with--platform nvidia, or a plugin's accelerator in an image that has the plugin.--image,--image-pull-policy,--container-runtimeand--log-levelapply to engine mode too.--profilenames the kubernetes deployment profile only; engine mode's numerics profile is--runtime-profile.--udsis gone.
Looper calls its models from the Router
A Looper algorithm, a prompt helper and context recovery now make their model
calls inside the Router, in either gateway mode, instead of sending them back
through the gateway. Each call runs the decision's plugins and goes to the
model's backend_refs with the provider model's timeouts and retries. The
responses are the same, with one exception below.
- A failed model reads shorter. Where a Fusion or Router Flow response lists
a model that failed, its
errorisanswered 503(with the status),timed out,cancelled,invalid responseorfailed, without transport details. - A model the Router calls needs
backend_refs. A decision whose Looper algorithm, prompt helper or context recovery calls a model withoutproviders.models[].backend_refsfails to load, and the error names the decision and the model.vllm-sr validatereports the same error. If an external gateway owns the backends (listeners: []), point that model'sbackend_refsat the gateway's OpenAI-compatible address. global.integrations.looper.endpointis deprecated. The Router ignores it and logslooper_endpoint_deprecated, andvllm-sr config migrateremoves it. The next release drops the field.
Envoy mode changes too
These apply with --gateway extproc and behind a gateway integration:
- Retries. With
retry_countset, a retry prefers an endpoint it has not tried yet (Envoy'sprevious_hosts), as in standalone mode.retry_countwithoutretry_ontakes the default retry conditions instead of failing validation. - Per-decision overrides reach Envoy. The Router sends a decision's
timeouts and retries as Envoy's per-request headers
(
x-envoy-upstream-rq-timeout-ms,x-envoy-upstream-rq-per-try-timeout-ms,x-envoy-max-retries,x-envoy-retry-on,x-envoy-retriable-status-codes). Every Envoy configuration the CLI renders or the repository ships lets ext_proc set exactly those five headers (mutation_rules); a custom Envoy configuration needs the same rule for the overrides to apply. header_mutationcannot set those five headers. Loading the config drops such entries with a warning. Envoy ignored them before, so no working configuration changes.- Duplicate listener names are rejected, as Envoy already did.
- A fallback answer carries the decision headers. A response served by
cross-model fallback carries
x-vsr-selected-decision,x-vsr-selected-algorithm,x-vsr-selected-recipeandx-vsr-routing-latency-mslike any routed response, as in standalone mode. It used to name only the serving model (x-vsr-selected-model) and the attempt count (x-vsr-fallback-attempts).
Removed
- Fleet Simulator (
vllm-sr-sim): its package, its image, its PyPI and release jobs, its make targets and its documentation. Installed copies keep working; the v0.3 documentation still describes them. Upgrade and Rollback shows how to remove the sidecar container an earliervllm-sr servestarted. - OpenClaw: the agent integration is gone from the CLI, the Dashboard and
the Helm chart's documented variables. That covers the Dashboard's OpenClaw
page, the Playground's HireClaw mode and ClawRoom, the
/api/openclaw/*endpoints, theopenclaw.readandopenclaw.managepermissions, andvllm-sr config import --from openclaw. For this release the Dashboard still accepts its-openclaw*flags and ignores them, like theOPENCLAW_*variables, and logs oneDEPRECATEDline naming those it was given. Remove them from custom manifests: the next release no longer accepts the flags, and an unknown flag stops the Dashboard at startup.vllm-sr serveno longer mounts the container runtime socket into the Dashboard.
Renamed
| Before | Now | The old name |
|---|---|---|
--target k8s | --target kubernetes | works for this release, with a warning |
--runtime docker|podman | --container-runtime docker|podman | works for this release, with a warning |
image extproc | image vllm-sr | published with the same digests for this release |
image extproc-rocm | image vllm-sr-rocm | published with the same digests for this release |
make docker-build-extproc | make docker-build-vllm-sr (-rocm, -cuda) | removed |