Skip to main content
Version: Latest

Release note: standalone mode is the default

The first release that includes #4623 serves the OpenAI-compatible API from the Router itself. vllm-sr serve starts no Envoy container by default: the Router container publishes the config's listeners, answers /health and /ready on them, and proxies to the model backends. The routing decisions, the upstream requests and the responses are the same as behind Envoy. See Gateway Modes.

Keep the previous stack​

vllm-sr serve --gateway extproc

--gateway extproc starts the Envoy container in front of the Router exactly as earlier releases did, with the same Envoy configuration. Choose it for Envoy features the standalone Router does not have yet, such as token-bucket rate limiting, mTLS, JWT or OIDC, or advanced route matching. Envoy mode also sustains about 5–10% more requests per second at 32 or more concurrent clients; below that, standalone mode answers faster (see the design doc's results).

What changes for a standalone stack​

  • No Envoy container. vllm-sr status and vllm-sr logs report the Router and the Dashboard; vllm-sr logs envoy and --envoy-image need --gateway extproc. The Dashboard's Playground and readiness checks use the Router's listener.
  • Proxy-control headers from clients are dropped. x-envoy-internal and Envoy's retry, timeout and tracing controls (such as x-envoy-max-retries and x-envoy-upstream-rq-timeout-ms) never reach a backend, as behind Envoy.
  • Identity headers from clients are dropped, both x-authz-* and the names global.services.authz.identity sets, because no authenticator stands in front of the Router. If a proxy or your application sets them, turn on listeners[].identity.trust_headers (optionally with trusted_peers). Without a trusting listener, a config with a decision on an authz signal, a rate limit rule that matches user or group, or authz providers is refused at startup, and memory, router replay and the per-user selection algorithms treat every request as anonymous.
  • Listener changes need a restart. A reload that changes a listener's address, port, timeout or tls paths, or adds or removes a listener, is rejected as restart_required; run vllm-sr serve again.

New​

  • listeners[].tls (cert_file, key_file) serves a listener over TLS in standalone mode, and a renewed certificate in those files serves new connections without a restart. --gateway extproc refuses it rather than serve the listener in cleartext.
  • --gateway standalone|extproc and --platform cpu|amd|nvidia work on the kubernetes target too. The CLI writes them into the generated Helm values as gateway.mode, the image repository and a GPU request.
  • vllm-sr serve --help lists its options by group, and an option of another group is an error that says where it applies.
  • The Router binary takes -gateway standalone. Its default stays extproc, so a manifest that runs it without the flag behaves as before.
  • On macOS, --platform amd|nvidia fails with a clear message: Docker's Linux VM gets no GPU there, so the docker target runs the CPU image.
  • Timeouts, retries and fallback, per model and per decision. providers.models[].reliability gains connect, total, idle, per-try and first-byte timeouts, retriable status codes, back-off, Retry-After and retry budgets. A decision's reliability and fallback blocks override them for the requests it routes, in both gateway modes; first_byte_timeout is standalone only. See Tune timeouts, retries, and endpoint health and Fall back to another model.
  • Configuration versions and rollback. Every accepted change activates a numbered version, and each response names it in x-vsr-config-version. A rejected change leaves the active version serving and reports why. The last ten activations survive a restart, and POST /api/v1/config/rollback activates a recorded version as a new one. See Configuration Management.

Kubernetes runs standalone too​

  • The Helm chart sets gateway.mode: standalone by default. The Router serves the listeners in its config, the Service exposes their ports, and the probes check /ready and /health on them. The default listener is http-8899. gateway.tls.secretName mounts a TLS Secret as a volume, so a rotated certificate serves new connections without a restart.
  • The Operator runs the Router standalone when spec.gateway is unset: the Router serves port 8801 itself, the port the Operator's Envoy sidecar served before. The sidecar is gone, and the Operator deletes its ConfigMap once the rollout completes. spec.gateway still selects the Gateway integration, with the Router serving ext_proc.
  • Keep a gateway in front. If Envoy Gateway, Agent Router, Istio, KServe, llm-d or another Envoy-based gateway calls the Router over ext_proc, upgrade with --set gateway.mode=extproc; the integration values files under deploy/kubernetes/ set it. An upgrade whose live config still has the old default listeners (grpc-50051, http-8080) fails to render in standalone mode, before anything changes. helm rollback restores the previous release.

One router image​

vllm-sr, vllm-sr-rocm and vllm-sr-cuda serve vllm-sr serve, the Helm chart and the Operator, in either gateway mode, so Kubernetes gains CUDA. With no arguments, or with Router flags, the image runs the Router on /app/config/config.yaml, as the former extproc image did.

The Dashboard needs no container socket​

vllm-sr serve in an empty directory still opens setup in the Dashboard, and now keeps waiting: when you activate a config there, the CLI starts the Router (and Envoy with --gateway extproc) from it. If you stop the command first, the next vllm-sr serve starts the Router, and vllm-sr status says that setup is complete. The Dashboard no longer starts, stops or execs containers, and no container socket is mounted into it.

A change the running Router hot-reloads applies as before. A change the running containers can't take is saved and waits for the CLI: one the Router answers restart_required for (in standalone mode, a listener's set, address, port, timeout or TLS certificate), or, with --gateway extproc, one that changes Envoy's generated configuration. The Dashboard answers "Restart required: run vllm-sr serve to apply.", vllm-sr status reports it, and the next vllm-sr serve recreates the containers from the saved config.

Recipe activation follows the same rule. A Recipe that changes the listeners, the managed storage or the Router management API used to make the Dashboard recreate the containers itself. Now the activation, once confirmed, is committed and answered 202 with status: restart_required and that message, and the next vllm-sr serve creates the containers from the Recipe's config, starting any storage it adds. Storage the Recipe stops using keeps its data and runs until vllm-sr stop. Deactivation works the same way.

Engine mode runs in a container​

vllm-sr serve MODEL runs the model runtime in the foreground, in a container from the router image of --platform (vllm-sr, vllm-sr-rocm or vllm-sr-cuda), so pip install vllm-sr and Docker or Podman are all it needs.

  • --host and --port say where the host publishes the runtime (default 127.0.0.1:8100).
  • Local package directories are mounted read-only, and downloads persist in ~/.cache/vllm-sr/models (VLLM_SR_ENGINE_CACHE_DIR moves it).
  • --device takes what the image runs: cpu, rocm[:N] with --platform amd, cuda[:N] with --platform nvidia, or a plugin's accelerator in an image that has the plugin.
  • --image, --image-pull-policy, --container-runtime and --log-level apply to engine mode too.
  • --profile names the kubernetes deployment profile only; engine mode's numerics profile is --runtime-profile. --uds is gone.

Looper calls its models from the Router​

A Looper algorithm, a prompt helper and context recovery now make their model calls inside the Router, in either gateway mode, instead of sending them back through the gateway. Each call runs the decision's plugins and goes to the model's backend_refs with the provider model's timeouts and retries. The responses are the same, with one exception below.

  • A failed model reads shorter. Where a Fusion or Router Flow response lists a model that failed, its error is answered 503 (with the status), timed out, cancelled, invalid response or failed, without transport details.
  • A model the Router calls needs backend_refs. A decision whose Looper algorithm, prompt helper or context recovery calls a model without providers.models[].backend_refs fails to load, and the error names the decision and the model. vllm-sr validate reports the same error. If an external gateway owns the backends (listeners: []), point that model's backend_refs at the gateway's OpenAI-compatible address.
  • global.integrations.looper.endpoint is deprecated. The Router ignores it and logs looper_endpoint_deprecated, and vllm-sr config migrate removes it. The next release drops the field.

Envoy mode changes too​

These apply with --gateway extproc and behind a gateway integration:

  • Retries. With retry_count set, a retry prefers an endpoint it has not tried yet (Envoy's previous_hosts), as in standalone mode. retry_count without retry_on takes the default retry conditions instead of failing validation.
  • Per-decision overrides reach Envoy. The Router sends a decision's timeouts and retries as Envoy's per-request headers (x-envoy-upstream-rq-timeout-ms, x-envoy-upstream-rq-per-try-timeout-ms, x-envoy-max-retries, x-envoy-retry-on, x-envoy-retriable-status-codes). Every Envoy configuration the CLI renders or the repository ships lets ext_proc set exactly those five headers (mutation_rules); a custom Envoy configuration needs the same rule for the overrides to apply.
  • header_mutation cannot set those five headers. Loading the config drops such entries with a warning. Envoy ignored them before, so no working configuration changes.
  • Duplicate listener names are rejected, as Envoy already did.
  • A fallback answer carries the decision headers. A response served by cross-model fallback carries x-vsr-selected-decision, x-vsr-selected-algorithm, x-vsr-selected-recipe and x-vsr-routing-latency-ms like any routed response, as in standalone mode. It used to name only the serving model (x-vsr-selected-model) and the attempt count (x-vsr-fallback-attempts).

Removed​

  • Fleet Simulator (vllm-sr-sim): its package, its image, its PyPI and release jobs, its make targets and its documentation. Installed copies keep working; the v0.3 documentation still describes them. Upgrade and Rollback shows how to remove the sidecar container an earlier vllm-sr serve started.
  • OpenClaw: the agent integration is gone from the CLI, the Dashboard and the Helm chart's documented variables. That covers the Dashboard's OpenClaw page, the Playground's HireClaw mode and ClawRoom, the /api/openclaw/* endpoints, the openclaw.read and openclaw.manage permissions, and vllm-sr config import --from openclaw. For this release the Dashboard still accepts its -openclaw* flags and ignores them, like the OPENCLAW_* variables, and logs one DEPRECATED line naming those it was given. Remove them from custom manifests: the next release no longer accepts the flags, and an unknown flag stops the Dashboard at startup. vllm-sr serve no longer mounts the container runtime socket into the Dashboard.

Renamed​

BeforeNowThe old name
--target k8s--target kubernetesworks for this release, with a warning
--runtime docker|podman--container-runtime docker|podmanworks for this release, with a warning
image extprocimage vllm-srpublished with the same digests for this release
image extproc-rocmimage vllm-sr-rocmpublished with the same digests for this release
make docker-build-extprocmake docker-build-vllm-sr (-rocm, -cuda)removed