Gateway Modes
Gateway mode chooses where client traffic enters; the target chooses where
the stack runs. These are separate from --engine, which serves native
judgment requests without Chat routing. Omitting --engine starts Router
mode. See component architecture for
the shared frontend and model runtime.
| Mode | Client traffic enters at | Use it for |
|---|---|---|
standalone (default) | the Router, which serves the OpenAI-compatible API on the config's listeners | one host, development, edge, most self-hosting |
extproc | an Envoy-based gateway in front of the Router, which serves ext_proc | Envoy features such as rate limiting, mTLS and advanced route matching, or a gateway you already run |
| Target | standalone | extproc |
|---|---|---|
docker (default) | the Router container serves the listeners; there is no Envoy container | the Envoy container the CLI starts, in front of the Router, as in earlier releases |
kubernetes | the Router pods serve the listeners, and the Service exposes them | the Router serves ext_proc for your Envoy Gateway, Envoy AI Gateway, Istio or KServe |
vllm-sr serve # standalone, docker
vllm-sr serve --gateway extproc # Envoy in front, docker
vllm-sr serve --target kubernetes --config config.yaml
Start with standalone unless you need one of the Envoy features above. On
docker, switching is a restart on the active configuration: run
vllm-sr serve --gateway extproc, and plain vllm-sr serve to come back. In
standalone mode there is no Envoy container, so vllm-sr logs envoy and the
x-envoy-* response headers belong to extproc only.
Both modes run the same routing core, so a request gets the same decision, the
same upstream request and the same response transformations in either. The
design lists the few
deliberate differences, such as the x-envoy-* headers that only Envoy adds.
Earlier releases always put Envoy in front of the Router. Standalone is now the
default; vllm-sr serve --gateway extproc restores Envoy ingress.
See the release note.
What a standalone Router does
-
Listeners: each entry of
listenersis served with HTTP/1.1 and HTTP/2 (h2c in cleartext), with itstimeoutas the idle timeout, a 500 MiB request body limit and at most 50,000 connections, the limits of the Envoy template. -
API keys: with
api_keysset, a client sends one of them asAuthorization: Bearer <key>orapi-key: <key>; other requests get an OpenAI-style 401. The key is removed before the request reaches a provider. -
Model allow-list: with
modelsset, the listener accepts only those request models (see Model allow-list). -
TLS:
tlsserves the listener over TLS 1.2 or later, with HTTP/2 or HTTP/1.1 negotiated by ALPN. Relative paths are relative to the config file's directory. The Router reloads the key pair when its files change, so a renewed certificate (a rotated Kubernetes secret, cert-manager) serves new connections without a restart; a pair that fails to load leaves the previous one serving.--gateway extprocdoes not serve it.listeners:- name: https-8443address: 0.0.0.0port: 8443tls:cert_file: certs/tls.crtkey_file: certs/tls.key -
The edge's trust boundary: client-sent identity headers (
x-authz-*, and the namesglobal.services.authz.identitysets), unless the listener trusts them (see Identity headers), and the proxy-control headers only a trusted proxy may set never reach routing or a backend. Those arex-envoy-internaland Envoy's retry, timeout and tracing controls, such asx-envoy-max-retries,x-envoy-retry-onandx-envoy-upstream-rq-timeout-ms. Envoy-based layers behind the Router (sidecars, gateways in front of model servers) obey them from a caller they trust, which the Router is, so a client could otherwise set retries and timeouts there; the Router's reliability policy stays the one retry and timeout authority. The list is the one Envoy strips from external requests, and it stays that narrow. -
Probes:
GET /healthanswers while the process runs, andGET /readyonce the routing core can take traffic. Prometheus metrics stay on the Router's metrics port (9190). -
Reloads: the Router reloads its config in place. A change to a listener's address, port, timeout or
tlspaths, or a new or removed listener, is rejected asrestart_requireduntil the Router restarts; API keys and model allow-lists reload in place.
Model allow-list
A listener's models lists the only request model values it accepts. Use it
to give a public key access to the router's auto model and nothing else, while
an internal listener keeps every model:
listeners:
- name: dashboard-internal # first listener: the Dashboard Playground uses it
address: 127.0.0.1
port: 8898
- name: public
address: 0.0.0.0
port: 8899
api_keys: ["${WORKSHOP_KEY}"]
models: [vllm-sr/auto]
- Names match exactly (case-sensitive, after trimming the request value), and aliases are not expanded: list every name clients may send. Empty or absent, the listener accepts every model.
- The check runs after the API key check and before any signal, cache or
decision, on the model the Router parsed for routing. Any other model,
including a provider model that would otherwise pass through, gets
403with{"error": {"code": "model_not_allowed", ...}}in the client's protocol. A request without a model gets400 model_required. GET /v1/modelson the listener lists only the allowed names the catalog has.- The model calls a decision makes in process (Looper and request-graph hops)
are not client requests and are not restricted, so
vllm-sr/autocan still reach every provider model its decisions name. - The listener ignores the
x-vsr-skip-processingopt-out even whenglobal.router.skip_processing.enabledis on, because a skipped request would bypass the check. --gateway extprocrejects a listener withmodelsas unsupported: the Envoy listener the CLI generates does not enforce it yet.
Identity headers
By default a standalone listener drops the identity headers a client sends, because nothing in front of the Router has authenticated it: anyone could claim any user. Behind an authenticating proxy, or for a trusted application server that sets the user itself, let the listener keep them:
listeners:
- name: http-8899
address: 0.0.0.0
port: 8899
identity:
trust_headers: true
# Optional: keep them only on connections from the proxy's network.
trusted_peers: ["10.0.0.0/8"]
trust_headerskeeps thex-authz-*headers and the namesglobal.services.authz.identitysets.trusted_peers(CIDRs) keeps them only when the connection's peer address is in one of the networks; the Router never readsX-Forwarded-Forfor this. Empty, every peer of a trusting listener is trusted.- Each listener decides for its own requests, and a reload applies a change. Requests on a listener that doesn't trust identity are anonymous.
Memory, router replay and the per-user selection algorithms (gmtrouter,
rl_driven) record the user of each request. Without a trusting listener they
still load and treat every request as anonymous, and the Router logs one
warning at startup that names them.
What needs extproc
Policy that enforces access by a client identity needs an identity source. The
Router refuses it at startup and on reload when no listener sets
identity.trust_headers, and the error names that option and
--gateway extproc:
- a decision on an
authz(role binding) signal; - a rate limit rule that matches
userorgroup; global.services.authz.providers, which resolve per-user API keys.
Turn on identity.trust_headers on a listener behind an authenticating proxy,
or serve them with --gateway extproc behind a gateway that authenticates
clients. Token-bucket rate limiting, mTLS, JWT or OIDC, and advanced route
matching are planned for standalone mode; until then they need Envoy as well.
Targets and platforms
--platform auto detects the execution backend on the selected target.
Choose cpu, rocm or cuda explicitly to use vllm-sr, vllm-sr-rocm
or vllm-sr-cuda, respectively.
- docker:
rocmpasses the ROCm devices through, andcudathe NVIDIA GPUs. Use--device-idsto select host GPUs for the default model deployment. - kubernetes: the generated Helm values set
gateway.mode, the image repository for the platform, and GPU resources (amd.com/gpuornvidia.com/gpu) for the configured placements. The cluster needs the matching device plugin. Use allocation ordinals such asrocm:0in YAML; host--device-idsapplies only to Docker.--target k8sis the old name of--target kubernetesand works for this release only.
Without the CLI, the Helm chart takes the same value, gateway.mode (default
standalone), and the Operator runs the Router standalone unless
spec.gateway names a Gateway, which selects extproc. To move a release that
an Envoy-based gateway calls, see
Upgrade and Rollback.
macOS
On macOS the docker target is CPU only: the built-in models run on the CPU in
the arm64 image, because Apple's virtualization gives Docker's Linux VM no Metal
or GPU compute. --platform rocm and --platform cuda fail there with a clear
message. GPU support through the host is tracked in
#4636.
- Comfortable on the CPU: the 307M Vela task models (domain, PII, jailbreak and safety guards, embeddings, reranking), Vela Omni Nano, and the Decision 2.0 Kai-0.6B and Eos-0.8B decision models.
- Larger models: a model needs about 4 bytes per parameter on the CPU, so a 2B decision model needs about 8 GB. Raise the memory of Docker's VM (Docker Desktop: Settings, Resources) above what the config's models need, or pick a smaller model. See Choose a model.
Options of vllm-sr serve
vllm-sr serve --help groups options by deployment target. Router and Engine
run the same instance frontend: --engine (-e) disables routing at startup;
every invocation without it starts Router mode and retains saved routing. Both retain model
management, Dashboard and the explicitly published System One API.
| Group | Applies to | Options |
|---|---|---|
| Instance and models | docker, kubernetes | MODEL, --engine, --revision, --data-parallel-size, --runtime-profile, --platform, --image, --log-level |
| Instance configuration | docker, kubernetes | --config, --target, --gateway, --minimal, --readonly, --algorithm |
| Container options | docker | --image-pull-policy, --container-runtime |
| Docker target | docker | --router-image, --envoy-image (with --gateway extproc), --dashboard-image, --startup-timeout, --replace-active-config, --recipe-env, --device-ids |
| Kubernetes target | kubernetes | --namespace, --context, --profile, --chart-dir |
Multiple models, replica placement, listener ports and API grants use the canonical config. Model resource flags update the configured default deployment; MODEL and all resource overrides are optional.
--container-runtime (docker or podman) replaces --runtime, which works
for this release only, on serve, status, logs, stop and dashboard.