Quickstart
Install vLLM Semantic Router, connect one model, and send your first routed request. To serve typed decision questions without a Chat backend, follow the System One quickstart instead. Both paths use the same frontend and model runtime.
vllm-sr serve runs a local stack in Docker: the Router, which serves an
OpenAI-compatible API on port 8899 itself (standalone mode, the default), the
Dashboard on port 8700, and the
model runtime that runs the Router's own
classifiers and embedding models. The models that answer your users run
elsewhere, such as Ollama, a vLLM server or a hosted API, and you connect them
in the Dashboard.
These pages follow main. Standalone mode and the built-in model runtime came
after vllm-sr 0.4.0, the current stable release, which puts Envoy in front
of the Router and has no --gateway option. To follow these pages today,
use the development channel. The installation tabs below select it: the curl
installer uses --channel dev, and pip/uv allow prereleases. Stable-release
upgrade instructions are separate from this main quickstart.
Requirements
| Host | You need |
|---|---|
| Every host | Linux, macOS or WSL2; Docker (Linux can use Podman instead); Python 3.10 or newer; about 5 GB of free disk for the images, plus 1–1.5 GB for each built-in model your routes use |
| CPU | Nothing else. The vllm-sr image runs every built-in model on the CPU; plan about 1.3 GB of memory for each 307M task model your routes use |
| AMD Instinct MI300X or MI325X | The ROCm driver on the host, Docker access to /dev/kfd and /dev/dri, and --platform rocm. Its vllm-sr-rocm image is a 6.5 GB download and takes about 20 GB of disk |
| NVIDIA | The NVIDIA Container Toolkit and --platform cuda (works, not yet validated) |
On macOS the Docker target runs on the CPU only; see Gateway Modes.
For the curl installer, pass --runtime podman to force Podman or
--runtime skip to skip container-runtime preparation. For example:
curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev --runtime skip
These are installer options. vllm-sr serve --container-runtime selects a
container runtime (docker or podman); skip is not a serve runtime.
Install
- curl
- pip
- uv
- Agent
curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade --pre vllm-sr
uv tool install --upgrade --prerelease allow vllm-sr
Copy this prompt into your coding agent:
Install the development channel of vLLM Semantic Router, then configure and verify it on this machine by following the official skill: https://vllm-sr.ai/install/agent/vllm-sr/SKILL.md
The prompt points to the public, self-contained vLLM SR agent skill. See Install with an agent for the workflow and safety boundaries.
Verify the CLI:
vllm-sr --version
Start the stack
The curl installer starts the stack for you. After a pip or uv install, start it yourself:
vllm-sr serve # Auto-detect the execution target
vllm-sr serve --platform rocm # AMD GPUs
The first start pulls the images, which takes a few minutes. With no
config.yaml in the current directory, the stack starts in setup mode: the
Dashboard runs, and the Router waits for a configuration.
Set up in the Dashboard
Open http://localhost:8700. The Dashboard listens on
127.0.0.1; on a remote host, run ssh -L 8700:127.0.0.1:8700 <host> first.
- Create the first administrator: a name, an email and a password.
- Connect a model: its name, its provider (vLLM, Ollama, OpenAI
Compatible or Anthropic) and its address as the Router container sees it,
for example
host.docker.internal:11434for Ollama on the same host. For a first local model, follow Configure models with Ollama. - Choose routing: From scratch makes one default route to that model.
- Review and Activate.
vllm-sr serve keeps waiting during setup. After you activate, it starts the
Router from that configuration, prints the endpoints and exits. If you stopped
it first, run vllm-sr serve again; until then vllm-sr status says that
setup is complete. Agents can do the same work through the CLI and the Router
management API without using the Dashboard.
Send a request
curl -s -D - http://localhost:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'x-vsr-debug: true' \
-d '{
"model": "vllm-sr/auto",
"messages": [{"role": "user", "content": "Hello!"}]
}'
vllm-sr/auto lets the Router choose. The response headers say what it chose:
x-vsr-selected-decision names the route and x-vsr-selected-model the
model; with x-vsr-debug: true, the x-vsr-matched-* headers list the
signals that matched. See Router headers.
Operate the stack
vllm-sr status # what runs, and whether setup or a restart is pending
vllm-sr logs router -f # Router logs
vllm-sr stop
Later changes the Router hot-reloads apply at once. A change the running
containers can't take, such as a listener's new port, is saved and the
Dashboard answers "Restart required: run vllm-sr serve to apply."; the next
vllm-sr serve applies it, and vllm-sr status reports it until then. See
Configuration Management.
vllm-sr serve --gateway extproc puts an Envoy container in front of the
Router, as releases before standalone mode did.
Gateway Modes explains when you need it.
What the installer leaves behind
install.sh writes these paths (defaults - --install-root and --bin-dir
move them):
~/.local/share/vllm-sr/venv/: the CLI's virtual environment.~/.local/share/vllm-sr/runtime.env: the container runtime chosen at install time.vllm-srreads it when it selects the runtime;CONTAINER_RUNTIMEoverrides it for one run, and a runtime missing fromPATHfalls back to auto-detection with a warning. Delete the file to return to auto-detection.~/.local/share/vllm-sr/first-run.log: kept only when the firstvllm-sr servewaits for setup or fails.~/.local/bin/vllm-sr: the launcher, which pinsVLLM_SR_INSTALL_ROOT.- A completion line in your shell rc file:
vllm-sr completion installappends an eval line to~/.bashrcor~/.zshrc, or writes~/.config/fish/completions/vllm-sr.fishon fish. Once the launcher is gone, every new bash or zsh shell reports the missing command, so remove the line - or the fish completions file - too. Shell Completion has the details.
Run vllm-sr stop while the CLI is still installed: it removes the stack's
containers and networks. Then delete the paths above.
Deleting those paths removes the install, but leaves in place:
- system packages the installer added when missing: Python and its venv
support; in serve mode, Homebrew
dockerandcolimaon macOS, or the Docker package, its enabled service and yourdockergroup membership on Linux; - the images
vllm-sr servepulled, and the Redis and Postgres data volumes the stack wrote; - the state the first
vllm-sr servewrote in the directory it ran from:config.yaml,.vllm-sr/(Dashboard data, recipe store, logs) andmodels/. The installer starts that first serve from the directory the installer ran in, falling back to$HOMEwhen it is not writable.