Skip to main content
Version: Latest (unreleased)

Quickstart

Install vLLM Semantic Router, connect one model, and send your first routed request. To serve typed decision questions without a Chat backend, follow the System One quickstart instead. Both paths use the same frontend and model runtime.

vllm-sr serve runs a local stack in Docker: the Router, which serves an OpenAI-compatible API on port 8899 itself (standalone mode, the default), the Dashboard on port 8700, and the model runtime that runs the Router's own classifiers and embedding models. The models that answer your users run elsewhere, such as Ollama, a vLLM server or a hosted API, and you connect them in the Dashboard.

Release channel

These pages follow main. Standalone mode and the built-in model runtime came after vllm-sr 0.4.0, the current stable release, which puts Envoy in front of the Router and has no --gateway option. To follow these pages today, use the development channel. The installation tabs below select it: the curl installer uses --channel dev, and pip/uv allow prereleases. Stable-release upgrade instructions are separate from this main quickstart.

Requirements​

HostYou need
Every hostLinux, macOS or WSL2; Docker (Linux can use Podman instead); Python 3.10 or newer; about 5 GB of free disk for the images, plus 1–1.5 GB for each built-in model your routes use
CPUNothing else. The vllm-sr image runs every built-in model on the CPU; plan about 1.3 GB of memory for each 307M task model your routes use
AMD Instinct MI300X or MI325XThe ROCm driver on the host, Docker access to /dev/kfd and /dev/dri, and --platform rocm. Its vllm-sr-rocm image is a 6.5 GB download and takes about 20 GB of disk
NVIDIAThe NVIDIA Container Toolkit and --platform cuda (works, not yet validated)

On macOS the Docker target runs on the CPU only; see Gateway Modes.

For the curl installer, pass --runtime podman to force Podman or --runtime skip to skip container-runtime preparation. For example:

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev --runtime skip

These are installer options. vllm-sr serve --container-runtime selects a container runtime (docker or podman); skip is not a serve runtime.

Install​

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel dev

Verify the CLI:

vllm-sr --version

Start the stack​

The curl installer starts the stack for you. After a pip or uv install, start it yourself:

vllm-sr serve # Auto-detect the execution target
vllm-sr serve --platform rocm # AMD GPUs

The first start pulls the images, which takes a few minutes. With no config.yaml in the current directory, the stack starts in setup mode: the Dashboard runs, and the Router waits for a configuration.

Set up in the Dashboard​

Open http://localhost:8700. The Dashboard listens on 127.0.0.1; on a remote host, run ssh -L 8700:127.0.0.1:8700 <host> first.

  1. Create the first administrator: a name, an email and a password.
  2. Connect a model: its name, its provider (vLLM, Ollama, OpenAI Compatible or Anthropic) and its address as the Router container sees it, for example host.docker.internal:11434 for Ollama on the same host. For a first local model, follow Configure models with Ollama.
  3. Choose routing: From scratch makes one default route to that model.
  4. Review and Activate.

vllm-sr serve keeps waiting during setup. After you activate, it starts the Router from that configuration, prints the endpoints and exits. If you stopped it first, run vllm-sr serve again; until then vllm-sr status says that setup is complete. Agents can do the same work through the CLI and the Router management API without using the Dashboard.

Send a request​

curl -s -D - http://localhost:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'x-vsr-debug: true' \
-d '{
"model": "vllm-sr/auto",
"messages": [{"role": "user", "content": "Hello!"}]
}'

vllm-sr/auto lets the Router choose. The response headers say what it chose: x-vsr-selected-decision names the route and x-vsr-selected-model the model; with x-vsr-debug: true, the x-vsr-matched-* headers list the signals that matched. See Router headers.

Operate the stack​

vllm-sr status # what runs, and whether setup or a restart is pending
vllm-sr logs router -f # Router logs
vllm-sr stop

Later changes the Router hot-reloads apply at once. A change the running containers can't take, such as a listener's new port, is saved and the Dashboard answers "Restart required: run vllm-sr serve to apply."; the next vllm-sr serve applies it, and vllm-sr status reports it until then. See Configuration Management.

vllm-sr serve --gateway extproc puts an Envoy container in front of the Router, as releases before standalone mode did. Gateway Modes explains when you need it.

What the installer leaves behind​

install.sh writes these paths (defaults - --install-root and --bin-dir move them):

  • ~/.local/share/vllm-sr/venv/: the CLI's virtual environment.
  • ~/.local/share/vllm-sr/runtime.env: the container runtime chosen at install time. vllm-sr reads it when it selects the runtime; CONTAINER_RUNTIME overrides it for one run, and a runtime missing from PATH falls back to auto-detection with a warning. Delete the file to return to auto-detection.
  • ~/.local/share/vllm-sr/first-run.log: kept only when the first vllm-sr serve waits for setup or fails.
  • ~/.local/bin/vllm-sr: the launcher, which pins VLLM_SR_INSTALL_ROOT.
  • A completion line in your shell rc file: vllm-sr completion install appends an eval line to ~/.bashrc or ~/.zshrc, or writes ~/.config/fish/completions/vllm-sr.fish on fish. Once the launcher is gone, every new bash or zsh shell reports the missing command, so remove the line - or the fish completions file - too. Shell Completion has the details.

Run vllm-sr stop while the CLI is still installed: it removes the stack's containers and networks. Then delete the paths above.

Deleting those paths removes the install, but leaves in place:

  • system packages the installer added when missing: Python and its venv support; in serve mode, Homebrew docker and colima on macOS, or the Docker package, its enabled service and your docker group membership on Linux;
  • the images vllm-sr serve pulled, and the Redis and Postgres data volumes the stack wrote;
  • the state the first vllm-sr serve wrote in the directory it ran from: config.yaml, .vllm-sr/ (Dashboard data, recipe store, logs) and models/. The installer starts that first serve from the directory the installer ran in, falling back to $HOME when it is not writable.

Next​