Choose a Deployment
Choose two things independently:
- where Semantic Router runs; and
- where the model backends run.
Chat provider models run behind reachable endpoints on the same host, in a cluster, or at a hosted API. vLLM-SR does not provision those generation backends. Its own decision, classifier, embedding, and reranking models use managed or attached model-runtime deployments. For decision serving without Chat backends, start in Engine mode.
Choose the Router topology
| Need | Recommended path | Start here |
|---|---|---|
| Evaluate locally or run on one host | Docker stack managed by the CLI | Deploy with Docker |
| Deploy a complete canonical config through GitOps | Helm | CLI and Helm workflow |
| Let Kubernetes reconcile Router resources and discovery | Kubernetes Operator | Kubernetes Operator |
| Attach routing policy to an existing gateway | Gateway integration | Gateways |
| Let another platform own model replicas and scheduling | Inference-platform integration | Inference Platforms |
Gateway and inference-platform integrations do not replace Router policy. They connect semantic model selection to infrastructure that owns traffic or model lifecycle.
On Docker and Kubernetes alike, the Router serves clients itself by default. Gateway Modes explains when to put an Envoy-based gateway in front of it instead.
Before committing to a path, check its project-maintained status, recurring test evidence, and external ownership boundary in the Deployment Support.
Choose the model backend
| Backend situation | Start here |
|---|---|
| You already have a reachable model or provider endpoint | Protocol Compatibility, then Backend Target Compatibility |
| You want a small local model for evaluation | Local model with Ollama |
| You want to serve models on AMD Instinct | AMD ROCm |
| You want to serve models or accelerate Router-side models on NVIDIA | NVIDIA CUDA |
| A Kubernetes platform owns model deployment and replicas | Inference Platforms |
Hardware is an overlay, not a separate Router topology. A GPU-backed model server can connect to either a Docker or Kubernetes Router deployment. Keep the Router on CPU unless measurements show that its local embeddings or classifiers benefit from GPU acceleration.
Test a model endpoint directly before sending traffic through the Router. A configured URL is not proof that the backend implements the selected wire protocol or supports the recipe's context, modality, and tool requirements.
Before production
Before exposing a deployment:
- pin the Router, model server, model, and integration versions together;
- validate buffered, streaming, failure, and rollback behavior through the actual data plane;
- move credentials into a secret manager; and
- review Configuration, Security Hardening, Data and Storage, and Upgrade and Rollback.