Deploy with NVIDIA CUDA
The model server and Semantic Router are separate services. A common deployment keeps the Router on CPU and gives the NVIDIA GPU to vLLM. Use the Router's CUDA image when its local embeddings or classifiers also need GPU acceleration.
--platform nvidia affects the local Router stack only. It selects the CUDA
Router image, passes NVIDIA GPUs into the Router container, and changes its
generated runtime configuration so supported local signal models prefer CUDA.
It does not download a language model or start a vLLM server.
Prerequisites
- Linux and an NVIDIA GPU supported by the vLLM release you plan to run;
- an x86-64 host when using the current Semantic Router CUDA image;
- a GPU of compute capability 7.0 or newer (Volta and later) for the Router
image, which compiles its local models for that minimum; build with
CUDA_COMPUTE_CAP=<value>to target an older or newer floor; - an NVIDIA driver, and GPU passthrough into the Router container. The CUDA Router image links the driver library directly, so it does not start without passthrough even when every Router-side model is configured for CPU;
- Docker and NVIDIA Container Toolkit;
- enough GPU memory for the vLLM model, KV cache, and any Router-side models; and
- a complete Semantic Router configuration with a reachable model endpoint.
Use the current vLLM NVIDIA requirements for supported hardware. Install and configure the runtime with the NVIDIA Container Toolkit guide, then verify both the host driver and container access:
nvidia-smi
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
Do not continue until the container command can see the expected GPUs. The second command is NVIDIA's sample workload.
Start and verify a vLLM backend
The following example follows the official
vLLM Docker deployment.
It publishes an OpenAI-compatible endpoint on port 8000 and keeps downloaded
model files in a named volume:
docker volume create vllm-huggingface-cache
docker run -d \
--name vllm-nvidia \
--runtime nvidia \
--gpus all \
--ipc=host \
-p 8000:8000 \
-v vllm-huggingface-cache:/root/.cache/huggingface \
vllm/vllm-openai:latest \
--model Qwen/Qwen3-0.6B
Choose a model and vLLM arguments that fit the available GPUs. Pass
HF_TOKEN as an environment variable when a model requires authentication;
do not put the token in an image, command history, or Router config. Pin the
vLLM image and model revision for a controlled deployment instead of relying
on latest.
Wait for model loading to finish, then test vLLM before adding the Router:
curl --fail http://127.0.0.1:8000/health
curl --fail http://127.0.0.1:8000/v1/models
curl --fail http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [{"role": "user", "content": "Reply with: ready"}],
"max_tokens": 16
}'
Router validation cannot prove that a backend can load a model or generate a response, so fix any direct vLLM error before continuing.
Connect the backend
Bind the served model in your canonical config. For the local Docker stack,
the Router can reach a host-published port through host.docker.internal:
providers:
defaults:
model: local/qwen
models:
- name: local/qwen
provider_model_id: Qwen/Qwen3-0.6B
api_format: openai
backend_refs:
- name: nvidia-vllm
endpoint: host.docker.internal:8000
protocol: http
provider: vllm
weight: 1
This is a provider fragment, not a complete Router config. Add the matching
model card and route to your existing recipe, or configure the endpoint in the
Dashboard. The provider_model_id must match a model returned by vLLM's
/v1/models endpoint. See Configuration for a complete
minimal document and
Models, Entrypoints, and Serving
for backend binding and routing policy.
The example publishes port 8000 on the host for direct testing. Restrict that
port with host networking controls, or use private service discovery in a
production deployment. Do not expose an unauthenticated vLLM endpoint to an
untrusted network.
Run the Router on NVIDIA
If vLLM should own all GPU memory, keep the Router on CPU:
vllm-sr config validate --config config.yaml
vllm-sr serve --config config.yaml
To run supported Router-side local embeddings and classifiers on CUDA, use
--platform nvidia. A stable CLI selects the matching published release image
(for example, CLI 0.4.0 uses vllm-sr-cuda:v0.4.0). Development CLI builds
use :latest unless an image is specified explicitly:
vllm-sr config validate --config config.yaml
vllm-sr serve --platform nvidia --config config.yaml
For a source checkout, build the maintained CUDA image first and explicitly
select its latest tag. An editable CLI installation with a stable package
version otherwise selects the release tag. With the image override,
ifnotpresent reuses the local build while allowing the CLI to obtain missing
companion images:
VLLM_SR_PLATFORM=nvidia make vllm-sr-build
VLLM_SR_IMAGE=ghcr.io/vllm-project/semantic-router/vllm-sr-cuda:latest \
vllm-sr serve \
--platform nvidia \
--config config.yaml \
--image-pull-policy ifnotpresent
If you override the build tag or registry, set VLLM_SR_IMAGE to the actual
built image and use VLLM_SR_DASHBOARD_IMAGE for a custom companion image.
Pin a digest when deployments require an immutable image identity. If the Router shares a GPU with vLLM, measure memory and latency under representative concurrency; moving small, batch-one signal models to CUDA does not always improve end-to-end latency.
Verify the routed path
Check the local stack and Router logs:
vllm-sr status
vllm-sr logs router | grep model_binding_ready
nvidia-smi
Every prepared Router-side model reports the device it runs on, so a GPU
deployment shows "device":"cuda:0" for the configured classifiers and
nvidia-smi lists the Router process. A model reporting "device":"cpu" runs
on the CPU regardless of GPU passthrough. Then send a request through an
entrypoint exposed by that recipe. Replace
vllm-sr/auto if your config uses another public model name:
curl --fail --include http://127.0.0.1:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"messages": [{"role": "user", "content": "Explain prefix caching briefly."}],
"max_tokens": 64
}'
A successful direct vLLM request proves the model server works. A successful routed request proves the Router, recipe, and backend binding work together.
Troubleshooting
Docker rejects --gpus all
Configure Docker with nvidia-ctk, restart Docker, and repeat NVIDIA's sample
container command. Debug the container runtime before debugging either vLLM or
Semantic Router. This is not optional for the CUDA Router image: it links the
driver library, so without working passthrough the container exits at startup
with libcuda.so.1: cannot open shared object file.
The Router uses the CPU
Confirm that --platform nvidia selected the vllm-sr-cuda image and that
VLLM_SR_NVIDIA_PRESERVE_CPU is not enabled. Check the generated runtime
configuration and startup logs, not only the source recipe. A recipe without a
local signal model has nothing to move to CUDA.
vLLM or the Router runs out of GPU memory
The vLLM model, KV cache, and Router-side models compete for the same device memory. Leave the Router on CPU, reduce vLLM memory or concurrency settings, or place the services on separate GPUs. Do not assume that silent CPU fallback meets the same latency target.
The backend works directly but routed requests fail
Check that the backend endpoint is reachable from the Router container and
that provider_model_id exactly matches /v1/models. Inside the Router
container, localhost:8000 refers to the Router itself; use
host.docker.internal:8000, container DNS on a shared network, or a reachable
service address.
Kubernetes does not schedule a GPU
--platform nvidia is a local-container shortcut. For Kubernetes, choose the
CUDA image and configure GPU resources, the NVIDIA device plugin, and node
placement through Helm values or the Operator. See
Configuration Workflows for the deployment
boundary.