Skip to main content
Version: Latest

Deploy with NVIDIA CUDA

The model server and Semantic Router are separate services. A common deployment keeps the Router on CPU and gives the NVIDIA GPU to vLLM. Use the Router's CUDA image when its local embeddings or classifiers also need GPU acceleration.

--platform nvidia affects the local Router stack only. It selects the CUDA Router image, passes NVIDIA GPUs into the Router container, and changes its generated runtime configuration so supported local signal models prefer CUDA. It does not download a language model or start a vLLM server.

Prerequisites​

  • Linux and an NVIDIA GPU supported by the vLLM release you plan to run;
  • an x86-64 host when using the current Semantic Router CUDA image;
  • a GPU of compute capability 7.0 or newer (Volta and later) for the Router image, which compiles its local models for that minimum; build with CUDA_COMPUTE_CAP=<value> to target an older or newer floor;
  • an NVIDIA driver, and GPU passthrough into the Router container. The CUDA Router image links the driver library directly, so it does not start without passthrough even when every Router-side model is configured for CPU;
  • Docker and NVIDIA Container Toolkit;
  • enough GPU memory for the vLLM model, KV cache, and any Router-side models; and
  • a complete Semantic Router configuration with a reachable model endpoint.

Use the current vLLM NVIDIA requirements for supported hardware. Install and configure the runtime with the NVIDIA Container Toolkit guide, then verify both the host driver and container access:

nvidia-smi
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi

Do not continue until the container command can see the expected GPUs. The second command is NVIDIA's sample workload.

Start and verify a vLLM backend​

The following example follows the official vLLM Docker deployment. It publishes an OpenAI-compatible endpoint on port 8000 and keeps downloaded model files in a named volume:

docker volume create vllm-huggingface-cache

docker run -d \
--name vllm-nvidia \
--runtime nvidia \
--gpus all \
--ipc=host \
-p 8000:8000 \
-v vllm-huggingface-cache:/root/.cache/huggingface \
vllm/vllm-openai:latest \
--model Qwen/Qwen3-0.6B

Choose a model and vLLM arguments that fit the available GPUs. Pass HF_TOKEN as an environment variable when a model requires authentication; do not put the token in an image, command history, or Router config. Pin the vLLM image and model revision for a controlled deployment instead of relying on latest.

Wait for model loading to finish, then test vLLM before adding the Router:

curl --fail http://127.0.0.1:8000/health
curl --fail http://127.0.0.1:8000/v1/models

curl --fail http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [{"role": "user", "content": "Reply with: ready"}],
"max_tokens": 16
}'

Router validation cannot prove that a backend can load a model or generate a response, so fix any direct vLLM error before continuing.

Connect the backend​

Bind the served model in your canonical config. For the local Docker stack, the Router can reach a host-published port through host.docker.internal:

providers:
defaults:
model: local/qwen
models:
- name: local/qwen
provider_model_id: Qwen/Qwen3-0.6B
api_format: openai
backend_refs:
- name: nvidia-vllm
endpoint: host.docker.internal:8000
protocol: http
provider: vllm
weight: 1

This is a provider fragment, not a complete Router config. Add the matching model card and route to your existing recipe, or configure the endpoint in the Dashboard. The provider_model_id must match a model returned by vLLM's /v1/models endpoint. See Configuration for a complete minimal document and Models, Entrypoints, and Serving for backend binding and routing policy.

The example publishes port 8000 on the host for direct testing. Restrict that port with host networking controls, or use private service discovery in a production deployment. Do not expose an unauthenticated vLLM endpoint to an untrusted network.

Run the Router on NVIDIA​

If vLLM should own all GPU memory, keep the Router on CPU:

vllm-sr config validate --config config.yaml
vllm-sr serve --config config.yaml

To run supported Router-side local embeddings and classifiers on CUDA, use --platform nvidia. A stable CLI selects the matching published release image (for example, CLI 0.4.0 uses vllm-sr-cuda:v0.4.0). Development CLI builds use :latest unless an image is specified explicitly:

vllm-sr config validate --config config.yaml
vllm-sr serve --platform nvidia --config config.yaml

For a source checkout, build the maintained CUDA image first and explicitly select its latest tag. An editable CLI installation with a stable package version otherwise selects the release tag. With the image override, ifnotpresent reuses the local build while allowing the CLI to obtain missing companion images:

VLLM_SR_PLATFORM=nvidia make vllm-sr-build
VLLM_SR_IMAGE=ghcr.io/vllm-project/semantic-router/vllm-sr-cuda:latest \
vllm-sr serve \
--platform nvidia \
--config config.yaml \
--image-pull-policy ifnotpresent

If you override the build tag or registry, set VLLM_SR_IMAGE to the actual built image and use VLLM_SR_DASHBOARD_IMAGE for a custom companion image.

Pin a digest when deployments require an immutable image identity. If the Router shares a GPU with vLLM, measure memory and latency under representative concurrency; moving small, batch-one signal models to CUDA does not always improve end-to-end latency.

Verify the routed path​

Check the local stack and Router logs:

vllm-sr status
vllm-sr logs router | grep model_binding_ready
nvidia-smi

Every prepared Router-side model reports the device it runs on, so a GPU deployment shows "device":"cuda:0" for the configured classifiers and nvidia-smi lists the Router process. A model reporting "device":"cpu" runs on the CPU regardless of GPU passthrough. Then send a request through an entrypoint exposed by that recipe. Replace vllm-sr/auto if your config uses another public model name:

curl --fail --include http://127.0.0.1:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"messages": [{"role": "user", "content": "Explain prefix caching briefly."}],
"max_tokens": 64
}'

A successful direct vLLM request proves the model server works. A successful routed request proves the Router, recipe, and backend binding work together.

Troubleshooting​

Docker rejects --gpus all​

Configure Docker with nvidia-ctk, restart Docker, and repeat NVIDIA's sample container command. Debug the container runtime before debugging either vLLM or Semantic Router. This is not optional for the CUDA Router image: it links the driver library, so without working passthrough the container exits at startup with libcuda.so.1: cannot open shared object file.

The Router uses the CPU​

Confirm that --platform nvidia selected the vllm-sr-cuda image and that VLLM_SR_NVIDIA_PRESERVE_CPU is not enabled. Check the generated runtime configuration and startup logs, not only the source recipe. A recipe without a local signal model has nothing to move to CUDA.

vLLM or the Router runs out of GPU memory​

The vLLM model, KV cache, and Router-side models compete for the same device memory. Leave the Router on CPU, reduce vLLM memory or concurrency settings, or place the services on separate GPUs. Do not assume that silent CPU fallback meets the same latency target.

The backend works directly but routed requests fail​

Check that the backend endpoint is reachable from the Router container and that provider_model_id exactly matches /v1/models. Inside the Router container, localhost:8000 refers to the Router itself; use host.docker.internal:8000, container DNS on a shared network, or a reachable service address.

Kubernetes does not schedule a GPU​

--platform nvidia is a local-container shortcut. For Kubernetes, choose the CUDA image and configure GPU resources, the NVIDIA device plugin, and node placement through Helm values or the Operator. See Configuration Workflows for the deployment boundary.