Skip to main content
Version: Latest

Troubleshooting and FAQ

Start with what the runtime says about itself. For a runtime you started:

curl -s localhost:8100/health
curl -s localhost:8100/v1/models

For the router's managed runtimes, look at the router's metrics and logs:

curl -s localhost:9190/metrics | grep '^vsr_model_runtime'

vsr_model_runtime_ready{deployment="..."} 1 means the deployment answers. The router log names each managed runtime process and why it stopped.

A signal never matches​

The model is probably not ready yet, or its answers arrive too late.

  1. Check vsr_model_runtime_ready for the deployment. While it is 0, every signal that uses it is unknown and does not match.

  2. Check vsr_model_runtime_unknown_answers_total by reason. timeout means the answer came after the signal's timeout_ms; raise it, or use a smaller model or a GPU. unavailable means the runtime was not ready or not reachable. overloaded means its queue was full.

  3. Send the same text with debugging on and read the x-vsr-matched-* headers:

    curl -s -D - -o /dev/null localhost:8899/v1/chat/completions \
    -H 'content-type: application/json' -H 'x-vsr-debug: true' \
    -d '{"model": "auto", "messages": [{"role": "user", "content": "your text"}]}'

The runtime stays in loading or warming​

  • First start: the model is downloading. Large models take minutes. The router log and GET /health show the stage.
  • No network: the runtime downloads from the Hugging Face Hub. Offline, copy the model into the cache first, or point artifact at a local copy.
  • warming for a long time on CPU: the self-check runs a few requests through the model. Large decision models on a CPU are slow; use a GPU or vllm-sr/Decision-2.0-Kai-0.6B.
  • loading with "retrying after ..." in the reason: the model failed to load for a reason that can pass, such as a busy GPU, too little free memory or an interrupted download. The runtime tries again up to five times, 5 seconds apart at first and doubling, while the other models in its process keep serving (--load-attempts, --load-retry-seconds). A damaged package or a failed self-check is reported as failed at once.

The runtime reports failed​

GET /v1/models shows the reason for each model. When the router runs the model, its log carries the same reason, for example model runtime is not ready: model @domain_classifier failed to load: .... When every model of a runtime process the router runs has failed, the router restarts that process (1 second at first, up to 60 seconds apart), so a passing cause such as a busy GPU or a full disk clears on its own. A model that fails beside models that still serve is retried by the runtime itself, so its process keeps running. A task model that still fails after three tries stops the router from starting. The common reasons:

Reason saysDo this
a file hash does not matchThe download is damaged or the repository changed. Delete that model from the cache and start again.
a revision is requiredA repository that is not built in needs revision with a 40-character commit.
access denied, gated or privateLog in with hf auth login or set HF_TOKEN for the account that has access. Vela 2.0 is a private preview.
does not fit, out of memoryUse a smaller model, a GPU with more memory, or give the model its own process.
device not availableThe named GPU does not exist or the installed PyTorch has no support for it. Use device: auto, or install the right PyTorch build.
built without LAPACKThe model needs LAPACK on the CPU, and this PyTorch (the ROCm image's) has none. Put the model on a GPU (device: rocm:0), or serve CPU models from the CPU image. The runtime does not retry it.
no family recognizes the packageThe model's architecture is not supported. See Choose a model.
a licence must be acceptedThe model's licence restricts use. Pass --accept-licence <id> after you checked that you may use it.

When one model of a process fails, the other models in that process keep serving.

A ready model's self-check says unverified​

GET /v1/models shows each model's self-check under golden. matched means its answers equal the released answers for that kind of device. unverified means the model serves, but the runtime can't vouch for that:

  • No reason, and golden.reference is empty: there are no released answers for this device, for example for a model that is not built in or a GPU without a validated record (CUDA). The runtime checked only that the answers are well formed and the same on every run.
  • reason starts with kernel choices not applied: on AMD Instinct MI300X and MI325X GPUs, Decision 2.0 and Vela 2.0 run their kernels with the configurations their release recorded with FLA 0.5.2, so every process gives the same answers. When those can't run here, the model still loads, the log says ... runs without its recorded kernel choices: ..., and its answers may differ from the released ones.
reason saysDo this
FLA is not installedInstall fla-core==0.5.2, or use the router's ROCm image, which has it.
they were recorded with FLA 0.5.2, not ...Install fla-core==0.5.2.
failed to importRepair the FLA installation; the message names the error.
has no autotuned kernel ...Install fla-core==0.5.2.

--autotune-cache keeps compiled kernels across restarts; it does not replace the recorded configurations.

The router refuses the configuration​

Message mentionsDo this
removed model execution fields, candle, ort, openvino, precision, variant, use_mmbert_32k, use_nli, polarity_guardRun vllm-sr config migrate. See Migrate.
a binding does not match the modelThe model has different labels, heads, dimensions or a smaller input limit than the feature needs. Bind a model made for that feature, or change input.
artifact must be a Hub repository ID or an absolute package pathUse owner/name for a Hub model or an absolute path for a local copy.
revision must be a 40-hex commitUse the full commit hash, not a branch or tag.

A long input is rejected​

Task deployments reject input longer than input.max_tokens instead of cutting it silently. Raise max_tokens up to the model's limit, or choose another policy:

global:
model_catalog:
deployments:
vela-pii:
provider: model_runtime
artifact: vllm-sr/Vela-1.0-Encoder-307M-PII
input:
max_tokens: 32768
overflow: window

truncate keeps the beginning of the text. window reads the whole text in overlapping windows and combines their results, which is what PII and safety scans should use so nothing is missed.

The router cannot reach an attached runtime​

  • The runtime listens on 127.0.0.1 unless you start it with --host 0.0.0.0.
  • From inside a container, localhost is the container itself. Use the host's address, a Docker network name or a Kubernetes Service.
  • curl <endpoint>/health from the router's machine must answer.

Requests are slower than expected​

  • Look at vllm_sr_runtime_request_duration_seconds and vllm_sr_runtime_queue_duration_seconds on the runtime's /metrics. Time spent queueing means the model is saturated: add a GPU, use a smaller model or another process.
  • On CPU, models of one process share the CPU threads. Start the runtime with --threads set to the cores you can give it.
  • Decision models on a GPU can use shared_context or batching; see Profiles.

FAQ​

Do I need a GPU? No. Every built-in task model and the smaller decision models run on a CPU.

Does the runtime send my data anywhere? No. It downloads model files from the Hugging Face Hub and never sends request text out. It does not log request text, and its metrics contain no request content.

Can I use a model that is not built in? Yes, for supported architectures: give a Hub repository with a revision, or an absolute local path. The runtime computes and reports the identity of models it does not know.

Is it safe to load models from the Hub? The runtime never runs code that ships inside a model repository, and it checks every file of a built-in model against its pinned hash before loading.

Can several routers share one runtime? Yes. Start it once and give every router the same endpoint.

Can one runtime serve several models? Yes. Pass several models to vllm-sr serve, or let the router group its deployments into processes.

What happens to my cached and stored vectors when I change the embedding model? They stay apart from new ones and are not reused. See Re-embed when the embedding model changes.

Where are models stored? In the Hugging Face cache (HF_HUB_CACHE) for a runtime you start, and under the router's model directory for managed runtimes; see the reference.