Common errors
Start with the first failing component instead of changing several settings at once:
vllm-sr status
vllm-sr logs router
vllm-sr logs envoy # only with --gateway extproc
vllm-sr config validate --config config.yaml
The examples below are fragments. Add them to the corresponding section of a complete canonical configuration and validate the result. Use config/config.yaml when you need the exhaustive field context.
Router cannot load the configuration
Failed to create ExtProc server
This is a top-level startup failure. The useful cause normally appears later on the same log line or immediately before it.
Check that the file exists and is readable, then run validation outside the container:
test -r config.yaml
vllm-sr config validate --config config.yaml
Follow the field path in the validation error. Do not add missing fields to a random nested block; canonical fields are location-sensitive.
failed to read config file
The process cannot open the path it received. Check:
- whether
--configis relative to the current working directory; - whether the same path exists inside the Router container;
- the file and parent-directory permissions; and
- whether a managed Recipe generated its runtime config in a different workspace.
Use vllm-sr status to identify the active workspace before inspecting
container mounts.
Entrypoint / Recipe Validation
vllm-sr config validate (or vllm-sr validate) includes a repair hint for common multi-recipe wiring errors.
Unknown recipe
Entrypoint references unknown recipe 'missing-recipe'
Hint: Change this to the name of a recipe defined under recipes.
Broken:
entrypoints:
- model_names: [my-model]
recipe: missing-recipe
recipes:
- name: production
Corrected:
entrypoints:
- model_names: [my-model]
recipe: production
recipes:
- name: production
Duplicate recipe name
Duplicate recipe name 'production'
Hint: Rename one recipe so every recipe has a unique name.
Give each recipe a distinct name, then update entrypoints that refer to the
renamed recipe:
recipes:
- name: production
- name: staging
entrypoints:
- model_names: [my-model]
recipe: production
Default entrypoint name collision
Entrypoint model 'vllm-sr/auto' is mapped more than once
vllm-sr/auto publishes the default recipe unless an explicit entrypoint
replaces it. Giving that same name to a named recipe creates a duplicate.
Broken:
entrypoints:
- model_names: [vllm-sr/auto]
recipe: production
Corrected:
entrypoints:
- model_names: [customer-production]
recipe: production
Entrypoint name collides with a backend model
A public recipe entrypoint must have its own name. The Router rejects an entrypoint whose name also identifies a backend model, including Fusion, ReMoM and Flow entrypoints. There is no special dispatch namespace for these algorithms.
Broken, if openai/gpt-oss-20b is a configured backend:
entrypoints:
- model_names: [openai/gpt-oss-20b]
recipe: flow
Corrected, with Flow requests sent to vllm-sr/flow and the flow recipe
containing decisions with algorithm.type: workflows:
entrypoints:
- model_names: [vllm-sr/flow]
recipe: flow
See the entrypoints and recipes tutorial and recipes tutorial for complete examples.
A request returns a routing error code
Look the code up in routing errors, then
find the Router's own reason in its log under the request's x-request-id:
vllm-sr logs router | grep '<x-request-id>'
For no_route, the entrypoint_routing_no_selection line names the model,
the recipe, and any decision that matched. Fusion, Flow and ReMoM are
selected by decisions inside that recipe; their public names are ordinary
entrypoints. Check the entrypoint mapping and recipe rules first.
Response cache cannot start
Backend configuration is required
Errors such as these mean backend_type selected a backend without its
matching configuration:
milvus configuration is required for Milvus cache backend
qdrant configuration is required for Qdrant cache backend
Milvus example:
global:
stores:
response_cache:
enabled: true
backend_type: milvus
milvus:
connection:
host: milvus
port: 19530
collection:
name: response_cache
Qdrant example:
global:
stores:
response_cache:
enabled: true
backend_type: qdrant
qdrant:
host: qdrant
port: 6334
collection_name: response_cache
The backend hostname must be resolvable from the Router container or pod, not only from the host.
Index or collection is missing
An error ending with auto-creation is disabled means the backing service is
reachable but the required index is absent. Choose one operating model:
- provision the index or collection before Router startup; or
- enable development-time creation for that backend.
Redis and Valkey use development.auto_create_index; Milvus uses
development.auto_create_collection. For example:
global:
stores:
response_cache:
enabled: true
backend_type: redis
redis:
# Add the connection, index, and search settings from the runtime example.
development:
auto_create_index: true
See the checked-in response-cache examples for complete backend blocks. Production deployments commonly provision schema separately and leave automatic creation disabled.
Cache hits are unexpectedly rare
First confirm that the decision enables the response_cache plugin and inspect
x-vsr-cache-hit or replay diagnostics. If embeddings are healthy but near
duplicates miss, test a lower similarity threshold on representative traffic:
global:
stores:
response_cache:
similarity_threshold: 0.75
A per-decision override belongs in the plugin configuration:
routing:
decisions:
- name: cached-route
plugins:
- type: response_cache
configuration:
enabled: true
semantic:
similarity_threshold: 0.70
Lower thresholds increase false-match risk. Evaluate answer equivalence before rolling them out, and use the management API's response-cache statistics and test endpoints when diagnosing the backend.
Responses API store cannot connect
An error such as:
failed to connect to Redis: redis ping failed
comes from the Responses API store, not the response cache. Check the separate service block:
global:
services:
response_api:
enabled: true
store_backend: redis
redis:
address: redis:6379
db: 0
Use store_backend: memory only for local work where losing stored responses
and conversation chains on restart is acceptable.
A PII route matches unexpectedly
With debug logging enabled, a matched PII rule includes the denied entity types:
[Signal Computation] PII rule "<name>" matched: denied_entities=[<types>]
If the type is allowed by policy, add it to that signal's allowlist. If the detector is producing low-confidence false positives, evaluate a higher threshold:
routing:
signals:
pii:
- name: pii-policy
threshold: 0.90
pii_types_allowed:
- GPE
- ORGANIZATION
Changing a privacy threshold changes false-negative risk. Validate it against a labeled dataset rather than a few hand-written prompts.
A jailbreak route matches unexpectedly
Contrastive classifier matches use a debug message beginning with:
[Signal Computation] Contrastive jailbreak rule "<name>" matched
Other jailbreak classifiers do not emit that exact phrase. Use replay or an
x-vsr-debug: true request to inspect matched signals and the selected
decision.
To reduce false positives, evaluate a higher threshold on the affected signal:
routing:
signals:
jailbreak:
- name: jailbreak-standard
threshold: 0.85
If a route should not depend on jailbreak detection, remove that condition from the route. Disabling the classifier globally changes every decision that uses it.
MCP category classifier cannot start
These errors identify an incomplete transport:
command is required for stdio transport
URL is required for HTTP transport
Use one transport and disable the local domain classifier when MCP should own category classification.
Stdio example:
global:
model_catalog:
modules:
classifier:
domain:
enabled: false
mcp:
enabled: true
transport_type: stdio
command: /app/bin/category-server
tool_name: classify_text
The executable and all arguments must exist inside the Router runtime.
Streamable HTTP example:
global:
model_catalog:
modules:
classifier:
domain:
enabled: false
mcp:
enabled: true
transport_type: streamable-http
url: http://mcp-server:8080/mcp
tool_name: classify_text
Test that URL from the Router network. If the server exposes another tool
name, set tool_name exactly or omit it to allow discovery of a recognized
classification tool.
Provider backend has no address
The validation error:
providers.models[<model>].backend_refs requires endpoint or base_url
means a backend reference cannot be resolved:
providers:
models:
- name: local-model
provider_model_id: local-model
api_format: openai
backend_refs:
- name: local-vllm
endpoint: 10.0.0.1:8000
protocol: http
provider: vllm
Use base_url when the provider requires a complete API root such as
https://provider.example/v1. Use a hostname reachable from the Router
network; localhost refers to the Router container itself.
The version segment in that API root is accepted for every Provider ID.
Providers whose own operation path already carries the version, such as
anthropic and minimax, do not repeat it, so both
https://provider.example/v1 and https://provider.example resolve to the
same upstream path.
See Container connectivity for end-to-end checks.
A classifier or embedding model cannot load
Every model runs in the model runtime. When a
model cannot load, its feature reports unknown results and the runtime records
why. Check the router's vsr_model_runtime_ready{deployment="..."} metric and
the router log, or ask a runtime you run yourself:
curl -s localhost:8100/v1/models
Each model's status and reason say what failed: a damaged download, a
missing revision, a private repository without a token, a device that does
not exist, or a model too large for its device.
Troubleshooting and FAQ lists each
reason and the fix.
Container image has no matching platform
ImagePullBackOff with no matching manifest means the selected tag or digest
does not contain an image for the node architecture.
Inspect the exact reference from the deployment:
docker buildx imagetools inspect <registry>/<image>:<tag>
If the tag is a multi-platform index, pin its index digest rather than one
architecture's child manifest. If it is single-platform, use a release that
publishes the required architecture or build and publish the image through
your approved pipeline. Do not switch to an unpinned latest tag as a
long-term immutability workaround.
Classification confidence is too low
If requests frequently fall back after domain classification, confirm the loaded model and category mapping before changing the threshold. Then evaluate a lower value on labeled traffic:
global:
model_catalog:
modules:
classifier:
domain:
threshold: 0.50
Lowering the threshold may increase wrong-domain routes. Report per-category precision and recall at the chosen operating point.
Diagnostic commands
# Validate the source configuration.
vllm-sr config validate --config config.yaml
# Identify the active local stack and component state.
vllm-sr status
# Read component logs without depending on generated container names.
vllm-sr logs router
vllm-sr logs envoy # only with --gateway extproc
# Check the public listener and model catalog.
curl -sS http://localhost:8899/v1/models
# Check management health and metrics.
curl -sS http://localhost:8080/health
curl -sS http://localhost:9190/metrics | head
If a provider succeeds from the host but fails from the Router, continue with Container connectivity. For image and artifact downloads, see Restricted network environments.