API Reference
Packages
vllm.ai/v1alpha1
Package v1alpha1 contains API Schema definitions for the vllm v1alpha1 API group
Resource Types
APIConfig
APIConfig defines API configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
batch_classification BatchClassificationConfig | Optional: {} |
AutoscalingSpec
AutoscalingSpec defines autoscaling configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled indicates if HPA is enabled | false | Optional: {} |
minReplicas integer | MinReplicas is the minimum number of replicas | 1 | Optional: {} |
maxReplicas integer | MaxReplicas is the maximum number of replicas | 10 | Optional: {} |
targetCPUUtilizationPercentage integer | TargetCPUUtilizationPercentage is the target CPU percentage | 80 | Optional: {} |
targetMemoryUtilizationPercentage integer | TargetMemoryUtilizationPercentage is the target memory percentage | Optional: {} |
BatchClassificationConfig
BatchClassificationConfig defines batch classification configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
max_batch_size integer | 100 | Optional: {} | |
concurrency_threshold integer | 5 | Optional: {} | |
max_concurrency integer | 8 | Optional: {} | |
metrics BatchMetricsConfig | Optional: {} |
BatchMetricsConfig
BatchMetricsConfig defines batch classification metrics configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | true | Optional: {} | |
detailed_goroutine_tracking boolean | true | Optional: {} | |
high_resolution_timing boolean | false | Optional: {} | |
sample_rate string | Sample rate for metrics (0.0-1.0). Stored as string to avoid float precision issues. | 1.0 | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} |
duration_buckets string array | Duration buckets for histograms. Stored as strings to avoid float precision issues. Example: ["0.001", "0.005", "0.01", "0.025", "0.05", "0.1", "0.25", "0.5", "1", "2.5", "5", "10", "30"] | Optional: {} | |
size_buckets integer array | Optional: {} |
CategoryModelConfig
CategoryModelConfig defines category model configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
model_id string | Optional: {} | ||
threshold string | Classification threshold (0.0-1.0). Stored as string to avoid float precision issues. | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} | |
use_cpu boolean | Optional: {} | ||
category_mapping_path string | Optional: {} |
ClassifierConfig
ClassifierConfig defines classifier configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
category_model CategoryModelConfig | Optional: {} | ||
pii_model PIIModelConfig | Optional: {} |
ComplexityCandidates
ComplexityCandidates defines candidate examples for complexity classification
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
candidates string array | List of candidate phrases or examples |
ComplexityModelConfig
ComplexityModelConfig configures how the complexity signal produces its score. It mirrors global.model_catalog.modules.complexity in the router config and is passed through field for field.
The contract requirement sits here rather than on RemoteClassifierBackendConfig because it is a property of this consumer, not of the block: complexity reads two response shapes, so guessing wrong would surface per request instead of at admission. A consumer that reads one shape
- categories does - keeps the field optional and defaults it.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
backend RemoteClassifierBackendConfig | Backend names a remote scoring model. Its absence keeps local prototype scoring; when set, the signal never reads the rules' hard/easy candidates. It sits on the module rather than on a rule because routing signals are replaced wholesale per recipe, so a per-rule backend would vanish under any recipe that did not repeat it. | Optional: {} |
ComplexityRulesConfig
ComplexityRulesConfig defines complexity-based signal classification.
The CEL rules below reject at admission the boundary combinations the Router refuses at config load. Without them the API server accepts the object and the Router crashloops on it, which turns a typo into an outage instead of a rejected write. They are per-object and static; anything needing the model catalog - whether backend.model resolves, for instance - stays with the Router's validator, which remains the single source of truth for the rest.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
prototype_scoring PrototypeScoringConfig | PrototypeScoring replaces the family prototype-scoring configuration for this rule. Absence inherits the family; a present object is a complete override, including an empty object. Defaults remain Router-owned. | Optional: {} | |
name string | Name of the complexity rule (e.g., "code-complexity", "reasoning-complexity") | ||
description string | Description of what this rule classifies | Optional: {} | |
threshold string | Threshold for the local prototype-scoring path (0.0-1.0), stored as a string to avoid float precision issues. The local margin is hard-minus-easy and centred on zero, so the threshold is symmetric: a margin above it is "hard", below its negative is "easy", and in between is "medium". It does not apply under a score.v1 backend, whose score is in the model's own units; state a boundary pair instead. | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} | |
hard_above string | HardAbove and EasyBelow are the two cut points for a score where a higher value is harder, in the scoring model's own units - so no [0,1] pattern applies and negative values are valid. Both are required together, and the pair is mutually exclusive with Threshold and with HardBelow/EasyAbove. Stored as strings to avoid float precision issues. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} | |
easy_below string | EasyBelow is the lower cut point of the harder-when-higher pair: a score below it is "easy", and anything between EasyBelow and HardAbove is "medium". It must be below HardAbove, and both are required together. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} | |
hard_below string | HardBelow and EasyAbove are the pair for a score where a lower value is harder - a model predicting the chance of a correct answer, say. They require a score.v1 backend: the local margin is harder-when-higher by construction, and inverting it locally means swapping the candidate lists. Stored as strings to avoid float precision issues. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} | |
easy_above string | EasyAbove is the upper cut point of the harder-when-lower pair: a score above it is "easy", and anything between HardBelow and EasyAbove is "medium". It must be above HardBelow, and both are required together. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} | |
hard ComplexityCandidates | Hard candidates represent complex/difficult examples. Read only by the local path; a remote backend never consults them, so they are optional. | Optional: {} | |
easy ComplexityCandidates | Easy candidates represent simple/easy examples. Read only by the local path; a remote backend never consults them, so they are optional. | Optional: {} | |
composer RuleComposition | Composer allows filtering based on other signals (e.g., only apply this rule if domain:medical) | Optional: {} |
CompositionCondition
CompositionCondition defines a single composition condition
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | Type of signal to check (e.g., "domain", "language", "category") | ||
name string | Name of the specific signal/rule value to match |
ConfigSpec
ConfigSpec defines the semantic router configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
routing JSON | Routing contains canonical v0.3 routing configuration under config.routing. It is intentionally preserved as an object so the operator can pass through the router-owned signal, projection, decision, and algorithm contract without lagging behind every router schema addition. | Type: object Optional: {} | |
decision_model DecisionModelBinding | DecisionModel selects the declared deployment that answers default judgment tasks. Omitted uses the Router's primary deployment. Artifact identity belongs in model_deployments; the reference never infers a model family or alias. | Optional: {} | |
model_deployments JSON | ModelDeployments contains canonical global.model_catalog.deployments. The router validates provider, device, precision and task compatibility. | Type: object Optional: {} | |
model_admission JSON | ModelAdmission contains canonical global.model_catalog.admission budgets. Keys name deployments or the router's existing admission consumers. | Type: object Optional: {} | |
embedding_models EmbeddingModelsConfig | Embedding models configuration (qwen3, gemma, mmbert) | Optional: {} | |
response_cache SemanticCacheConfig | Response cache configuration. | Optional: {} | |
semantic_cache SemanticCacheConfig | SemanticCache is the deprecated response-cache field. | Optional: {} | |
tools ToolsConfig | Tools configuration | Optional: {} | |
prompt_guard PromptGuardConfig | Prompt guard configuration | Optional: {} | |
classifier ClassifierConfig | Classifier configuration | Optional: {} | |
complexity_rules ComplexityRulesConfig array | Complexity rules for complexity-aware routing | Optional: {} | |
complexity_model ComplexityModelConfig | ComplexityModel says how the complexity signal produces its score. Absent, the signal scores locally against each rule's hard/easy candidates. With a backend, a remote model produces the score and the candidates are never read. Mirrors global.model_catalog.modules.complexity in the router config. | Optional: {} | |
external_models ExternalModelConfig array | ExternalModels declares the remote models that classifier backends ( classifier.pii.backend.model, complexity_model.backend.model) andthe prompt guard protocol refer to by name. Mirrors global.model_catalog.external[] in the router config field for field; the router's own validator decides whether a backend resolves against it. | Optional: {} | |
strategy string | Decision routing strategy ("priority" for priority-based matching) | Enum: [priority] Optional: {} | |
decisions DecisionConfig array | Routing decisions based on signals (domain, complexity, etc.) | Optional: {} | |
reasoning_effort string | ReasoningEffort is the default reasoning effort for model bindings that do not select a different effort. The selected model family validates the value because built-in and custom families may expose different ladders. | Optional: {} | |
api APIConfig | API configuration | Optional: {} | |
observability ObservabilityConfig | Observability configuration | Optional: {} | |
streamed_body StreamedBodyConfig | StreamedBody enables streamed request body handling. Mirrors global.router.streamed_body; the gateway must send bodies to ExtProc in STREAMED or FullDuplexStreamed mode for it to take effect. | Optional: {} |
DecisionConfig
DecisionConfig defines a routing decision
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name is the unique identifier for this decision | ||
description string | Description provides information about what this decision handles | Optional: {} | |
priority integer | Priority is used for decision ordering - higher priority decisions are evaluated first | Optional: {} | |
rules RuleCombinationConfig | Rules defines the combination of conditions using AND/OR logic | ||
modelRefs ModelRefConfig array | ModelRefs contains model references for this decision | Optional: {} | |
preferred_endpoints string array | PreferredEndpoints specifies which vLLM endpoints to prefer for this decision | Optional: {} | |
plugins RawExtension array | Plugins contains policy configurations applied after rule matching | Optional: {} | |
algorithm JSON | Algorithm configures base model selection for this decision. It is preserved as a router-owned object so supported algorithms can evolve without requiring the operator CRD to duplicate every nested field. | Type: object Optional: {} | |
reliability DecisionReliabilityConfig | Reliability overrides the timeouts and retries of the provider model that serves this decision's requests | Optional: {} | |
fallback DecisionFallbackConfig | Fallback overrides the cross-model fallback policy for this decision's requests, over the recipe's and the router's | Optional: {} |
DecisionFallbackConfig
DecisionFallbackConfig is a decision's fallback block, with the router configuration's field names and meaning. A field left out keeps the recipe's or the router's value.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled turns cross-model fallback on or off for this decision | Optional: {} | |
max_attempts integer | MaxAttempts bounds the candidates tried, the first one included | Minimum: 0 Optional: {} | |
total_timeout string | TotalTimeout bounds the whole fallback chain | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
per_attempt_timeout string | PerAttemptTimeout bounds each candidate's attempt | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
retryable_status_codes integer array | RetryableStatusCodes are the statuses that move on to the next candidate | items:Maximum: 599 items:Minimum: 100 Optional: {} |
DecisionModelBinding
DecisionModelBinding selects one canonical model deployment.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
deployment string | Deployment is an exact key in model_deployments. | MinLength: 1 |
DecisionReliabilityConfig
DecisionReliabilityConfig is a decision's reliability block, with the router configuration's field names and meaning. Durations are Go durations such as "30s"; "0s" turns a timeout off.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
total_timeout string | TotalTimeout bounds the whole call: every attempt and the response | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
per_try_timeout string | PerTryTimeout bounds each attempt until its response starts | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
idle_timeout string | IdleTimeout bounds the wait for more of a streamed response (standalone mode only) | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
first_byte_timeout string | FirstByteTimeout bounds the wait for the first response byte (standalone mode only) | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
retry_count integer | RetryCount is the number of retries after the first attempt | Maximum: 5 Minimum: 0 Optional: {} | |
retry_on string | RetryOn adds retry conditions, with Envoy's names (5xx, reset, ...) | Optional: {} | |
retriable_status_codes integer array | RetriableStatusCodes adds statuses retried under retriable-status-codes | items:Maximum: 599 items:Minimum: 100 Optional: {} | |
retry_back_off_base string | RetryBackOffBase is the base of the randomized exponential wait between retries (standalone mode only) | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
retry_back_off_max string | RetryBackOffMax caps that wait (standalone mode only) | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} | |
retry_after_max string | RetryAfterMax honors a response's Retry-After up to this bound (standalone mode only) | Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$ Optional: {} |
EmbeddingEndpointConfig
EmbeddingEndpointConfig defines an external OpenAI-compatible embedding endpoint.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
base_url string | BaseURL is the base URL for the embedding endpoint, typically ending in /v1. | Optional: {} | |
model string | Model is the embedding model name sent to the external provider. | Optional: {} | |
api_key_env string | APIKeyEnv names the environment variable containing the provider API key. | Optional: {} | |
timeout_seconds integer | TimeoutSeconds is the request timeout for embedding calls. | Minimum: 0 Optional: {} | |
max_retries integer | MaxRetries is the maximum number of retry attempts for embedding calls. | Minimum: 0 Optional: {} | |
max_response_bytes integer | MaxResponseBytes caps the size of each embedding response body. | Minimum: 0 Optional: {} | |
dimensions integer | Dimensions requests a provider-side output dimension when supported. | Minimum: 1 Optional: {} |
EmbeddingModelsConfig
EmbeddingModelsConfig defines configuration for embedding models
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
qwen3_model_path string | Path to Qwen3-Embedding-0.6B model directory Qwen3 provides 32K context and high quality embeddings (1024 dimensions) | Optional: {} | |
mmbert_model_path string | Path to mmBERT 2D Matryoshka embedding model directory Supports layer early exit (3/6/11/22) and dimension reduction (64-768) | Optional: {} | |
use_cpu boolean | Use CPU for inference (default: true) | true | Optional: {} |
embedding_config HNSWEmbeddingConfig | Embedding configuration for embedding-based classification | Optional: {} | |
endpoint EmbeddingEndpointConfig | Endpoint configures an external embedding provider endpoint. The API key should be injected into the semantic router pod environment and referenced by APIKeyEnv rather than stored directly in the CR. | Optional: {} |
ExporterConfig
ExporterConfig defines exporter configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | otlp | Optional: {} | |
endpoint string | jaeger:4317 | Optional: {} | |
insecure boolean | true | Optional: {} |
ExternalModelConfig
ExternalModelConfig is one entry of global.model_catalog.external[]: a remote model a classifier backend or the prompt guard can name. Field names are the router's YAML keys so the generic typed conversion carries them unchanged.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name is the catalog name a backend block refers to in its model field. | MinLength: 1 | |
model_role string | ModelRole is what the model is used for; classifier backends require "classification", the prompt guard protocol requires "guardrail". | MinLength: 1 | |
llm_model_name string | ModelName is the model identifier the remote service expects, and the value a token_spans.v1 envelope's model member must equal. | MinLength: 1 | |
llm_endpoint ExternalModelEndpoint | Endpoint is where the remote model is reached. | ||
llm_timeout_seconds integer | TimeoutSeconds bounds one call when the backend block sets no deadline. | Minimum: 1 Optional: {} |
ExternalModelEndpoint
ExternalModelEndpoint is the address of a remote classification model.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
address string | MinLength: 1 | ||
port integer | Maximum: 65535 Minimum: 1 | ||
protocol string | Enum: [http https] Optional: {} |
GatewayReference
GatewayReference references an existing Gateway
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the Gateway | MinLength: 1 | |
namespace string | Namespace of the Gateway | MinLength: 1 |
GatewaySpec
GatewaySpec defines Gateway API integration configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
existingRef GatewayReference | ExistingRef references an existing Gateway that calls the Router over ext_proc. Setting it selects extproc mode. | Optional: {} |
HNSWCacheConfig
HNSWCacheConfig defines HNSW index configuration for hybrid/in-memory backends.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
use_hnsw boolean | UseHNSW enables HNSW indexing for faster similarity search | false | Optional: {} |
hnsw_m integer | M is the number of bi-directional links per node | 16 | Minimum: 2 Optional: {} |
hnsw_ef_construction integer | EfConstruction is the size of dynamic candidate list during construction | 200 | Minimum: 1 Optional: {} |
max_memory_entries integer | MaxMemoryEntries limits in-memory entries for hybrid backend | 1000 | Minimum: 0 Optional: {} |
HNSWEmbeddingConfig
HNSWEmbeddingConfig contains settings for embedding classification with HNSW indexing
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
backend string | Backend selects the embedding provider backend: the built-in model runtime (the default) or an external OpenAI-compatible endpoint. | Enum: [model_runtime openai_compatible] Optional: {} | |
model_type string | ModelType specifies which embedding model to use Options: "qwen3" (1024-dim, 32K context), "gemma" (768-dim, 8K context), "mmbert" (64-768-dim, multilingual), "remote" (external provider) | Enum: [qwen3 gemma mmbert remote] Optional: {} | |
preload_embeddings boolean | PreloadEmbeddings enables precomputing candidate embeddings at startup | true | Optional: {} |
target_dimension integer | TargetDimension is the embedding dimension to use (default: 768) For mmBERT, supported local dimensions are 64, 128, 256, 512, 768. External providers may use other positive dimensions such as 1024, 1536, or 3072. | Minimum: 1 Optional: {} | |
target_layer integer | TargetLayer controls mmBERT early exit and is used only when ModelType is "mmbert". Lower layers reduce encoder work but may reduce quality; layer 22 uses the full encoder depth. Evaluate the latency and quality trade-off on representative deployment data. | Enum: [3 6 11 22] Optional: {} | |
enable_soft_matching boolean | EnableSoftMatching allows below-threshold matches when no rule meets its threshold. | false | Optional: {} |
min_score_threshold string | MinScoreThreshold for matching (0.0-1.0). Stored as string to avoid float precision issues. | 0.5 | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} |
ImageSpec
ImageSpec defines the container image configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
repository string | Repository is the container image repository | ghcr.io/vllm-project/semantic-router/vllm-sr | Optional: {} |
tag string | Tag is the container image tag | latest | Optional: {} |
pullPolicy PullPolicy | PullPolicy is the image pull policy | IfNotPresent | Enum: [Always Never IfNotPresent] Optional: {} |
imageRegistry string | ImageRegistry is an optional registry prefix | Optional: {} |
IngressHost
IngressHost defines an ingress host
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
host string | Optional: {} | ||
paths IngressPath array | Optional: {} |
IngressPath
IngressPath defines an ingress path
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
path string | Optional: {} | ||
pathType string | Optional: {} | ||
servicePort integer | Optional: {} |
IngressSpec
IngressSpec defines ingress configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled indicates if ingress is enabled | false | Optional: {} |
className string | ClassName is the ingress class name | Optional: {} | |
annotations object (keys:string, values:string) | Annotations for ingress | Optional: {} | |
hosts IngressHost array | Hosts configuration | Optional: {} | |
tls IngressTLS array | TLS configuration | Optional: {} |
IngressTLS
IngressTLS defines ingress TLS configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
secretName string | Optional: {} | ||
hosts string array | Optional: {} |
LoRAAdapterSpec
LoRAAdapterSpec defines one LoRA adapter exposed by a VLLMEndpoint model.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name is the unique adapter identifier referenced by decision.modelRefs[].lora_name. | MaxLength: 100 MinLength: 1 | |
description string | Description provides a short human-readable summary for UI and docs surfaces. | MaxLength: 500 Optional: {} |
MetricsPortSpec
MetricsPortSpec extends PortSpec with enable flag
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
port integer | Port is the service port | Maximum: 65535 Minimum: 1 Optional: {} | |
targetPort integer | TargetPort is the container port | Maximum: 65535 Minimum: 1 Optional: {} | |
protocol Protocol | Protocol is the port protocol | TCP | Optional: {} |
enabled boolean | Enabled indicates if metrics should be exposed | true | Optional: {} |
MilvusCacheAuth
MilvusCacheAuth defines Milvus authentication.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled controls whether to use authentication | false | Optional: {} |
username string | Username for Milvus authentication | Optional: {} | |
password string | Password for Milvus authentication (plaintext - consider using PasswordSecretRef instead) | Optional: {} | |
password_secret_ref SecretKeySelector | PasswordSecretRef references a Secret containing the Milvus password Preferred over plaintext Password field for security | Optional: {} |
MilvusCacheCollection
MilvusCacheCollection defines Milvus collection configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the Milvus collection | semantic_cache | Optional: {} |
description string | Description of the collection | Semantic cache for LLM request-response pairs | Optional: {} |
vector_field MilvusCacheVectorField | VectorField configuration for embeddings | Optional: {} | |
index MilvusCacheCollectionIndex | Index configuration for the collection | Optional: {} |
MilvusCacheCollectionIndex
MilvusCacheCollectionIndex defines collection index settings.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | Type of index algorithm | HNSW | Enum: [HNSW IVF_FLAT IVF_SQ8 IVF_PQ] Optional: {} |
params MilvusCacheIndexParams | Params for the index | Optional: {} |
MilvusCacheConfig
MilvusCacheConfig defines Milvus cache backend configuration. Configure these settings when using Milvus as the semantic cache backend.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
connection MilvusCacheConnection | Connection settings for Milvus server | Optional: {} | |
collection MilvusCacheCollection | Collection settings for Milvus | Optional: {} | |
search MilvusCacheSearch | Search settings for Milvus queries | Optional: {} | |
development MilvusCacheDevelopment | Development settings for Milvus cache | Optional: {} |
MilvusCacheConnection
MilvusCacheConnection defines Milvus connection parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
host string | Host is the Milvus server hostname or IP address | Optional: {} | |
port integer | Port is the Milvus server port | 19530 | Maximum: 65535 Minimum: 1 Optional: {} |
database string | Database name in Milvus | semantic_router_cache | Optional: {} |
timeout integer | Timeout for Milvus operations in seconds | 30 | Minimum: 0 Optional: {} |
auth MilvusCacheAuth | Auth configuration for Milvus authentication | Optional: {} | |
tls MilvusCacheTLS | TLS configuration for secure Milvus connections | Optional: {} |
MilvusCacheDevelopment
MilvusCacheDevelopment defines development-mode settings.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
drop_collection_on_startup boolean | DropCollectionOnStartup clears the collection when router starts (for testing) | false | Optional: {} |
auto_create_collection boolean | AutoCreateCollection automatically creates the collection if it doesn't exist | true | Optional: {} |
MilvusCacheIndexParams
MilvusCacheIndexParams defines index parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
M integer | M is the number of bi-directional links for HNSW | 16 | Minimum: 2 Optional: {} |
efConstruction integer | EfConstruction for HNSW index building | 64 | Minimum: 1 Optional: {} |
MilvusCacheSearch
MilvusCacheSearch defines Milvus search parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
params MilvusCacheSearchParams | Params for search operations | Optional: {} | |
topk integer | TopK is the number of results to return | 10 | Minimum: 1 Optional: {} |
consistency_level string | ConsistencyLevel for search operations Options: "Strong", "Session", "Bounded", "Eventually" | Session | Enum: [Strong Session Bounded Eventually] Optional: {} |
MilvusCacheSearchParams
MilvusCacheSearchParams defines search-time parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
ef integer | Ef is the search-time HNSW parameter | 64 | Minimum: 1 Optional: {} |
MilvusCacheTLS
MilvusCacheTLS defines TLS settings for Milvus connections.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled controls whether to use TLS | false | Optional: {} |
cert_file string | CertFile is the path to client certificate file | Optional: {} | |
key_file string | KeyFile is the path to client key file | Optional: {} | |
ca_file string | CAFile is the path to CA certificate file | Optional: {} |
MilvusCacheVectorField
MilvusCacheVectorField defines vector field configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the vector field | embedding | Optional: {} |
dimension integer | Dimension of the embedding vectors | Minimum: 1 Optional: {} | |
metric_type string | MetricType for vector similarity Options: "IP" (inner product), "L2", "COSINE" | IP | Enum: [IP L2 COSINE] Optional: {} |
ModelReasoningSpec
ModelReasoningSpec selects a catalog reasoning family or defines the request projection for a custom self-hosted model. Family and inline fields are mutually exclusive and are validated by the Router's canonical compiler.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
family string | Optional: {} | ||
type string | Enum: [chat_template_kwargs reasoning_effort reasoning_mode top_level_reasoning_effort] Optional: {} | ||
parameter string | Optional: {} | ||
activationParameter string | Optional: {} | ||
effortFlags object (keys:string, values:string) | EffortFlags maps a logical effort to a boolean chat-template parameter. | Optional: {} | |
levels string array | Optional: {} | ||
default string | Optional: {} | ||
modes string array | items:Enum: [enabled disabled adaptive] Optional: {} | ||
defaultMode string | Enum: [enabled disabled adaptive] Optional: {} | ||
disabled string | Optional: {} |
ModelRefConfig
ModelRefConfig defines a model reference for routing
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
model string | Model name to route to | ||
lora_name string | LoRAName is the optional LoRA adapter name | Optional: {} | |
use_reasoning boolean | UseReasoning enables reasoning mode for this model | Optional: {} | |
reasoning_mode string | ReasoningMode selects the model's reasoning activation mode when the family supports more than a boolean switch. | Enum: [enabled disabled adaptive] Optional: {} | |
reasoning_effort string | ReasoningEffort selects one of the model family's declared effort levels. | Optional: {} |
ObservabilityConfig
ObservabilityConfig defines observability configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
tracing TracingConfig | Optional: {} |
OpenShiftFeaturesStatus
OpenShiftFeaturesStatus tracks OpenShift-specific feature status
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
routesEnabled boolean | RoutesEnabled indicates if OpenShift Routes are enabled | ||
routeHostname string | RouteHostname is the hostname of the created Route | Optional: {} |
OpenShiftSpec
OpenShiftSpec defines OpenShift-specific configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
routes RouteConfig | Routes configuration for OpenShift Routes | Optional: {} |
PIIModelConfig
PIIModelConfig defines PII model configuration.
The contract rule sits on the consumer, as on ComplexityModelConfig: the shared backend block lists every contract any consumer reads, and each consumer narrows it to what it can parse, so a mismatch is refused at admission instead of by the router at load.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
max_sequence_length integer | MaxSequenceLength is the total tokenized input budget, including special tokens. Omission or zero preserves the 512-token legacy limit. | Minimum: 0 Optional: {} | |
window PromptGuardWindowConfig | Window scans original content tokens with explicit overlap. Omission or null leaves window selection unchanged; no CRD defaults are injected. | Optional: {} | |
model_id string | Optional: {} | ||
threshold string | Detection threshold (0.0-1.0). Stored as string to avoid float precision issues. | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} | |
use_cpu boolean | Optional: {} | ||
pii_mapping_path string | Optional: {} | ||
backend RemoteClassifierBackendConfig | Backend names a remote token classifier speaking token_spans.v1. Its absence keeps local PII inference. The local selectors this replaces are model_id and use_cpu above. Explicit token windows are only supported by the local model. | Optional: {} | |
on_error string | OnError selects what a PII backend failure, or a provider-declared truncation, does to the rule that consumed it: allow (default) treats the content as not matching, block matches it as classification_error. | Enum: [allow block] Optional: {} |
PersistenceSpec
PersistenceSpec defines persistence configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled indicates if persistence is enabled | true | Optional: {} |
storageClassName string | StorageClassName is the storage class name | standard | Optional: {} |
accessMode PersistentVolumeAccessMode | AccessMode is the access mode | ReadWriteOnce | Optional: {} |
size string | Size is the storage size | 10Gi | Optional: {} |
existingClaim string | ExistingClaim is an existing PVC to use | Optional: {} | |
annotations object (keys:string, values:string) | Annotations for the PVC | Optional: {} |
PortSpec
PortSpec defines a service port configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
port integer | Port is the service port | Maximum: 65535 Minimum: 1 Optional: {} | |
targetPort integer | TargetPort is the container port | Maximum: 65535 Minimum: 1 Optional: {} | |
protocol Protocol | Protocol is the port protocol | TCP | Optional: {} |
ProbeSpec
ProbeSpec defines probe configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled indicates if the probe is enabled | true | Optional: {} |
initialDelaySeconds integer | InitialDelaySeconds before probe starts | Optional: {} | |
periodSeconds integer | PeriodSeconds between probes | Optional: {} | |
timeoutSeconds integer | TimeoutSeconds for probe | Optional: {} | |
failureThreshold integer | FailureThreshold for probe | Optional: {} |
PromptGuardConfig
PromptGuardConfig defines prompt guard configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
backend RemoteClassifierBackendConfig | Backend selects a named external classifier and its typed result contract. | Optional: {} | |
max_sequence_length integer | MaxSequenceLength limits the total tokenized input, including special tokens. Omission or zero retains the 512-token budget. The model loader validates the requested budget against the loaded model's capacity. | Minimum: 0 Optional: {} | |
window PromptGuardWindowConfig | Window enables explicit scanning of all input tokens. Omission or null keeps whole-input inference. Only the local model supports it. | Optional: {} | |
enabled boolean | true | Optional: {} | |
model_id string | ModelID binds the guard to one model; empty runs the decision model. | Optional: {} | |
threshold string | Jailbreak detection threshold (0.0-1.0). Stored as string to avoid float precision issues; empty takes the threshold calibrated for the model. | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} | |
use_cpu boolean | true | Optional: {} | |
jailbreak_mapping_path string | Optional: {} | ||
positive_labels string array | PositiveLabels lists the jailbreak_mapping labels that count as unsafe, for a custom backend whose positive class isn't named "jailbreak" (e.g. "INJECTION", "malicious"). Defaults to ["jailbreak"] when unset. | Optional: {} | |
on_error string | OnError selects what a prompt-guard classifier failure does to the rule that failed to evaluate. "allow" (the default) tolerates the failure and treats the content as not matching; "block" treats it as a positive detection, because an inference failure means the content could not be verified safe. Without this field on the CRD the setting is pruned by the API server and an operator-managed deployment silently fails open. | Enum: [allow block] Optional: {} |
PromptGuardWindowConfig
PromptGuardWindowConfig scans original content tokens with overlap. The native tokenizer also checks that special tokens leave enough content room.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
size integer | Size is the inference window budget, including special tokens. | Minimum: 1 | |
overlap integer | Overlap counts content tokens shared by consecutive windows. | Minimum: 0 Optional: {} |
PrototypeScoringConfig
PrototypeScoringConfig overrides prototype-bank construction and scoring for one embedding-backed signal rule. Router-owned defaults apply only after translation; the operator preserves an absent override and an empty object.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled controls prototype clustering. False retains every candidate. | Optional: {} | |
cluster_similarity_threshold string | ClusterSimilarityThreshold is the clustering similarity threshold. Stored as a numeric string, like other fractional operator config fields. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} | |
max_prototypes integer | MaxPrototypes caps the number of cluster representatives. | Optional: {} | |
best_weight string | BestWeight weights the best prototype against the top-M mean. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} | |
top_m integer | TopM is the number of highest-scoring prototypes included in the mean. | Optional: {} | |
margin_threshold string | MarginThreshold is the minimum winner-versus-runner-up score margin. | Pattern: ^-?[0-9]+(\.[0-9]+)?$ Optional: {} |
QdrantCacheConfig
QdrantCacheConfig defines Qdrant cache backend configuration. Configure these settings when using Qdrant as the semantic cache backend.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
host string | Host is the Qdrant server hostname or IP address | Optional: {} | |
port integer | Port is the Qdrant gRPC port | 6334 | Maximum: 65535 Minimum: 1 Optional: {} |
api_key string | APIKey for Qdrant authentication | Optional: {} | |
use_tls boolean | UseTLS enables TLS for the Qdrant connection | false | Optional: {} |
collection_name string | CollectionName is the Qdrant collection to use for semantic cache | semantic_cache | Optional: {} |
connect_timeout integer | ConnectTimeout is the timeout in seconds for Qdrant connection | 10 | Optional: {} |
RedisCacheConfig
RedisCacheConfig defines Redis cache backend configuration. Configure these settings when using Redis as the semantic cache backend.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
connection RedisCacheConnection | Connection settings for Redis server | Optional: {} | |
index RedisCacheIndex | Index settings for Redis vector search | Optional: {} | |
search RedisCacheSearch | Search settings for Redis queries | Optional: {} | |
development RedisCacheDevelopment | Development settings for Redis cache | Optional: {} |
RedisCacheConnection
RedisCacheConnection defines Redis connection parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
host string | Host is the Redis server hostname or IP address Example: "redis.default.svc.cluster.local" | Optional: {} | |
port integer | Port is the Redis server port | 6379 | Maximum: 65535 Minimum: 1 Optional: {} |
database integer | Database is the Redis database number to use | 0 | Minimum: 0 Optional: {} |
password string | Password for Redis authentication (plaintext - consider using PasswordSecretRef instead) | Optional: {} | |
password_secret_ref SecretKeySelector | PasswordSecretRef references a Secret containing the Redis password Preferred over plaintext Password field for security | Optional: {} | |
timeout integer | Timeout for Redis operations in seconds | 30 | Minimum: 0 Optional: {} |
tls RedisCacheTLS | TLS configuration for secure Redis connections | Optional: {} |
RedisCacheDevelopment
RedisCacheDevelopment defines development-mode settings.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
drop_index_on_startup boolean | DropIndexOnStartup clears the index when router starts (for testing) | false | Optional: {} |
auto_create_index boolean | AutoCreateIndex automatically creates the index if it doesn't exist | true | Optional: {} |
RedisCacheIndex
RedisCacheIndex defines Redis vector index configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the Redis index | semantic_cache_idx | Optional: {} |
prefix string | Prefix for Redis keys | doc: | Optional: {} |
vector_field RedisCacheVectorField | VectorField configuration for embeddings | Optional: {} | |
index_type string | IndexType specifies the index algorithm Options: "HNSW" (recommended), "FLAT" | HNSW | Enum: [HNSW FLAT] Optional: {} |
params RedisCacheIndexParams | Params for HNSW index | Optional: {} |
RedisCacheIndexParams
RedisCacheIndexParams defines HNSW index parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
M integer | M is the number of bi-directional links per node Higher values = better recall, more memory | 16 | Minimum: 2 Optional: {} |
efConstruction integer | EfConstruction is the size of dynamic candidate list during construction Higher values = better quality, slower indexing | 64 | Minimum: 1 Optional: {} |
RedisCacheSearch
RedisCacheSearch defines Redis search parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
topk integer | TopK is the number of results to return from vector search | 1 | Minimum: 1 Optional: {} |
RedisCacheTLS
RedisCacheTLS defines TLS settings for Redis connections.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled controls whether to use TLS for Redis connection | false | Optional: {} |
cert_file string | CertFile is the path to client certificate file | Optional: {} | |
key_file string | KeyFile is the path to client key file | Optional: {} | |
ca_file string | CAFile is the path to CA certificate file | Optional: {} |
RedisCacheVectorField
RedisCacheVectorField defines vector field configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the vector field | embedding | Optional: {} |
dimension integer | Dimension of the embedding vectors For BERT: 384, for Qwen3: 1024, for Gemma: 768 | Minimum: 1 Optional: {} | |
metric_type string | MetricType for vector similarity Options: "COSINE", "IP" (inner product), "L2" (Euclidean) | COSINE | Enum: [COSINE IP L2] Optional: {} |
RemoteClassifierBackendConfig
RemoteClassifierBackendConfig is the shared remote-classifier block. How the remote is called (protocol), what shape it answers with (contract), which catalog entry it is (model) and how long to wait (deadline) are independent axes rather than one enumeration. It mirrors the router's backend block field for field so the operator passes it through unchanged.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
protocol string | Protocol is how the remote is called. | Enum: [http_classify http_chat] | |
contract string | Contract is the response shape the signal reads. Complexity reads two - score.v1, one regression number interpreted through each rule's boundaries, and label_distribution.v1, hard/easy/medium probabilities - so the router requires it there rather than guessing per request. PII reads token_spans.v1, entity spans with code-point offsets. Prompt guard http_chat reads label_decision.v1, a verdict without invented probability. | Enum: [score.v1 label_distribution.v1 token_spans.v1 label_decision.v1] Optional: {} | |
model string | Model is the name of an entry in the external model catalog. | MinLength: 1 | |
deadline_ms integer | DeadlineMs bounds one remote call. Defaults to the router's value. | Minimum: 1 Optional: {} |
ResourceConfig
ResourceConfig defines resource configuration for tracing
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
service_name string | vllm-semantic-router | Optional: {} | |
service_version string | v0.1.0 | Optional: {} | |
deployment_environment string | development | Optional: {} |
RouteConfig
RouteConfig defines OpenShift Route configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled specifies whether to create an OpenShift Route | false | Optional: {} |
hostname string | Hostname for the Route (optional - OpenShift generates if empty) | Optional: {} | |
tls RouteTLSConfig | TLS configuration for the Route | Optional: {} |
RouteTLSConfig
RouteTLSConfig defines TLS configuration for OpenShift Routes
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
termination string | Termination type (edge, passthrough, reencrypt) | edge | Enum: [edge passthrough reencrypt] Optional: {} |
insecureEdgeTerminationPolicy string | InsecureEdgeTerminationPolicy for HTTP traffic | Redirect | Enum: [Allow Redirect None] Optional: {} |
RuleCombinationConfig
RuleCombinationConfig defines how to combine multiple rule conditions
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
operator string | Operator specifies how to combine conditions: "AND", "OR", or "NOT". NOT is strictly unary: it takes exactly one child condition and negates its result. Compose NOR/NAND by nesting NOT around OR/AND. | Enum: [AND OR NOT] | |
on_unknown string | OnUnknown resolves a terminal unknown result after the rule tree is evaluated. | Enum: [no_match match fail_request] Optional: {} | |
conditions RuleConditionConfig array | Conditions is the list of rule references to evaluate |
RuleComposition
RuleComposition defines how to compose/filter rules based on other signals
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
operator string | Operator for combining conditions (AND, OR, NOT). NOT is strictly unary and negates its single child. | Enum: [AND OR NOT] | |
conditions CompositionCondition array | List of conditions that must be met |
RuleConditionConfig
RuleConditionConfig references a specific rule by type and name
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | Type specifies the signal or projection type referenced by this condition. | Enum: [keyword embedding domain fact_check user_feedback reask preference language context structure complexity modality authz jailbreak pii kb conversation event projection] | |
name string | Name is the name of the rule to reference |
SamplingConfig
SamplingConfig defines sampling configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | always_on | Optional: {} | |
rate string | Sampling rate (0.0-1.0). Stored as string to avoid float precision issues. | 1.0 | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} |
SemanticCacheConfig
SemanticCacheConfig defines semantic cache configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled controls whether semantic caching is active | true | Optional: {} |
backend_type string | BackendType specifies the cache backend to use Options: "memory" (default), "redis", "valkey", "milvus", "qdrant", "hybrid" | memory | Enum: [memory redis valkey milvus qdrant hybrid] Optional: {} |
similarity_threshold string | Similarity threshold for cache hits (0.0-1.0). Stored as string to avoid float precision issues. | 0.8 | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} |
max_entries integer | MaxEntries is the maximum number of cache entries (for memory/hybrid backends) | 1000 | Optional: {} |
ttl_seconds integer | TTLSeconds is the time-to-live for cache entries in seconds | 3600 | Optional: {} |
eviction_policy string | EvictionPolicy for in-memory cache ("fifo", "lru", "lfu") | fifo | Enum: [fifo lru lfu] Optional: {} |
redis RedisCacheConfig | Redis configuration (required when backend_type is "redis") | Optional: {} | |
valkey ValkeyCacheConfig | Valkey configuration (required when backend_type is "valkey") | Optional: {} | |
milvus MilvusCacheConfig | Milvus configuration (required when backend_type is "milvus") | Optional: {} | |
qdrant QdrantCacheConfig | Qdrant configuration (required when backend_type is "qdrant") | Optional: {} | |
embedding_model string | EmbeddingModel specifies which embedding model to use for semantic similarity Options: "mmbert" (default), "bert", "qwen3", "gemma" | mmbert | Enum: [bert qwen3 gemma mmbert] Optional: {} |
hnsw HNSWCacheConfig | HNSW configuration for hybrid/in-memory backends | Optional: {} |
SemanticRouter
SemanticRouter is the Schema for the semanticrouters API
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
apiVersion string | vllm.ai/v1alpha1 | ||
kind string | SemanticRouter | ||
metadata ObjectMeta | Refer to Kubernetes API documentation for fields of metadata. | ||
spec SemanticRouterSpec | |||
status SemanticRouterStatus |
SemanticRouterList
SemanticRouterList contains a list of SemanticRouter
| Field | Description | Default | Validation |
|---|---|---|---|
apiVersion string | vllm.ai/v1alpha1 | ||
kind string | SemanticRouterList | ||
metadata ListMeta | Refer to Kubernetes API documentation for fields of metadata. | ||
items SemanticRouter array |
SemanticRouterSpec
SemanticRouterSpec defines the desired state of SemanticRouter
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
image ImageSpec | Image configuration | Optional: {} | |
replicas integer | Number of replicas | 1 | Minimum: 0 Optional: {} |
imagePullSecrets LocalObjectReference array | ImagePullSecrets for private registries | Optional: {} | |
serviceAccount ServiceAccountSpec | ServiceAccount configuration | Optional: {} | |
service ServiceSpec | Service configuration | Optional: {} | |
resources ResourceRequirements | Resource requirements | Optional: {} | |
persistence PersistenceSpec | Persistence configuration | Optional: {} | |
config ConfigSpec | Configuration overrides merged into the canonical v0.3 config.yaml. Router-wide runtime overrides land under config.global.router/services/stores/ integrations/model_catalog, with model-backed modules nested under config.global.model_catalog.modules. Provider defaults land under config.providers.defaults. | Optional: {} | |
toolsDb ToolEntry array | Tools database configuration | Optional: {} | |
vllmEndpoints VLLMEndpointSpec array | VLLMEndpoints is a Kubernetes-native backend discovery adapter. It generates canonical config.providers.models[].backend_refs and config.routing.modelCards entries. | Optional: {} | |
autoscaling AutoscalingSpec | Autoscaling configuration | Optional: {} | |
startupProbe ProbeSpec | Probes configuration | Optional: {} | |
livenessProbe ProbeSpec | Optional: {} | ||
readinessProbe ProbeSpec | Optional: {} | ||
securityContext SecurityContext | Security context | Optional: {} | |
podSecurityContext PodSecurityContext | Pod security context | Optional: {} | |
podAnnotations object (keys:string, values:string) | Pod annotations | Optional: {} | |
nodeSelector object (keys:string, values:string) | Node selector | Optional: {} | |
tolerations Toleration array | Tolerations | Optional: {} | |
affinity Affinity | Affinity | Optional: {} | |
env EnvVar array | Environment variables | Optional: {} | |
args string array | Router arguments, after the gateway mode flags the Operator passes (-gateway=standalone -listener-address=0.0.0.0, or -gateway=extproc). The gateway mode follows spec.gateway, so args may not set those flags. | MaxItems: 64 items:MaxLength: 4096 Optional: {} | |
gateway GatewaySpec | Gateway selects what serves client traffic. Omitted, the Router runs standalone: it serves the OpenAI-compatible API on port 8801 itself, with no Envoy, and the Service exposes that port. With existingRef, the Router serves ext_proc gRPC on port 50051 for that Gateway, whose ext_proc policy and routes you manage. | Optional: {} | |
openshift OpenShiftSpec | OpenShift-specific features | Optional: {} | |
ingress IngressSpec | Ingress configuration | Optional: {} |
SemanticRouterStatus
SemanticRouterStatus defines the observed state of SemanticRouter
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
conditions Condition array | Conditions represent the latest available observations of the SemanticRouter's state | Optional: {} | |
observedGeneration integer | ObservedGeneration reflects the generation of the most recently observed SemanticRouter | Optional: {} | |
replicas integer | Replicas is the current number of replicas | Optional: {} | |
readyReplicas integer | ReadyReplicas is the number of ready replicas | Optional: {} | |
phase string | Phase represents the current phase of the SemanticRouter | Optional: {} | |
gatewayMode string | GatewayMode is what serves client traffic: standalone (the Router's own listener on port 8801) or gateway-integration (the Gateway in spec.gateway, with the Router serving ext_proc) | Optional: {} | |
openshiftFeatures OpenShiftFeaturesStatus | OpenShiftFeatures tracks OpenShift-specific feature status | Optional: {} |
ServiceAccountSpec
ServiceAccountSpec defines service account configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
create boolean | Create specifies whether to create a service account | true | Optional: {} |
name string | Name of the service account to use | Optional: {} | |
annotations object (keys:string, values:string) | Annotations for the service account | Optional: {} |
ServiceBackend
ServiceBackend defines a direct Kubernetes service backend
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Service name | MinLength: 1 | |
namespace string | Service namespace (defaults to same namespace) | Optional: {} | |
port integer | Service port | Minimum: 1 |
ServiceSpec
ServiceSpec defines the service configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type ServiceType | Type is the service type | ClusterIP | Enum: [ClusterIP NodePort LoadBalancer] Optional: {} |
grpc PortSpec | GRPC port configuration | Optional: {} | |
api PortSpec | API port configuration | Optional: {} | |
metrics MetricsPortSpec | Metrics port configuration | Optional: {} |
StreamedBodyConfig
StreamedBodyConfig defines streamed request body handling.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled accumulates request body chunks before routing at end-of-stream. | Optional: {} | |
max_bytes integer | MaxBytes caps the accumulated body size. A larger body is rejected and the ExtProc stream ends; the downstream response follows the gateway's ExtProc failure policy. Zero disables the limit. | Minimum: 0 Optional: {} | |
timeout_sec integer | TimeoutSec caps how long body accumulation may take. A slower body is rejected and the ExtProc stream ends; the downstream response follows the gateway's ExtProc failure policy. Zero disables the limit. | Minimum: 0 Optional: {} |
Tool
Tool defines a tool function
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | Enum: [function] Optional: {} | ||
function ToolFunction | Optional: {} |
ToolEntry
ToolEntry defines a tool entry in the tools database
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
tool Tool | Optional: {} | ||
description string | Optional: {} | ||
category string | Optional: {} | ||
tags string array | Optional: {} |
ToolFunction
ToolFunction defines a tool function details
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Optional: {} | ||
description string | Optional: {} | ||
parameters ToolParameters | Optional: {} |
ToolParameters
ToolParameters defines tool function parameters
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | Optional: {} | ||
properties JSON | Type: object Optional: {} | ||
required string array | Optional: {} |
ToolsConfig
ToolsConfig defines tools configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | true | Optional: {} | |
top_k integer | 3 | Optional: {} | |
similarity_threshold string | Similarity threshold for tool selection (0.0-1.0). Stored as string to avoid float precision issues. | 0.2 | Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$ Optional: {} |
tools_db_path string | config/tools_db.json | Optional: {} | |
fallback_to_empty boolean | true | Optional: {} |
TracingConfig
TracingConfig defines tracing configuration
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | false | Optional: {} | |
provider string | opentelemetry | Optional: {} | |
exporter ExporterConfig | Optional: {} | ||
sampling SamplingConfig | Optional: {} | ||
resource ResourceConfig | Optional: {} |
VLLMBackend
VLLMBackend specifies how to reach the vLLM service
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
type string | Type of backend: kserve, llamastack, or service | Enum: [kserve llamastack service] | |
inferenceServiceName string | For type=kserve: InferenceService name for auto-discovery | Optional: {} | |
discoveryLabels object (keys:string, values:string) | For type=llamastack: Labels to match services | Optional: {} | |
service ServiceBackend | For type=service: Direct service configuration | Optional: {} |
VLLMEndpointSpec
VLLMEndpointSpec defines a vLLM model backend endpoint
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the backend ref generated under config.providers.models[].backend_refs | MinLength: 1 | |
model string | Model name as reported by vLLM (e.g., "Model-A", "llama3-8b") | MinLength: 1 | |
catalog string | Catalog optionally selects a repository built-in Model Card. Model remains the request-facing alias. | Optional: {} | |
reasoning ModelReasoningSpec | Reasoning optionally selects a built-in family or defines inline wire behavior for this self-hosted model. Catalog-backed models normally omit it. | Optional: {} | |
loras LoRAAdapterSpec array | LoRAs declares the LoRA adapters exposed for this logical model in routing.modelCards. | MaxItems: 50 Optional: {} | |
backend VLLMBackend | Backend configuration | ||
weight integer | Weight for load balancing (default: 1) | 1 | Optional: {} |
ValkeyCacheConfig
ValkeyCacheConfig defines Valkey cache backend configuration. Configure these settings when using Valkey as the semantic cache backend.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
connection ValkeyCacheConnection | Connection settings for Valkey server | Optional: {} | |
index ValkeyCacheIndex | Index settings for Valkey vector search | Optional: {} | |
search ValkeyCacheSearch | Search settings for Valkey queries | Optional: {} | |
development ValkeyCacheDevelopment | Development settings for Valkey cache | Optional: {} |
ValkeyCacheConnection
ValkeyCacheConnection defines Valkey connection parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
host string | Host is the Valkey server hostname or IP address Example: "valkey.default.svc.cluster.local" | Optional: {} | |
port integer | Port is the Valkey server port | 6379 | Maximum: 65535 Minimum: 1 Optional: {} |
database integer | Database is the Valkey database number to use | 0 | Minimum: 0 Optional: {} |
password string | Password for Valkey authentication (plaintext - consider using PasswordSecretRef instead) | Optional: {} | |
password_secret_ref SecretKeySelector | PasswordSecretRef references a Secret containing the Valkey password Preferred over plaintext Password field for security | Optional: {} | |
timeout integer | Timeout for Valkey operations in seconds | 30 | Minimum: 0 Optional: {} |
tls ValkeyCacheTLS | TLS configuration for secure Valkey connections | Optional: {} |
ValkeyCacheDevelopment
ValkeyCacheDevelopment defines development-mode settings.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
drop_index_on_startup boolean | DropIndexOnStartup clears the index when router starts (for testing) | false | Optional: {} |
auto_create_index boolean | AutoCreateIndex automatically creates the index if it doesn't exist | true | Optional: {} |
ValkeyCacheIndex
ValkeyCacheIndex defines Valkey vector index configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the Valkey index | semantic_cache_idx | Optional: {} |
prefix string | Prefix for Valkey keys | doc: | Optional: {} |
vector_field ValkeyCacheVectorField | VectorField configuration for embeddings | Optional: {} | |
index_type string | IndexType specifies the index algorithm Options: "HNSW" (recommended), "FLAT" | HNSW | Enum: [HNSW FLAT] Optional: {} |
params ValkeyCacheIndexParams | Params for HNSW index | Optional: {} |
ValkeyCacheIndexParams
ValkeyCacheIndexParams defines HNSW index parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
M integer | M is the number of bi-directional links per node Higher values = better recall, more memory | 16 | Minimum: 2 Optional: {} |
efConstruction integer | EfConstruction is the size of dynamic candidate list during construction Higher values = better quality, slower indexing | 64 | Minimum: 1 Optional: {} |
ValkeyCacheSearch
ValkeyCacheSearch defines Valkey search parameters.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
topk integer | TopK is the number of results to return from vector search | 1 | Minimum: 1 Optional: {} |
ValkeyCacheTLS
ValkeyCacheTLS defines TLS settings for Valkey connections.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
enabled boolean | Enabled controls whether to use TLS for Valkey connection | false | Optional: {} |
cert_file string | CertFile is the path to client certificate file | Optional: {} | |
key_file string | KeyFile is the path to client key file | Optional: {} | |
ca_file string | CAFile is the path to CA certificate file | Optional: {} |
ValkeyCacheVectorField
ValkeyCacheVectorField defines vector field configuration.
Appears in:
| Field | Description | Default | Validation |
|---|---|---|---|
name string | Name of the vector field | embedding | Optional: {} |
dimension integer | Dimension of the embedding vectors For BERT: 384, for Qwen3: 1024, for Gemma: 768 | Minimum: 1 Optional: {} | |
metric_type string | MetricType for vector similarity Options: "COSINE", "IP" (inner product), "L2" (Euclidean) | COSINE | Enum: [COSINE IP L2] Optional: {} |