跳到主要内容
版本:最新版(未发布)

API Reference

Packages​

vllm.ai/v1alpha1​

Package v1alpha1 contains API Schema definitions for the vllm v1alpha1 API group

Resource Types​

APIConfig​

APIConfig defines API configuration

Appears in:

FieldDescriptionDefaultValidation
batch_classification BatchClassificationConfigOptional: {}

AutoscalingSpec​

AutoscalingSpec defines autoscaling configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled indicates if HPA is enabledfalseOptional: {}
minReplicas integerMinReplicas is the minimum number of replicas1Optional: {}
maxReplicas integerMaxReplicas is the maximum number of replicas10Optional: {}
targetCPUUtilizationPercentage integerTargetCPUUtilizationPercentage is the target CPU percentage80Optional: {}
targetMemoryUtilizationPercentage integerTargetMemoryUtilizationPercentage is the target memory percentageOptional: {}

BatchClassificationConfig​

BatchClassificationConfig defines batch classification configuration

Appears in:

FieldDescriptionDefaultValidation
max_batch_size integer100Optional: {}
concurrency_threshold integer5Optional: {}
max_concurrency integer8Optional: {}
metrics BatchMetricsConfigOptional: {}

BatchMetricsConfig​

BatchMetricsConfig defines batch classification metrics configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleantrueOptional: {}
detailed_goroutine_tracking booleantrueOptional: {}
high_resolution_timing booleanfalseOptional: {}
sample_rate stringSample rate for metrics (0.0-1.0). Stored as string to avoid float precision issues.1.0Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
duration_buckets string arrayDuration buckets for histograms. Stored as strings to avoid float precision issues.
Example: ["0.001", "0.005", "0.01", "0.025", "0.05", "0.1", "0.25", "0.5", "1", "2.5", "5", "10", "30"]
Optional: {}
size_buckets integer arrayOptional: {}

CategoryModelConfig​

CategoryModelConfig defines category model configuration

Appears in:

FieldDescriptionDefaultValidation
model_id stringOptional: {}
threshold stringClassification threshold (0.0-1.0). Stored as string to avoid float precision issues.Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
use_cpu booleanOptional: {}
category_mapping_path stringOptional: {}

ClassifierConfig​

ClassifierConfig defines classifier configuration

Appears in:

FieldDescriptionDefaultValidation
category_model CategoryModelConfigOptional: {}
pii_model PIIModelConfigOptional: {}

ComplexityCandidates​

ComplexityCandidates defines candidate examples for complexity classification

Appears in:

FieldDescriptionDefaultValidation
candidates string arrayList of candidate phrases or examples

ComplexityModelConfig​

ComplexityModelConfig configures how the complexity signal produces its score. It mirrors global.model_catalog.modules.complexity in the router config and is passed through field for field.

The contract requirement sits here rather than on RemoteClassifierBackendConfig because it is a property of this consumer, not of the block: complexity reads two response shapes, so guessing wrong would surface per request instead of at admission. A consumer that reads one shape

  • categories does - keeps the field optional and defaults it.

Appears in:

FieldDescriptionDefaultValidation
backend RemoteClassifierBackendConfigBackend names a remote scoring model. Its absence keeps local prototype
scoring; when set, the signal never reads the rules' hard/easy
candidates. It sits on the module rather than on a rule because routing
signals are replaced wholesale per recipe, so a per-rule backend would
vanish under any recipe that did not repeat it.
Optional: {}

ComplexityRulesConfig​

ComplexityRulesConfig defines complexity-based signal classification.

The CEL rules below reject at admission the boundary combinations the Router refuses at config load. Without them the API server accepts the object and the Router crashloops on it, which turns a typo into an outage instead of a rejected write. They are per-object and static; anything needing the model catalog - whether backend.model resolves, for instance - stays with the Router's validator, which remains the single source of truth for the rest.

Appears in:

FieldDescriptionDefaultValidation
prototype_scoring PrototypeScoringConfigPrototypeScoring replaces the family prototype-scoring configuration for
this rule. Absence inherits the family; a present object is a complete
override, including an empty object. Defaults remain Router-owned.
Optional: {}
name stringName of the complexity rule (e.g., "code-complexity", "reasoning-complexity")
description stringDescription of what this rule classifiesOptional: {}
threshold stringThreshold for the local prototype-scoring path (0.0-1.0), stored as a
string to avoid float precision issues. The local margin is
hard-minus-easy and centred on zero, so the threshold is symmetric: a
margin above it is "hard", below its negative is "easy", and in between
is "medium". It does not apply under a score.v1 backend, whose score is
in the model's own units; state a boundary pair instead.
Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
hard_above stringHardAbove and EasyBelow are the two cut points for a score where a
higher value is harder, in the scoring model's own units - so no [0,1]
pattern applies and negative values are valid. Both are required
together, and the pair is mutually exclusive with Threshold and with
HardBelow/EasyAbove. Stored as strings to avoid float precision issues.
Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}
easy_below stringEasyBelow is the lower cut point of the harder-when-higher pair: a score
below it is "easy", and anything between EasyBelow and HardAbove is
"medium". It must be below HardAbove, and both are required together.
Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}
hard_below stringHardBelow and EasyAbove are the pair for a score where a lower value is
harder - a model predicting the chance of a correct answer, say. They
require a score.v1 backend: the local margin is harder-when-higher by
construction, and inverting it locally means swapping the candidate
lists. Stored as strings to avoid float precision issues.
Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}
easy_above stringEasyAbove is the upper cut point of the harder-when-lower pair: a score
above it is "easy", and anything between HardBelow and EasyAbove is
"medium". It must be above HardBelow, and both are required together.
Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}
hard ComplexityCandidatesHard candidates represent complex/difficult examples. Read only by the
local path; a remote backend never consults them, so they are optional.
Optional: {}
easy ComplexityCandidatesEasy candidates represent simple/easy examples. Read only by the local
path; a remote backend never consults them, so they are optional.
Optional: {}
composer RuleCompositionComposer allows filtering based on other signals (e.g., only apply this rule if domain:medical)Optional: {}

CompositionCondition​

CompositionCondition defines a single composition condition

Appears in:

FieldDescriptionDefaultValidation
type stringType of signal to check (e.g., "domain", "language", "category")
name stringName of the specific signal/rule value to match

ConfigSpec​

ConfigSpec defines the semantic router configuration

Appears in:

FieldDescriptionDefaultValidation
routing JSONRouting contains canonical v0.3 routing configuration under config.routing.
It is intentionally preserved as an object so the operator can pass through
the router-owned signal, projection, decision, and algorithm contract without
lagging behind every router schema addition.
Type: object
Optional: {}
decision_model DecisionModelBindingDecisionModel selects the declared deployment that answers default judgment tasks.
Omitted uses the Router's primary deployment. Artifact identity belongs in
model_deployments; the reference never infers a model family or alias.
Optional: {}
model_deployments JSONModelDeployments contains canonical global.model_catalog.deployments.
The router validates provider, device, precision and task compatibility.
Type: object
Optional: {}
model_admission JSONModelAdmission contains canonical global.model_catalog.admission budgets.
Keys name deployments or the router's existing admission consumers.
Type: object
Optional: {}
embedding_models EmbeddingModelsConfigEmbedding models configuration (qwen3, gemma, mmbert)Optional: {}
response_cache SemanticCacheConfigResponse cache configuration.Optional: {}
semantic_cache SemanticCacheConfigSemanticCache is the deprecated response-cache field.Optional: {}
tools ToolsConfigTools configurationOptional: {}
prompt_guard PromptGuardConfigPrompt guard configurationOptional: {}
classifier ClassifierConfigClassifier configurationOptional: {}
complexity_rules ComplexityRulesConfig arrayComplexity rules for complexity-aware routingOptional: {}
complexity_model ComplexityModelConfigComplexityModel says how the complexity signal produces its score.
Absent, the signal scores locally against each rule's hard/easy
candidates. With a backend, a remote model produces the score and the
candidates are never read. Mirrors
global.model_catalog.modules.complexity in the router config.
Optional: {}
external_models ExternalModelConfig arrayExternalModels declares the remote models that classifier backends
(classifier.pii.backend.model, complexity_model.backend.model) and
the prompt guard protocol refer to by name. Mirrors
global.model_catalog.external[] in the router config field for field;
the router's own validator decides whether a backend resolves against it.
Optional: {}
strategy stringDecision routing strategy ("priority" for priority-based matching)Enum: [priority]
Optional: {}
decisions DecisionConfig arrayRouting decisions based on signals (domain, complexity, etc.)Optional: {}
reasoning_effort stringReasoningEffort is the default reasoning effort for model bindings that do
not select a different effort. The selected model family validates the
value because built-in and custom families may expose different ladders.
Optional: {}
api APIConfigAPI configurationOptional: {}
observability ObservabilityConfigObservability configurationOptional: {}
streamed_body StreamedBodyConfigStreamedBody enables streamed request body handling. Mirrors
global.router.streamed_body; the gateway must send bodies to ExtProc in
STREAMED or FullDuplexStreamed mode for it to take effect.
Optional: {}

DecisionConfig​

DecisionConfig defines a routing decision

Appears in:

FieldDescriptionDefaultValidation
name stringName is the unique identifier for this decision
description stringDescription provides information about what this decision handlesOptional: {}
priority integerPriority is used for decision ordering - higher priority decisions are evaluated firstOptional: {}
rules RuleCombinationConfigRules defines the combination of conditions using AND/OR logic
modelRefs ModelRefConfig arrayModelRefs contains model references for this decisionOptional: {}
preferred_endpoints string arrayPreferredEndpoints specifies which vLLM endpoints to prefer for this decisionOptional: {}
plugins RawExtension arrayPlugins contains policy configurations applied after rule matchingOptional: {}
algorithm JSONAlgorithm configures base model selection for this decision. It is
preserved as a router-owned object so supported algorithms can evolve
without requiring the operator CRD to duplicate every nested field.
Type: object
Optional: {}
reliability DecisionReliabilityConfigReliability overrides the timeouts and retries of the provider model
that serves this decision's requests
Optional: {}
fallback DecisionFallbackConfigFallback overrides the cross-model fallback policy for this decision's
requests, over the recipe's and the router's
Optional: {}

DecisionFallbackConfig​

DecisionFallbackConfig is a decision's fallback block, with the router configuration's field names and meaning. A field left out keeps the recipe's or the router's value.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled turns cross-model fallback on or off for this decisionOptional: {}
max_attempts integerMaxAttempts bounds the candidates tried, the first one includedMinimum: 0
Optional: {}
total_timeout stringTotalTimeout bounds the whole fallback chainPattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
per_attempt_timeout stringPerAttemptTimeout bounds each candidate's attemptPattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
retryable_status_codes integer arrayRetryableStatusCodes are the statuses that move on to the next candidateitems:Maximum: 599
items:Minimum: 100
Optional: {}

DecisionModelBinding​

DecisionModelBinding selects one canonical model deployment.

Appears in:

FieldDescriptionDefaultValidation
deployment stringDeployment is an exact key in model_deployments.MinLength: 1

DecisionReliabilityConfig​

DecisionReliabilityConfig is a decision's reliability block, with the router configuration's field names and meaning. Durations are Go durations such as "30s"; "0s" turns a timeout off.

Appears in:

FieldDescriptionDefaultValidation
total_timeout stringTotalTimeout bounds the whole call: every attempt and the responsePattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
per_try_timeout stringPerTryTimeout bounds each attempt until its response startsPattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
idle_timeout stringIdleTimeout bounds the wait for more of a streamed response
(standalone mode only)
Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
first_byte_timeout stringFirstByteTimeout bounds the wait for the first response byte
(standalone mode only)
Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
retry_count integerRetryCount is the number of retries after the first attemptMaximum: 5
Minimum: 0
Optional: {}
retry_on stringRetryOn adds retry conditions, with Envoy's names (5xx, reset, ...)Optional: {}
retriable_status_codes integer arrayRetriableStatusCodes adds statuses retried under retriable-status-codesitems:Maximum: 599
items:Minimum: 100
Optional: {}
retry_back_off_base stringRetryBackOffBase is the base of the randomized exponential wait between
retries (standalone mode only)
Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
retry_back_off_max stringRetryBackOffMax caps that wait (standalone mode only)Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}
retry_after_max stringRetryAfterMax honors a response's Retry-After up to this bound
(standalone mode only)
Pattern: ^([0-9]+(\.[0-9]+)?(ns|us|ms|s|m|h))+$
Optional: {}

EmbeddingEndpointConfig​

EmbeddingEndpointConfig defines an external OpenAI-compatible embedding endpoint.

Appears in:

FieldDescriptionDefaultValidation
base_url stringBaseURL is the base URL for the embedding endpoint, typically ending in /v1.Optional: {}
model stringModel is the embedding model name sent to the external provider.Optional: {}
api_key_env stringAPIKeyEnv names the environment variable containing the provider API key.Optional: {}
timeout_seconds integerTimeoutSeconds is the request timeout for embedding calls.Minimum: 0
Optional: {}
max_retries integerMaxRetries is the maximum number of retry attempts for embedding calls.Minimum: 0
Optional: {}
max_response_bytes integerMaxResponseBytes caps the size of each embedding response body.Minimum: 0
Optional: {}
dimensions integerDimensions requests a provider-side output dimension when supported.Minimum: 1
Optional: {}

EmbeddingModelsConfig​

EmbeddingModelsConfig defines configuration for embedding models

Appears in:

FieldDescriptionDefaultValidation
qwen3_model_path stringPath to Qwen3-Embedding-0.6B model directory
Qwen3 provides 32K context and high quality embeddings (1024 dimensions)
Optional: {}
mmbert_model_path stringPath to mmBERT 2D Matryoshka embedding model directory
Supports layer early exit (3/6/11/22) and dimension reduction (64-768)
Optional: {}
use_cpu booleanUse CPU for inference (default: true)trueOptional: {}
embedding_config HNSWEmbeddingConfigEmbedding configuration for embedding-based classificationOptional: {}
endpoint EmbeddingEndpointConfigEndpoint configures an external embedding provider endpoint.
The API key should be injected into the semantic router pod environment
and referenced by APIKeyEnv rather than stored directly in the CR.
Optional: {}

ExporterConfig​

ExporterConfig defines exporter configuration

Appears in:

FieldDescriptionDefaultValidation
type stringotlpOptional: {}
endpoint stringjaeger:4317Optional: {}
insecure booleantrueOptional: {}

ExternalModelConfig​

ExternalModelConfig is one entry of global.model_catalog.external[]: a remote model a classifier backend or the prompt guard can name. Field names are the router's YAML keys so the generic typed conversion carries them unchanged.

Appears in:

FieldDescriptionDefaultValidation
name stringName is the catalog name a backend block refers to in its model field.MinLength: 1
model_role stringModelRole is what the model is used for; classifier backends require
"classification", the prompt guard protocol requires "guardrail".
MinLength: 1
llm_model_name stringModelName is the model identifier the remote service expects, and the
value a token_spans.v1 envelope's model member must equal.
MinLength: 1
llm_endpoint ExternalModelEndpointEndpoint is where the remote model is reached.
llm_timeout_seconds integerTimeoutSeconds bounds one call when the backend block sets no deadline.Minimum: 1
Optional: {}

ExternalModelEndpoint​

ExternalModelEndpoint is the address of a remote classification model.

Appears in:

FieldDescriptionDefaultValidation
address stringMinLength: 1
port integerMaximum: 65535
Minimum: 1
protocol stringEnum: [http https]
Optional: {}

GatewayReference​

GatewayReference references an existing Gateway

Appears in:

FieldDescriptionDefaultValidation
name stringName of the GatewayMinLength: 1
namespace stringNamespace of the GatewayMinLength: 1

GatewaySpec​

GatewaySpec defines Gateway API integration configuration

Appears in:

FieldDescriptionDefaultValidation
existingRef GatewayReferenceExistingRef references an existing Gateway that calls the Router over
ext_proc. Setting it selects extproc mode.
Optional: {}

HNSWCacheConfig​

HNSWCacheConfig defines HNSW index configuration for hybrid/in-memory backends.

Appears in:

FieldDescriptionDefaultValidation
use_hnsw booleanUseHNSW enables HNSW indexing for faster similarity searchfalseOptional: {}
hnsw_m integerM is the number of bi-directional links per node16Minimum: 2
Optional: {}
hnsw_ef_construction integerEfConstruction is the size of dynamic candidate list during construction200Minimum: 1
Optional: {}
max_memory_entries integerMaxMemoryEntries limits in-memory entries for hybrid backend1000Minimum: 0
Optional: {}

HNSWEmbeddingConfig​

HNSWEmbeddingConfig contains settings for embedding classification with HNSW indexing

Appears in:

FieldDescriptionDefaultValidation
backend stringBackend selects the embedding provider backend: the built-in model
runtime (the default) or an external OpenAI-compatible endpoint.
Enum: [model_runtime openai_compatible]
Optional: {}
model_type stringModelType specifies which embedding model to use
Options: "qwen3" (1024-dim, 32K context), "gemma" (768-dim, 8K context), "mmbert" (64-768-dim, multilingual), "remote" (external provider)
Enum: [qwen3 gemma mmbert remote]
Optional: {}
preload_embeddings booleanPreloadEmbeddings enables precomputing candidate embeddings at startuptrueOptional: {}
target_dimension integerTargetDimension is the embedding dimension to use (default: 768)
For mmBERT, supported local dimensions are 64, 128, 256, 512, 768.
External providers may use other positive dimensions such as 1024, 1536, or 3072.
Minimum: 1
Optional: {}
target_layer integerTargetLayer controls mmBERT early exit and is used only when ModelType is "mmbert".
Lower layers reduce encoder work but may reduce quality; layer 22 uses the full encoder depth.
Evaluate the latency and quality trade-off on representative deployment data.
Enum: [3 6 11 22]
Optional: {}
enable_soft_matching booleanEnableSoftMatching allows below-threshold matches when no rule meets its threshold.falseOptional: {}
min_score_threshold stringMinScoreThreshold for matching (0.0-1.0). Stored as string to avoid float precision issues.0.5Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}

ImageSpec​

ImageSpec defines the container image configuration

Appears in:

FieldDescriptionDefaultValidation
repository stringRepository is the container image repositoryghcr.io/vllm-project/semantic-router/vllm-srOptional: {}
tag stringTag is the container image taglatestOptional: {}
pullPolicy PullPolicyPullPolicy is the image pull policyIfNotPresentEnum: [Always Never IfNotPresent]
Optional: {}
imageRegistry stringImageRegistry is an optional registry prefixOptional: {}

IngressHost​

IngressHost defines an ingress host

Appears in:

FieldDescriptionDefaultValidation
host stringOptional: {}
paths IngressPath arrayOptional: {}

IngressPath​

IngressPath defines an ingress path

Appears in:

FieldDescriptionDefaultValidation
path stringOptional: {}
pathType stringOptional: {}
servicePort integerOptional: {}

IngressSpec​

IngressSpec defines ingress configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled indicates if ingress is enabledfalseOptional: {}
className stringClassName is the ingress class nameOptional: {}
annotations object (keys:string, values:string)Annotations for ingressOptional: {}
hosts IngressHost arrayHosts configurationOptional: {}
tls IngressTLS arrayTLS configurationOptional: {}

IngressTLS​

IngressTLS defines ingress TLS configuration

Appears in:

FieldDescriptionDefaultValidation
secretName stringOptional: {}
hosts string arrayOptional: {}

LoRAAdapterSpec​

LoRAAdapterSpec defines one LoRA adapter exposed by a VLLMEndpoint model.

Appears in:

FieldDescriptionDefaultValidation
name stringName is the unique adapter identifier referenced by decision.modelRefs[].lora_name.MaxLength: 100
MinLength: 1
description stringDescription provides a short human-readable summary for UI and docs surfaces.MaxLength: 500
Optional: {}

MetricsPortSpec​

MetricsPortSpec extends PortSpec with enable flag

Appears in:

FieldDescriptionDefaultValidation
port integerPort is the service portMaximum: 65535
Minimum: 1
Optional: {}
targetPort integerTargetPort is the container portMaximum: 65535
Minimum: 1
Optional: {}
protocol ProtocolProtocol is the port protocolTCPOptional: {}
enabled booleanEnabled indicates if metrics should be exposedtrueOptional: {}

MilvusCacheAuth​

MilvusCacheAuth defines Milvus authentication.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled controls whether to use authenticationfalseOptional: {}
username stringUsername for Milvus authenticationOptional: {}
password stringPassword for Milvus authentication (plaintext - consider using PasswordSecretRef instead)Optional: {}
password_secret_ref SecretKeySelectorPasswordSecretRef references a Secret containing the Milvus password
Preferred over plaintext Password field for security
Optional: {}

MilvusCacheCollection​

MilvusCacheCollection defines Milvus collection configuration.

Appears in:

FieldDescriptionDefaultValidation
name stringName of the Milvus collectionsemantic_cacheOptional: {}
description stringDescription of the collectionSemantic cache for LLM request-response pairsOptional: {}
vector_field MilvusCacheVectorFieldVectorField configuration for embeddingsOptional: {}
index MilvusCacheCollectionIndexIndex configuration for the collectionOptional: {}

MilvusCacheCollectionIndex​

MilvusCacheCollectionIndex defines collection index settings.

Appears in:

FieldDescriptionDefaultValidation
type stringType of index algorithmHNSWEnum: [HNSW IVF_FLAT IVF_SQ8 IVF_PQ]
Optional: {}
params MilvusCacheIndexParamsParams for the indexOptional: {}

MilvusCacheConfig​

MilvusCacheConfig defines Milvus cache backend configuration. Configure these settings when using Milvus as the semantic cache backend.

Appears in:

FieldDescriptionDefaultValidation
connection MilvusCacheConnectionConnection settings for Milvus serverOptional: {}
collection MilvusCacheCollectionCollection settings for MilvusOptional: {}
search MilvusCacheSearchSearch settings for Milvus queriesOptional: {}
development MilvusCacheDevelopmentDevelopment settings for Milvus cacheOptional: {}

MilvusCacheConnection​

MilvusCacheConnection defines Milvus connection parameters.

Appears in:

FieldDescriptionDefaultValidation
host stringHost is the Milvus server hostname or IP addressOptional: {}
port integerPort is the Milvus server port19530Maximum: 65535
Minimum: 1
Optional: {}
database stringDatabase name in Milvussemantic_router_cacheOptional: {}
timeout integerTimeout for Milvus operations in seconds30Minimum: 0
Optional: {}
auth MilvusCacheAuthAuth configuration for Milvus authenticationOptional: {}
tls MilvusCacheTLSTLS configuration for secure Milvus connectionsOptional: {}

MilvusCacheDevelopment​

MilvusCacheDevelopment defines development-mode settings.

Appears in:

FieldDescriptionDefaultValidation
drop_collection_on_startup booleanDropCollectionOnStartup clears the collection when router starts (for testing)falseOptional: {}
auto_create_collection booleanAutoCreateCollection automatically creates the collection if it doesn't existtrueOptional: {}

MilvusCacheIndexParams​

MilvusCacheIndexParams defines index parameters.

Appears in:

FieldDescriptionDefaultValidation
M integerM is the number of bi-directional links for HNSW16Minimum: 2
Optional: {}
efConstruction integerEfConstruction for HNSW index building64Minimum: 1
Optional: {}

MilvusCacheSearch​

MilvusCacheSearch defines Milvus search parameters.

Appears in:

FieldDescriptionDefaultValidation
params MilvusCacheSearchParamsParams for search operationsOptional: {}
topk integerTopK is the number of results to return10Minimum: 1
Optional: {}
consistency_level stringConsistencyLevel for search operations
Options: "Strong", "Session", "Bounded", "Eventually"
SessionEnum: [Strong Session Bounded Eventually]
Optional: {}

MilvusCacheSearchParams​

MilvusCacheSearchParams defines search-time parameters.

Appears in:

FieldDescriptionDefaultValidation
ef integerEf is the search-time HNSW parameter64Minimum: 1
Optional: {}

MilvusCacheTLS​

MilvusCacheTLS defines TLS settings for Milvus connections.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled controls whether to use TLSfalseOptional: {}
cert_file stringCertFile is the path to client certificate fileOptional: {}
key_file stringKeyFile is the path to client key fileOptional: {}
ca_file stringCAFile is the path to CA certificate fileOptional: {}

MilvusCacheVectorField​

MilvusCacheVectorField defines vector field configuration.

Appears in:

FieldDescriptionDefaultValidation
name stringName of the vector fieldembeddingOptional: {}
dimension integerDimension of the embedding vectorsMinimum: 1
Optional: {}
metric_type stringMetricType for vector similarity
Options: "IP" (inner product), "L2", "COSINE"
IPEnum: [IP L2 COSINE]
Optional: {}

ModelReasoningSpec​

ModelReasoningSpec selects a catalog reasoning family or defines the request projection for a custom self-hosted model. Family and inline fields are mutually exclusive and are validated by the Router's canonical compiler.

Appears in:

FieldDescriptionDefaultValidation
family stringOptional: {}
type stringEnum: [chat_template_kwargs reasoning_effort reasoning_mode top_level_reasoning_effort]
Optional: {}
parameter stringOptional: {}
activationParameter stringOptional: {}
effortFlags object (keys:string, values:string)EffortFlags maps a logical effort to a boolean chat-template parameter.Optional: {}
levels string arrayOptional: {}
default stringOptional: {}
modes string arrayitems:Enum: [enabled disabled adaptive]
Optional: {}
defaultMode stringEnum: [enabled disabled adaptive]
Optional: {}
disabled stringOptional: {}

ModelRefConfig​

ModelRefConfig defines a model reference for routing

Appears in:

FieldDescriptionDefaultValidation
model stringModel name to route to
lora_name stringLoRAName is the optional LoRA adapter nameOptional: {}
use_reasoning booleanUseReasoning enables reasoning mode for this modelOptional: {}
reasoning_mode stringReasoningMode selects the model's reasoning activation mode when the
family supports more than a boolean switch.
Enum: [enabled disabled adaptive]
Optional: {}
reasoning_effort stringReasoningEffort selects one of the model family's declared effort levels.Optional: {}

ObservabilityConfig​

ObservabilityConfig defines observability configuration

Appears in:

FieldDescriptionDefaultValidation
tracing TracingConfigOptional: {}

OpenShiftFeaturesStatus​

OpenShiftFeaturesStatus tracks OpenShift-specific feature status

Appears in:

FieldDescriptionDefaultValidation
routesEnabled booleanRoutesEnabled indicates if OpenShift Routes are enabled
routeHostname stringRouteHostname is the hostname of the created RouteOptional: {}

OpenShiftSpec​

OpenShiftSpec defines OpenShift-specific configuration

Appears in:

FieldDescriptionDefaultValidation
routes RouteConfigRoutes configuration for OpenShift RoutesOptional: {}

PIIModelConfig​

PIIModelConfig defines PII model configuration.

The contract rule sits on the consumer, as on ComplexityModelConfig: the shared backend block lists every contract any consumer reads, and each consumer narrows it to what it can parse, so a mismatch is refused at admission instead of by the router at load.

Appears in:

FieldDescriptionDefaultValidation
max_sequence_length integerMaxSequenceLength is the total tokenized input budget, including special
tokens. Omission or zero preserves the 512-token legacy limit.
Minimum: 0
Optional: {}
window PromptGuardWindowConfigWindow scans original content tokens with explicit overlap. Omission or
null leaves window selection unchanged; no CRD defaults are injected.
Optional: {}
model_id stringOptional: {}
threshold stringDetection threshold (0.0-1.0). Stored as string to avoid float precision issues.Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
use_cpu booleanOptional: {}
pii_mapping_path stringOptional: {}
backend RemoteClassifierBackendConfigBackend names a remote token classifier speaking token_spans.v1. Its
absence keeps local PII inference. The local selectors this replaces are
model_id and use_cpu above. Explicit token windows are only supported by
the local model.
Optional: {}
on_error stringOnError selects what a PII backend failure, or a provider-declared
truncation, does to the rule that consumed it: allow (default) treats the
content as not matching, block matches it as classification_error.
Enum: [allow block]
Optional: {}

PersistenceSpec​

PersistenceSpec defines persistence configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled indicates if persistence is enabledtrueOptional: {}
storageClassName stringStorageClassName is the storage class namestandardOptional: {}
accessMode PersistentVolumeAccessModeAccessMode is the access modeReadWriteOnceOptional: {}
size stringSize is the storage size10GiOptional: {}
existingClaim stringExistingClaim is an existing PVC to useOptional: {}
annotations object (keys:string, values:string)Annotations for the PVCOptional: {}

PortSpec​

PortSpec defines a service port configuration

Appears in:

FieldDescriptionDefaultValidation
port integerPort is the service portMaximum: 65535
Minimum: 1
Optional: {}
targetPort integerTargetPort is the container portMaximum: 65535
Minimum: 1
Optional: {}
protocol ProtocolProtocol is the port protocolTCPOptional: {}

ProbeSpec​

ProbeSpec defines probe configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled indicates if the probe is enabledtrueOptional: {}
initialDelaySeconds integerInitialDelaySeconds before probe startsOptional: {}
periodSeconds integerPeriodSeconds between probesOptional: {}
timeoutSeconds integerTimeoutSeconds for probeOptional: {}
failureThreshold integerFailureThreshold for probeOptional: {}

PromptGuardConfig​

PromptGuardConfig defines prompt guard configuration.

Appears in:

FieldDescriptionDefaultValidation
backend RemoteClassifierBackendConfigBackend selects a named external classifier and its typed result contract.Optional: {}
max_sequence_length integerMaxSequenceLength limits the total tokenized input, including special
tokens. Omission or zero retains the 512-token budget. The model loader
validates the requested budget against the loaded model's capacity.
Minimum: 0
Optional: {}
window PromptGuardWindowConfigWindow enables explicit scanning of all input tokens. Omission or null
keeps whole-input inference. Only the local model supports it.
Optional: {}
enabled booleantrueOptional: {}
model_id stringModelID binds the guard to one model; empty runs the decision model.Optional: {}
threshold stringJailbreak detection threshold (0.0-1.0). Stored as string to avoid float
precision issues; empty takes the threshold calibrated for the model.
Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
use_cpu booleantrueOptional: {}
jailbreak_mapping_path stringOptional: {}
positive_labels string arrayPositiveLabels lists the jailbreak_mapping labels that count as unsafe,
for a custom backend whose positive class isn't named "jailbreak"
(e.g. "INJECTION", "malicious"). Defaults to ["jailbreak"] when unset.
Optional: {}
on_error stringOnError selects what a prompt-guard classifier failure does to the rule
that failed to evaluate. "allow" (the default) tolerates the failure and
treats the content as not matching; "block" treats it as a positive
detection, because an inference failure means the content could not be
verified safe. Without this field on the CRD the setting is pruned by the
API server and an operator-managed deployment silently fails open.
Enum: [allow block]
Optional: {}

PromptGuardWindowConfig​

PromptGuardWindowConfig scans original content tokens with overlap. The native tokenizer also checks that special tokens leave enough content room.

Appears in:

FieldDescriptionDefaultValidation
size integerSize is the inference window budget, including special tokens.Minimum: 1
overlap integerOverlap counts content tokens shared by consecutive windows.Minimum: 0
Optional: {}

PrototypeScoringConfig​

PrototypeScoringConfig overrides prototype-bank construction and scoring for one embedding-backed signal rule. Router-owned defaults apply only after translation; the operator preserves an absent override and an empty object.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled controls prototype clustering. False retains every candidate.Optional: {}
cluster_similarity_threshold stringClusterSimilarityThreshold is the clustering similarity threshold.
Stored as a numeric string, like other fractional operator config fields.
Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}
max_prototypes integerMaxPrototypes caps the number of cluster representatives.Optional: {}
best_weight stringBestWeight weights the best prototype against the top-M mean.Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}
top_m integerTopM is the number of highest-scoring prototypes included in the mean.Optional: {}
margin_threshold stringMarginThreshold is the minimum winner-versus-runner-up score margin.Pattern: ^-?[0-9]+(\.[0-9]+)?$
Optional: {}

QdrantCacheConfig​

QdrantCacheConfig defines Qdrant cache backend configuration. Configure these settings when using Qdrant as the semantic cache backend.

Appears in:

FieldDescriptionDefaultValidation
host stringHost is the Qdrant server hostname or IP addressOptional: {}
port integerPort is the Qdrant gRPC port6334Maximum: 65535
Minimum: 1
Optional: {}
api_key stringAPIKey for Qdrant authenticationOptional: {}
use_tls booleanUseTLS enables TLS for the Qdrant connectionfalseOptional: {}
collection_name stringCollectionName is the Qdrant collection to use for semantic cachesemantic_cacheOptional: {}
connect_timeout integerConnectTimeout is the timeout in seconds for Qdrant connection10Optional: {}

RedisCacheConfig​

RedisCacheConfig defines Redis cache backend configuration. Configure these settings when using Redis as the semantic cache backend.

Appears in:

FieldDescriptionDefaultValidation
connection RedisCacheConnectionConnection settings for Redis serverOptional: {}
index RedisCacheIndexIndex settings for Redis vector searchOptional: {}
search RedisCacheSearchSearch settings for Redis queriesOptional: {}
development RedisCacheDevelopmentDevelopment settings for Redis cacheOptional: {}

RedisCacheConnection​

RedisCacheConnection defines Redis connection parameters.

Appears in:

FieldDescriptionDefaultValidation
host stringHost is the Redis server hostname or IP address
Example: "redis.default.svc.cluster.local"
Optional: {}
port integerPort is the Redis server port6379Maximum: 65535
Minimum: 1
Optional: {}
database integerDatabase is the Redis database number to use0Minimum: 0
Optional: {}
password stringPassword for Redis authentication (plaintext - consider using PasswordSecretRef instead)Optional: {}
password_secret_ref SecretKeySelectorPasswordSecretRef references a Secret containing the Redis password
Preferred over plaintext Password field for security
Optional: {}
timeout integerTimeout for Redis operations in seconds30Minimum: 0
Optional: {}
tls RedisCacheTLSTLS configuration for secure Redis connectionsOptional: {}

RedisCacheDevelopment​

RedisCacheDevelopment defines development-mode settings.

Appears in:

FieldDescriptionDefaultValidation
drop_index_on_startup booleanDropIndexOnStartup clears the index when router starts (for testing)falseOptional: {}
auto_create_index booleanAutoCreateIndex automatically creates the index if it doesn't existtrueOptional: {}

RedisCacheIndex​

RedisCacheIndex defines Redis vector index configuration.

Appears in:

FieldDescriptionDefaultValidation
name stringName of the Redis indexsemantic_cache_idxOptional: {}
prefix stringPrefix for Redis keysdoc:Optional: {}
vector_field RedisCacheVectorFieldVectorField configuration for embeddingsOptional: {}
index_type stringIndexType specifies the index algorithm
Options: "HNSW" (recommended), "FLAT"
HNSWEnum: [HNSW FLAT]
Optional: {}
params RedisCacheIndexParamsParams for HNSW indexOptional: {}

RedisCacheIndexParams​

RedisCacheIndexParams defines HNSW index parameters.

Appears in:

FieldDescriptionDefaultValidation
M integerM is the number of bi-directional links per node
Higher values = better recall, more memory
16Minimum: 2
Optional: {}
efConstruction integerEfConstruction is the size of dynamic candidate list during construction
Higher values = better quality, slower indexing
64Minimum: 1
Optional: {}

RedisCacheSearch​

RedisCacheSearch defines Redis search parameters.

Appears in:

FieldDescriptionDefaultValidation
topk integerTopK is the number of results to return from vector search1Minimum: 1
Optional: {}

RedisCacheTLS​

RedisCacheTLS defines TLS settings for Redis connections.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled controls whether to use TLS for Redis connectionfalseOptional: {}
cert_file stringCertFile is the path to client certificate fileOptional: {}
key_file stringKeyFile is the path to client key fileOptional: {}
ca_file stringCAFile is the path to CA certificate fileOptional: {}

RedisCacheVectorField​

RedisCacheVectorField defines vector field configuration.

Appears in:

FieldDescriptionDefaultValidation
name stringName of the vector fieldembeddingOptional: {}
dimension integerDimension of the embedding vectors
For BERT: 384, for Qwen3: 1024, for Gemma: 768
Minimum: 1
Optional: {}
metric_type stringMetricType for vector similarity
Options: "COSINE", "IP" (inner product), "L2" (Euclidean)
COSINEEnum: [COSINE IP L2]
Optional: {}

RemoteClassifierBackendConfig​

RemoteClassifierBackendConfig is the shared remote-classifier block. How the remote is called (protocol), what shape it answers with (contract), which catalog entry it is (model) and how long to wait (deadline) are independent axes rather than one enumeration. It mirrors the router's backend block field for field so the operator passes it through unchanged.

Appears in:

FieldDescriptionDefaultValidation
protocol stringProtocol is how the remote is called.Enum: [http_classify http_chat]
contract stringContract is the response shape the signal reads. Complexity reads two -
score.v1, one regression number interpreted through each rule's
boundaries, and label_distribution.v1, hard/easy/medium probabilities -
so the router requires it there rather than guessing per request. PII
reads token_spans.v1, entity spans with code-point offsets. Prompt guard
http_chat reads label_decision.v1, a verdict without invented probability.
Enum: [score.v1 label_distribution.v1 token_spans.v1 label_decision.v1]
Optional: {}
model stringModel is the name of an entry in the external model catalog.MinLength: 1
deadline_ms integerDeadlineMs bounds one remote call. Defaults to the router's value.Minimum: 1
Optional: {}

ResourceConfig​

ResourceConfig defines resource configuration for tracing

Appears in:

FieldDescriptionDefaultValidation
service_name stringvllm-semantic-routerOptional: {}
service_version stringv0.1.0Optional: {}
deployment_environment stringdevelopmentOptional: {}

RouteConfig​

RouteConfig defines OpenShift Route configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled specifies whether to create an OpenShift RoutefalseOptional: {}
hostname stringHostname for the Route (optional - OpenShift generates if empty)Optional: {}
tls RouteTLSConfigTLS configuration for the RouteOptional: {}

RouteTLSConfig​

RouteTLSConfig defines TLS configuration for OpenShift Routes

Appears in:

FieldDescriptionDefaultValidation
termination stringTermination type (edge, passthrough, reencrypt)edgeEnum: [edge passthrough reencrypt]
Optional: {}
insecureEdgeTerminationPolicy stringInsecureEdgeTerminationPolicy for HTTP trafficRedirectEnum: [Allow Redirect None]
Optional: {}

RuleCombinationConfig​

RuleCombinationConfig defines how to combine multiple rule conditions

Appears in:

FieldDescriptionDefaultValidation
operator stringOperator specifies how to combine conditions: "AND", "OR", or "NOT". NOT is strictly unary: it takes
exactly one child condition and negates its result. Compose NOR/NAND by nesting NOT around OR/AND.
Enum: [AND OR NOT]
on_unknown stringOnUnknown resolves a terminal unknown result after the rule tree is evaluated.Enum: [no_match match fail_request]
Optional: {}
conditions RuleConditionConfig arrayConditions is the list of rule references to evaluate

RuleComposition​

RuleComposition defines how to compose/filter rules based on other signals

Appears in:

FieldDescriptionDefaultValidation
operator stringOperator for combining conditions (AND, OR, NOT). NOT is strictly unary and negates its single child.Enum: [AND OR NOT]
conditions CompositionCondition arrayList of conditions that must be met

RuleConditionConfig​

RuleConditionConfig references a specific rule by type and name

Appears in:

FieldDescriptionDefaultValidation
type stringType specifies the signal or projection type referenced by this condition.Enum: [keyword embedding domain fact_check user_feedback reask preference language context structure complexity modality authz jailbreak pii kb conversation event projection]
name stringName is the name of the rule to reference

SamplingConfig​

SamplingConfig defines sampling configuration

Appears in:

FieldDescriptionDefaultValidation
type stringalways_onOptional: {}
rate stringSampling rate (0.0-1.0). Stored as string to avoid float precision issues.1.0Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}

SemanticCacheConfig​

SemanticCacheConfig defines semantic cache configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled controls whether semantic caching is activetrueOptional: {}
backend_type stringBackendType specifies the cache backend to use
Options: "memory" (default), "redis", "valkey", "milvus", "qdrant", "hybrid"
memoryEnum: [memory redis valkey milvus qdrant hybrid]
Optional: {}
similarity_threshold stringSimilarity threshold for cache hits (0.0-1.0). Stored as string to avoid float precision issues.0.8Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
max_entries integerMaxEntries is the maximum number of cache entries (for memory/hybrid backends)1000Optional: {}
ttl_seconds integerTTLSeconds is the time-to-live for cache entries in seconds3600Optional: {}
eviction_policy stringEvictionPolicy for in-memory cache ("fifo", "lru", "lfu")fifoEnum: [fifo lru lfu]
Optional: {}
redis RedisCacheConfigRedis configuration (required when backend_type is "redis")Optional: {}
valkey ValkeyCacheConfigValkey configuration (required when backend_type is "valkey")Optional: {}
milvus MilvusCacheConfigMilvus configuration (required when backend_type is "milvus")Optional: {}
qdrant QdrantCacheConfigQdrant configuration (required when backend_type is "qdrant")Optional: {}
embedding_model stringEmbeddingModel specifies which embedding model to use for semantic similarity
Options: "mmbert" (default), "bert", "qwen3", "gemma"
mmbertEnum: [bert qwen3 gemma mmbert]
Optional: {}
hnsw HNSWCacheConfigHNSW configuration for hybrid/in-memory backendsOptional: {}

SemanticRouter​

SemanticRouter is the Schema for the semanticrouters API

Appears in:

FieldDescriptionDefaultValidation
apiVersion stringvllm.ai/v1alpha1
kind stringSemanticRouter
metadata ObjectMetaRefer to Kubernetes API documentation for fields of metadata.
spec SemanticRouterSpec
status SemanticRouterStatus

SemanticRouterList​

SemanticRouterList contains a list of SemanticRouter

FieldDescriptionDefaultValidation
apiVersion stringvllm.ai/v1alpha1
kind stringSemanticRouterList
metadata ListMetaRefer to Kubernetes API documentation for fields of metadata.
items SemanticRouter array

SemanticRouterSpec​

SemanticRouterSpec defines the desired state of SemanticRouter

Appears in:

FieldDescriptionDefaultValidation
image ImageSpecImage configurationOptional: {}
replicas integerNumber of replicas1Minimum: 0
Optional: {}
imagePullSecrets LocalObjectReference arrayImagePullSecrets for private registriesOptional: {}
serviceAccount ServiceAccountSpecServiceAccount configurationOptional: {}
service ServiceSpecService configurationOptional: {}
resources ResourceRequirementsResource requirementsOptional: {}
persistence PersistenceSpecPersistence configurationOptional: {}
config ConfigSpecConfiguration overrides merged into the canonical v0.3 config.yaml.
Router-wide runtime overrides land under config.global.router/services/stores/
integrations/model_catalog, with model-backed modules nested under
config.global.model_catalog.modules. Provider defaults land under
config.providers.defaults.
Optional: {}
toolsDb ToolEntry arrayTools database configurationOptional: {}
vllmEndpoints VLLMEndpointSpec arrayVLLMEndpoints is a Kubernetes-native backend discovery adapter.
It generates canonical config.providers.models[].backend_refs
and config.routing.modelCards entries.
Optional: {}
autoscaling AutoscalingSpecAutoscaling configurationOptional: {}
startupProbe ProbeSpecProbes configurationOptional: {}
livenessProbe ProbeSpecOptional: {}
readinessProbe ProbeSpecOptional: {}
securityContext SecurityContextSecurity contextOptional: {}
podSecurityContext PodSecurityContextPod security contextOptional: {}
podAnnotations object (keys:string, values:string)Pod annotationsOptional: {}
nodeSelector object (keys:string, values:string)Node selectorOptional: {}
tolerations Toleration arrayTolerationsOptional: {}
affinity AffinityAffinityOptional: {}
env EnvVar arrayEnvironment variablesOptional: {}
args string arrayRouter arguments, after the gateway mode flags the Operator passes
(-gateway=standalone -listener-address=0.0.0.0, or -gateway=extproc).
The gateway mode follows spec.gateway, so args may not set those flags.
MaxItems: 64
items:MaxLength: 4096
Optional: {}
gateway GatewaySpecGateway selects what serves client traffic. Omitted, the Router runs
standalone: it serves the OpenAI-compatible API on port 8801 itself,
with no Envoy, and the Service exposes that port. With existingRef, the
Router serves ext_proc gRPC on port 50051 for that Gateway, whose
ext_proc policy and routes you manage.
Optional: {}
openshift OpenShiftSpecOpenShift-specific featuresOptional: {}
ingress IngressSpecIngress configurationOptional: {}

SemanticRouterStatus​

SemanticRouterStatus defines the observed state of SemanticRouter

Appears in:

FieldDescriptionDefaultValidation
conditions Condition arrayConditions represent the latest available observations of the SemanticRouter's stateOptional: {}
observedGeneration integerObservedGeneration reflects the generation of the most recently observed SemanticRouterOptional: {}
replicas integerReplicas is the current number of replicasOptional: {}
readyReplicas integerReadyReplicas is the number of ready replicasOptional: {}
phase stringPhase represents the current phase of the SemanticRouterOptional: {}
gatewayMode stringGatewayMode is what serves client traffic: standalone (the Router's own
listener on port 8801) or gateway-integration (the Gateway in
spec.gateway, with the Router serving ext_proc)
Optional: {}
openshiftFeatures OpenShiftFeaturesStatusOpenShiftFeatures tracks OpenShift-specific feature statusOptional: {}

ServiceAccountSpec​

ServiceAccountSpec defines service account configuration

Appears in:

FieldDescriptionDefaultValidation
create booleanCreate specifies whether to create a service accounttrueOptional: {}
name stringName of the service account to useOptional: {}
annotations object (keys:string, values:string)Annotations for the service accountOptional: {}

ServiceBackend​

ServiceBackend defines a direct Kubernetes service backend

Appears in:

FieldDescriptionDefaultValidation
name stringService nameMinLength: 1
namespace stringService namespace (defaults to same namespace)Optional: {}
port integerService portMinimum: 1

ServiceSpec​

ServiceSpec defines the service configuration

Appears in:

FieldDescriptionDefaultValidation
type ServiceTypeType is the service typeClusterIPEnum: [ClusterIP NodePort LoadBalancer]
Optional: {}
grpc PortSpecGRPC port configurationOptional: {}
api PortSpecAPI port configurationOptional: {}
metrics MetricsPortSpecMetrics port configurationOptional: {}

StreamedBodyConfig​

StreamedBodyConfig defines streamed request body handling.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled accumulates request body chunks before routing at end-of-stream.Optional: {}
max_bytes integerMaxBytes caps the accumulated body size. A larger body is rejected and the
ExtProc stream ends; the downstream response follows the gateway's ExtProc
failure policy. Zero disables the limit.
Minimum: 0
Optional: {}
timeout_sec integerTimeoutSec caps how long body accumulation may take. A slower body is
rejected and the ExtProc stream ends; the downstream response follows the
gateway's ExtProc failure policy. Zero disables the limit.
Minimum: 0
Optional: {}

Tool​

Tool defines a tool function

Appears in:

FieldDescriptionDefaultValidation
type stringEnum: [function]
Optional: {}
function ToolFunctionOptional: {}

ToolEntry​

ToolEntry defines a tool entry in the tools database

Appears in:

FieldDescriptionDefaultValidation
tool ToolOptional: {}
description stringOptional: {}
category stringOptional: {}
tags string arrayOptional: {}

ToolFunction​

ToolFunction defines a tool function details

Appears in:

FieldDescriptionDefaultValidation
name stringOptional: {}
description stringOptional: {}
parameters ToolParametersOptional: {}

ToolParameters​

ToolParameters defines tool function parameters

Appears in:

FieldDescriptionDefaultValidation
type stringOptional: {}
properties JSONType: object
Optional: {}
required string arrayOptional: {}

ToolsConfig​

ToolsConfig defines tools configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleantrueOptional: {}
top_k integer3Optional: {}
similarity_threshold stringSimilarity threshold for tool selection (0.0-1.0). Stored as string to avoid float precision issues.0.2Pattern: ^0(\.[0-9]+)?$|^1(\.0+)?$
Optional: {}
tools_db_path stringconfig/tools_db.jsonOptional: {}
fallback_to_empty booleantrueOptional: {}

TracingConfig​

TracingConfig defines tracing configuration

Appears in:

FieldDescriptionDefaultValidation
enabled booleanfalseOptional: {}
provider stringopentelemetryOptional: {}
exporter ExporterConfigOptional: {}
sampling SamplingConfigOptional: {}
resource ResourceConfigOptional: {}

VLLMBackend​

VLLMBackend specifies how to reach the vLLM service

Appears in:

FieldDescriptionDefaultValidation
type stringType of backend: kserve, llamastack, or serviceEnum: [kserve llamastack service]
inferenceServiceName stringFor type=kserve: InferenceService name for auto-discoveryOptional: {}
discoveryLabels object (keys:string, values:string)For type=llamastack: Labels to match servicesOptional: {}
service ServiceBackendFor type=service: Direct service configurationOptional: {}

VLLMEndpointSpec​

VLLMEndpointSpec defines a vLLM model backend endpoint

Appears in:

FieldDescriptionDefaultValidation
name stringName of the backend ref generated under config.providers.models[].backend_refsMinLength: 1
model stringModel name as reported by vLLM (e.g., "Model-A", "llama3-8b")MinLength: 1
catalog stringCatalog optionally selects a repository built-in Model Card. Model remains
the request-facing alias.
Optional: {}
reasoning ModelReasoningSpecReasoning optionally selects a built-in family or defines inline wire
behavior for this self-hosted model. Catalog-backed models normally omit it.
Optional: {}
loras LoRAAdapterSpec arrayLoRAs declares the LoRA adapters exposed for this logical model in routing.modelCards.MaxItems: 50
Optional: {}
backend VLLMBackendBackend configuration
weight integerWeight for load balancing (default: 1)1Optional: {}

ValkeyCacheConfig​

ValkeyCacheConfig defines Valkey cache backend configuration. Configure these settings when using Valkey as the semantic cache backend.

Appears in:

FieldDescriptionDefaultValidation
connection ValkeyCacheConnectionConnection settings for Valkey serverOptional: {}
index ValkeyCacheIndexIndex settings for Valkey vector searchOptional: {}
search ValkeyCacheSearchSearch settings for Valkey queriesOptional: {}
development ValkeyCacheDevelopmentDevelopment settings for Valkey cacheOptional: {}

ValkeyCacheConnection​

ValkeyCacheConnection defines Valkey connection parameters.

Appears in:

FieldDescriptionDefaultValidation
host stringHost is the Valkey server hostname or IP address
Example: "valkey.default.svc.cluster.local"
Optional: {}
port integerPort is the Valkey server port6379Maximum: 65535
Minimum: 1
Optional: {}
database integerDatabase is the Valkey database number to use0Minimum: 0
Optional: {}
password stringPassword for Valkey authentication (plaintext - consider using PasswordSecretRef instead)Optional: {}
password_secret_ref SecretKeySelectorPasswordSecretRef references a Secret containing the Valkey password
Preferred over plaintext Password field for security
Optional: {}
timeout integerTimeout for Valkey operations in seconds30Minimum: 0
Optional: {}
tls ValkeyCacheTLSTLS configuration for secure Valkey connectionsOptional: {}

ValkeyCacheDevelopment​

ValkeyCacheDevelopment defines development-mode settings.

Appears in:

FieldDescriptionDefaultValidation
drop_index_on_startup booleanDropIndexOnStartup clears the index when router starts (for testing)falseOptional: {}
auto_create_index booleanAutoCreateIndex automatically creates the index if it doesn't existtrueOptional: {}

ValkeyCacheIndex​

ValkeyCacheIndex defines Valkey vector index configuration.

Appears in:

FieldDescriptionDefaultValidation
name stringName of the Valkey indexsemantic_cache_idxOptional: {}
prefix stringPrefix for Valkey keysdoc:Optional: {}
vector_field ValkeyCacheVectorFieldVectorField configuration for embeddingsOptional: {}
index_type stringIndexType specifies the index algorithm
Options: "HNSW" (recommended), "FLAT"
HNSWEnum: [HNSW FLAT]
Optional: {}
params ValkeyCacheIndexParamsParams for HNSW indexOptional: {}

ValkeyCacheIndexParams​

ValkeyCacheIndexParams defines HNSW index parameters.

Appears in:

FieldDescriptionDefaultValidation
M integerM is the number of bi-directional links per node
Higher values = better recall, more memory
16Minimum: 2
Optional: {}
efConstruction integerEfConstruction is the size of dynamic candidate list during construction
Higher values = better quality, slower indexing
64Minimum: 1
Optional: {}

ValkeyCacheSearch​

ValkeyCacheSearch defines Valkey search parameters.

Appears in:

FieldDescriptionDefaultValidation
topk integerTopK is the number of results to return from vector search1Minimum: 1
Optional: {}

ValkeyCacheTLS​

ValkeyCacheTLS defines TLS settings for Valkey connections.

Appears in:

FieldDescriptionDefaultValidation
enabled booleanEnabled controls whether to use TLS for Valkey connectionfalseOptional: {}
cert_file stringCertFile is the path to client certificate fileOptional: {}
key_file stringKeyFile is the path to client key fileOptional: {}
ca_file stringCAFile is the path to CA certificate fileOptional: {}

ValkeyCacheVectorField​

ValkeyCacheVectorField defines vector field configuration.

Appears in:

FieldDescriptionDefaultValidation
name stringName of the vector fieldembeddingOptional: {}
dimension integerDimension of the embedding vectors
For BERT: 384, for Qwen3: 1024, for Gemma: 768
Minimum: 1
Optional: {}
metric_type stringMetricType for vector similarity
Options: "COSINE", "IP" (inner product), "L2" (Euclidean)
COSINEEnum: [COSINE IP L2]
Optional: {}