Skip to main content
Version: Latest (unreleased)

Stores and Tools

Overview​

This page covers the shared storage and tool blocks inside global:.

These settings back route-local plugins and router-wide tool behavior.

Key Advantages​

  • Centralizes shared backing stores instead of repeating them per route.
  • Keeps response cache, memory, retrieval, and tool catalogs consistent.
  • Lets route-local plugins stay small and focused.
  • Makes shared infrastructure dependencies explicit.

What Problem Does It Solve?​

Route-local plugins often depend on shared storage or tool state. If those dependencies are configured ad hoc inside each route, the system becomes inconsistent and harder to operate.

These global: blocks solve that by defining shared backing services once.

When to Use​

Use these blocks when:

  • multiple routes depend on the same response cache or memory backend
  • retrieval features need one shared vector store
  • the router should expose one shared tool catalog
  • backing-store configuration belongs to the whole router rather than one route

Configuration​

Response Cache​

global:
stores:
response_cache:
enabled: true
backend_type: memory
similarity_threshold: 0.8

Negation guard​

Bi-encoder similarity cannot tell "turn on dark mode" from "turn off dark mode": opposite-meaning queries often score above similarity_threshold while genuine paraphrases score below it, so raising the threshold does not reliably prevent false hits. Every cache backend applies a lexical guard before it serves a semantic candidate: it catches negation cues and known antonym swaps in near-identical English token sets. It needs no model, has no setting, and runs for the in-memory, Redis, Valkey, Milvus, Qdrant and hybrid caches. It does not cover cue-less, word-order-only or non-English changes.

Rejections are logged as cache_negation_reject, count as misses, and still surface the rejected score on x-vsr-cache-similarity. Remote and hybrid backends reject candidates without an original query and continue checking the bounded fetched candidate set. The same check also applies to hybrid's Milvus fallback.

Earlier releases configured it with polarity_guard, including an NLI tier (nli, lexical+nli) on the in-memory backend. The NLI model is retired and the router refuses polarity_guard; vllm-sr config migrate removes it.

Memory​

The memory store supports three backends: milvus (default), valkey, and qdrant.

Milvus backend (default):

global:
stores:
memory:
enabled: true
milvus:
address: milvus:19530
collection: agentic_memory
dimension: 256

Valkey backend (requires Valkey with Search module):

global:
stores:
memory:
enabled: true
backend: valkey
valkey:
host: valkey
port: 6379
dimension: 256
collection_prefix: "mem:"
index_name: mem_idx
metric_type: COSINE

Qdrant backend:

global:
stores:
memory:
enabled: true
backend: qdrant
qdrant:
host: qdrant
port: 6334
collection: agentic_memory
dimension: 256
embedding_model: mmbert
default_retrieval_limit: 5
default_similarity_threshold: 0.30

All three examples embed with mmbert (Vela Embedding), the default, so dimension is one of its sizes: 64, 128, 256, 512 or 768. With any other size the router logs Failed to create memory store: … Memory will be disabled and runs without memory. The Qdrant example's 0.30 threshold, on plain cosine scores, is a starting point, not a calibrated value. Check unrelated queries and corrected facts before using it with your data.

For full deployment instructions, see:

  • Valkey Agentic Memory — Docker, Kubernetes, config reference, tuning, and troubleshooting
  • Qdrant — Docker, Kubernetes, config reference, tuning, and troubleshooting
  • config/runtime/memory/ for backend-specific configuration references

When an external model with model_role: memory_rewrite is configured, its max_response_bytes limits each query-rewrite response. An omitted or non-positive value uses the 1 MiB default.

Hybrid search and reflection​

hybrid_mode and reflection.algorithm can be set globally under global.stores.memory and overridden in a decision's memory plugin. The router refuses to start with any other value, including one set in a recipe's decisions:

FieldAccepted valuesDefault
hybrid_modeweighted, rrf (exact, lowercase)weighted
reflection.algorithmheuristic, noopheuristic

hybrid_mode takes effect only with hybrid_search: true. Older configs that use hybrid_mode: rerank or algorithm: recency_semantic fail at startup. Replace them with weighted and heuristic, which is how those values already ran.

Write path bounds​

Automatic persistence first respects Memory enablement, retention policy, and an explicit auto_store: false on the selected decision's memory plugin. A Responses request may opt out, but cannot override these server restrictions. When policy permits persistence, the Responses request's auto_store takes precedence over the decision's value, followed by global.stores.memory. Only an omitted value falls back to the next level.

Response handling does not wait for Memory persistence to complete. Identity checks and capacity reservation precede the bounded history snapshot, which is taken while the response path still owns the conversation state; protocol encoding and writes run in the background. Other response-path Replay operations remain synchronous.

Configure global.stores.memory.persistence:

FieldMeaningDefault
timeout_secondsSeconds from reservation to timeout, including preparation, queue wait, and writing; 0–9,223,372,03630
concurrencyWorker slots, including preparation and writes; 0–648
queueReserved attempts waiting for a worker; 0–102464
shutdown_grace_secondsSeconds to drain writes on reload or shutdown before cancellation; 0–9,223,372,0365

Omit a field or set it to 0 to take the default. Negative values and values above the listed limits are rejected during configuration validation, before workers or queue storage are allocated at startup or reload. This also applies to the initial config_source: kubernetes document, before the controller loads routing CRDs; global resource bounds are not deferred.

Each persistence attempt has a shared 1 MiB payload budget for request history, retained Responses history, and the current assistant response. Assistant text is counted before think-tag stripping. History is also limited to 256 messages/items and 32 nested content levels; history and the current response share a 4096-node structural limit. Bounded length checks run before reserving persistence capacity; text assembly and history copying run only after admission. Exceeding a limit skips persistence with skipped / history_too_large and fail_open=true, without occupying persistence capacity or truncating the model response or history. Missing user identity skips preparation with skipped / memory_info_unavailable and fail_open=true. A response the jailbreak or hallucination policy blocks reports policy_blocked instead of persisting. Background contexts retain only span context and tracestate.

For requests with a Router Replay record, accepted attempts reserve capacity for scheduled and one terminal receipt, protecting both from queue saturation. Storage errors, shutdown drain expiry, or process crashes can still lose receipts. Exhausted persistence or receipt capacity rejects new writes with queue_full or receipt_queue_full; retired pools use shutting_down. These remain fail-open and are logged by request ID. Monitor llm_plugin_execution_total{plugin_type="memory_persistence", status="rejected"}.

Timeout and cancellation report one terminal outcome even while queued; cancelled jobs do not start. Native embedding calls cannot be interrupted, so active work retains its worker slot and resources until exit. Cancellation does not undo writes already accepted by a backend.

Vector Store​

global:
stores:
vector_store:
enabled: true
backend_type: milvus
metadata_store: postgres

Supported backends: memory, milvus, llama_stack, valkey, qdrant.

metadata_store controls the registry for vector-store and uploaded-file metadata. Use postgres for restart-safe local or production-like stacks; the CLI local runtime will provision Postgres and fill metadata_postgres connection defaults when metadata_store: postgres is set. Use memory only for ephemeral local experiments because store and file metadata is lost on router restart.

With embeddings from the model runtime, including Vela Embedding, each new vector store records the identity of the representation that created its vectors. Existing stores remain visible and their uploaded files are retained. Searching or attaching files to an incompatible or untagged store returns 409 EMBEDDING_REINDEX_REQUIRED. Create a new vector store and reattach the original uploaded file IDs to generate compatible vectors. Client metadata cannot replace the router-owned _router_embedding_identity field.

The same check applies to request-time RAG and cached retrieval results. The llama_stack backend embeds search queries remotely, so it cannot currently be combined with identity-bound runtime document embeddings. Use memory, milvus, valkey, or qdrant for that configuration. Remote provider identity verification is a separate capability.

Tools​

global:
integrations:
tools:
enabled: true
top_k: 3
tools_db_path: config/runtime/tools/tools_db.json

Data and Security​

  • Cache, memory, and vector stores can contain prompts, responses, embeddings, retrieved documents, or extracted memories. Configure authentication, encryption, retention, and tenant/user scope for the selected backend.
  • Embedding dimensions must match existing collections. Rebuild or migrate an index when the embedding model or dimension changes.
  • Tool retrieval controls what is shown to a model; it does not authorize tool execution. Enforce permissions at the tool service.
  • See complete backend examples and the full configuration contract in config/config.yaml.