Skip to main content
Version: Latest (unreleased)

Router management API

The router management API provides configuration, routing previews, plugin inspection, model diagnostics, storage, and observability. It listens on port 8080 by default and the local stack binds it to 127.0.0.1.

For model traffic, use the configured inference listener described in Router API.

Start with the live schema​

The running router generates its endpoint discovery and OpenAPI document from the routes it has registered. Use these pages for registered methods, path and query parameters, request-body fields, access policy, and response media:

PathPurpose
GET /api/v1Endpoint discovery with permission and sensitivity metadata
GET /openapi.jsonComplete OpenAPI 3.0 document
GET /openapi.json?path=...&method=...One valid path or operation document
GET /docsInteractive Swagger UI

This page groups the API by user task. The live OpenAPI document is the field-level source of truth for the version you are running.

curl -sS http://localhost:8080/health
curl -sS http://localhost:8080/openapi.json
curl -sS 'http://localhost:8080/openapi.json?path=/api/v1/config&method=PATCH'

Agents should call these Router endpoints directly; the Dashboard is not part of the discovery path. The Website also renders the generated contract in the searchable OpenAPI reference.

Resource boundaries​

CapabilityResponsibility
configCanonical configuration, recipes, validation, planning, and activation
routingEvaluate a request through the configured routing pipeline
pluginsDiscover plugin types, inspect recipe/decision bindings, and preview or probe behavior
inventoryInspect configured and prepared runtime resources
diagnosticsInvoke a specific prepared model or a routing classifier
storageManage stored data, response-cache partitions, and context-recovery scopes
observabilityRead metrics, replays, and management audit; submit outcome evidence
systemHealth, startup, readiness, and API discovery

Plugin configuration belongs to its recipe decision in the canonical document. Use /api/v1/config/recipes/{name} with If-Match to change it. There is no second plugin configuration database or independent plugin enable/disable state. /config/router is not registered; the canonical configuration API is /api/v1/config.

Access and authentication​

The local CLI keeps the management port on loopback. For a remote router, prefer a private network or an SSH tunnel instead of publishing the port:

ssh -N -L 8080:127.0.0.1:8080 router-host

Management authentication is disabled unless configured. To require bearer tokens, set global.services.management_api.auth.mode: bearer and define roles and token sources in the management API configuration. Then send:

Authorization: Bearer <token>

GET /health remains public. Other routes enforce their assigned permission when bearer authentication is enabled. Configuration and replay responses can also redact sensitive fields unless the principal has the corresponding detail permission.

Health and discovery​

MethodPathUse
GET/healthProcess liveness
GET/readyWhether startup has completed
GET/startup-statusStartup and model-download progress
GET/api/v1/statusVersioned replica-local startup and configuration observations
GET/api/v1Registered endpoint discovery
GET/openapi.jsonGenerated OpenAPI schema, optionally narrowed by exact path and method
GET/docsSwagger UI

Use /health for liveness and /ready for readiness. During model download or runtime preparation, a process can be healthy while /ready still returns 503.

Startup completes once every model deployment the Router manages for its configuration is ready, decision models included. While the Router waits for them, /ready and /startup-status report phase: loading_model_deployments, pending_models names the deployments that are not ready, and ready_models and total_models count them. /startup-status also lists them in model_deployments, each with its name, artifact, process, state, ready and, after a failure, reason, and keeps the list once startup is complete. A model that fails to load ends startup with phase: error. A configuration reload does not turn /ready back to 503: the previous configuration serves until the new one's models are ready. The standalone listeners and the ext_proc gRPC port open only once those models are ready, so the listener's /ready and the gRPC health service never report ready before the management /ready does.

For a Router with a runtime registry, /ready and /startup-status report that replica's observed startup state. Another replica's shared file or Redis record cannot change these responses. Until the local replica reports startup progress, both endpoints return 503.

Replica-local status​

GET /api/v1/status requires ready.read. It returns schema_version: "v1", an opaque instance_id, an observation timestamp, and conditions with True, False, or Unknown status. The instance identity changes when the Router registry is recreated.

The response reads in-process startup observations and the active configuration registry. It does not treat a shared Redis value as this replica's state. Configuration hashes identify the active snapshot and the most recent candidate observed by the activation coordinator. They do not claim to describe an unobserved external desired configuration.

ConditionMeaning
LiveThe process answered the request
StartupCompleteThe local startup writer reported completion or failure
ActiveConfigAvailableThe registry has a classification runtime that can still be acquired
ConfigConvergedThe active and latest observed configuration hashes match
ReadyKnown startup/runtime failures are false; required dependency and quorum readiness is otherwise unknown

The stable reason codes are ProcessRunning, StartupUnobserved, StartupIncomplete, StartupFailed, StartupComplete, NoActiveConfig, ActiveConfigAvailable, ConfigIdentityUnobserved, ConfigHashesMatch, ConfigActivationPending, ConfigActivationFailed, ConfigActivationSuperseded, and RequiredDependenciesUnobserved. Activation stages are diagnostic metadata, not additional readiness conditions.

A failed reload reports ConfigActivationFailed without discarding the last-known-good active hash. Successful publication changes the active hash and reports ConfigHashesMatch. Raw startup messages, credentials, and activation failure details are excluded.

This first contract does not evaluate every required dependency, recipe capability, or quorum. RequiredDependenciesUnobserved must not be treated as ready. HTTP 200 means the status document was retrieved, not that traffic can be served. The existing /health, /ready, and /startup-status contracts are unchanged; this endpoint does not replace their deployment probes. Fleet aggregation, per-replica persistence/TTL, and probe wiring remain separate work.

Inspect signals without an inference call​

The classification endpoints are useful when tuning signals or diagnosing why a decision did not match. They do not call a generation backend.

curl -sS http://localhost:8080/api/v1/diagnostics/classify/intent \
-H 'Content-Type: application/json' \
-d '{"text":"Write a Python function that merges two sorted lists."}'
MethodPathUse
POST/api/v1/diagnostics/classify/intentEvaluate intent/domain routing
POST/api/v1/diagnostics/classify/piiDetect configured PII types
POST/api/v1/diagnostics/classify/securityEvaluate jailbreak and security classification
POST/api/v1/diagnostics/classify/fact-checkDecide whether text needs fact checking
POST/api/v1/diagnostics/classify/user-feedbackClassify user feedback
POST/api/v1/diagnostics/classify/combinedRun intent, PII, and security classification
POST/api/v1/diagnostics/classify/batchRun a selected classifier over a batch
POST/api/v1/routing/previewEvaluate all configured signals
POST/api/v1/diagnostics/embeddingsGenerate configured text or image embeddings
POST/api/v1/diagnostics/similarityCompare a text pair
POST/api/v1/diagnostics/similarity/batchRun batch similarity matching

Names, scores, and matched rules depend on the active recipe. Use the live schema for each endpoint's supported input forms.

When the matched decision uses fast_response, Preview reports selection_status: not_required and selection_method: fast_response, with no selected_model. This immediate response needs no model assignment or candidate capability/context admission. The client-facing response model identifier does not imply that a generation backend was selected or called.

Guard and PII report input_limit in signal_errors when input exceeds their configured inference budget. Check the effective model and deployment limits before retrying. Other inference failures retain their bounded signal error codes; the configured unknown-signal policy determines the route outcome.

Standalone classification, embedding, and similarity diagnostics return 400 INVALID_INPUT when inference reports that a model's input budget was exceeded. The error retains the model's limit details; other inference failures remain server errors.

Diagnose a prepared model binding​

GET /api/v1/diagnostics/models?recipe=<name> lists the bindings actually prepared for that recipe. Select a returned name and its contract; the Router does not infer a model from a family name or silently load a fallback. Both recipe and binding are required for the typed invocation endpoints.

Operation under /api/v1/diagnostics/modelsPrepared taskVela use
POST /labelsLabel distribution, including configured window scanningDomain, Guard, FactCheck, Feedback, Modality
POST /label-scoresIndependent label scores and configured operating pointSafety, Hazard
POST /tokensToken spans with original UTF-8 byte offsetsPII
POST /embeddingsThe binding's configured embedding representationEmbedding
POST /rerankQuery/document relevance scores in input orderReranker

For example, take a label-distribution binding name from discovery:

curl -sS 'http://localhost:8080/api/v1/diagnostics/models?recipe=default'
curl -sS http://localhost:8080/api/v1/diagnostics/models/labels \
-H 'Content-Type: application/json' \
-d '{"recipe":"default","binding":"<prepared-binding-name>","text":"Explain this Python error."}'

Every invocation returns the actual binding identity, artifact revision, provider, device, precision, and task limits alongside the result. Calls retain the owning runtime generation until inference finishes, including during reload. An unprepared or foreign-recipe binding fails explicitly. Windowed Guard and PII keep their prepared scan geometry; Hazard keeps its operating-point thresholds and policy digest. The shared Vela Encoder is an artifact used by task bindings, so it does not need a separate inference endpoint.

The convenience classification, embedding, and similarity endpoints also accept an optional recipe. Omitting it retains their default-recipe behavior; a supplied recipe is resolved explicitly. The typed model endpoints are the preferred way to verify exactly which loaded model produced a result. Rerank batches obey global.services.api.batch_classification.max_batch_size (100 pairs when unset). Typed calls share their prepared resource's admission gate and have a two-minute request deadline; a model call retains its lease until the model runtime answers, even after cancellation. Actual coverage follows active configuration: discovering a task does not load it.

Inspect models and metrics​

MethodPathUse
GET/api/v1/inventory/modelsLoaded model inventory
GET/api/v1/inventory/classifierClassifier configuration and status
GET/api/v1/inventory/embedding-modelsLoaded embedding models
GET/v1/modelsOpenAI-compatible model list
GET/api/v1/observability/classification-metricsClassification counters and timing

Secrets in classifier information are redacted unless the caller has secret_view.

The model inventory reports successfully prepared task bindings in the active runtime generation, including each binding's recipe and effective provider, device, precision, and input limit in metadata. Shared artifacts may appear under several recipe bindings. Configured but unused models are not marked ready. During startup, the inventory can instead report pending artifact downloads. registry.repo_id preserves the Hugging Face source identity, including for a derived local graph whose declared source revision uniquely matches a registered checkpoint. Such a match is reported as metadata.registry_match: source_revision; it identifies the declared source, not byte-for-byte equivalence of the derived artifact. metadata.resource_id remains the physical runtime identity.

Input limits have separate meanings: forward_max_tokens is the model's single forward-pass capacity, input_max_tokens is the effective input budget, and windowed bindings also expose document_max_tokens, window_size, and window_overlap. Unknown values are omitted. The legacy max_sequence_length field retains its effective-input-limit meaning and must not be interpreted as the physical context window.

system.gpu_available means an active prepared binding uses local GPU execution; it does not indicate whether the host has unused GPU hardware.

Read and change router configuration​

Read the current canonical document and its ETag before making a change:

curl -i http://localhost:8080/api/v1/config \
-H "Authorization: Bearer ${VSR_MGMT_TOKEN}"
MethodPathUse
GET/api/v1/configRead the active canonical configuration
POST/api/v1/config/validateValidate and normalize without writing
POST/api/v1/config/planPlan the exact candidate and return the current/candidate ETags without writing
PATCH/api/v1/configMerge, validate, persist, and hot-reload an update
PUT/api/v1/configReplace, validate, persist, and hot-reload the document
GET/api/v1/config/versionsList the configuration history
POST/api/v1/config/rollbackActivate a recorded version again, as a new version
GET/api/v1/config/hashCompare persisted, generated, and active hashes

Recipe operations use the same canonical document:

MethodPathUse
GET/api/v1/config/recipesList default and named recipes and their entrypoints
POST/api/v1/config/recipes/validateValidate a recipe mutation without applying it
GET/api/v1/config/recipes/{name}Read one recipe
PUT/api/v1/config/recipes/{name}Create or replace one recipe
DELETE/api/v1/config/recipes/{name}Delete an unreferenced named recipe

Every config mutation, including rollback and Recipe PUT/DELETE, requires the exact current ETag in If-Match. The Router does not accept unguarded writes. A mutation response and GET /api/v1/config/hash use the same explicit runtime identity fields: source_config_hash, generated_runtime_hash, active_runtime_hash, and activation_status. Config mutations validate, keep the replaced document in the configuration history, and trigger reload; an active config still does not prove that upstream model backends are healthy. Check /ready and send a representative request after a change.

activation_status is active, pending, failed, or unknown. When a candidate has been attempted, activation includes its document hash, attempt, stage, timestamps, and a redacted failure detail. A failed candidate leaves the previous generation active. Config mutation responses return 503 with status: activation_failed when that failure is observed during the activation wait: the document was persisted, so read its current ETag and inspect or roll back that candidate instead of blindly retrying the write. The status is process-local and describes the latest attempted candidate. A newer file supersedes an older attempt.

Each activation takes the next configuration version. A mutation that activates returns it as config_version. A rejected one carries activation.reasons: each has a stage (parse, compile, validate, warm or activate), a code, an optional path into the document, and a redacted message. GET /api/v1/config names the version that serves in the x-vsr-config-version header and its document hash in x-vsr-config-hash; both can trail the persisted document while it activates or after it is rejected. GET /api/v1/config/hash adds active_version and last_rejection.

An activation rebuilds only what its changes reach. When the signals and the models they use are unchanged, the new version keeps the loaded classifiers and embedding models. Under --gateway standalone, it also keeps the upstream connection pools when the providers' backends and the listeners are unchanged. A request in standalone mode is served by the version it started on, from routing through its fallback chain to the last byte. Such a Router binds its listeners at startup, so a change to the set of listeners or to a listener's address, port or timeout is rejected with code restart_required at that listener's path; restart the Router to apply it. Listener api_keys change without a restart.

The Router records every activation, from any source, in one history of the last 10 versions (-config-history-limit changes it). Every configuration file has its own history, so a restart keeps it and Routers whose files share a directory never mix theirs: a workspace's config.yaml keeps it in .vllm-sr/config-backups beside the file, and any other file in a directory named after it inside that one. VLLM_SR_CONFIG_BACKUP_DIR moves it. The Helm chart keeps it on the models volume for one persistent replica, and each Pod keeps its own when replicas would share it; the other Kubernetes manifests keep it per Pod, since the ConfigMap is the document replicas share. Writers of one history take turns under a lock, so no version names two activations. A Router that restarts on the newest recorded document keeps its version. GET /api/v1/config/versions lists the history newest first. Each entry keeps version (its timestamp), timestamp, source and filename, and adds config_version, hash, active on the version that serves, and rollback_of on a rollback. POST /api/v1/config/rollback takes {"version": "<number or timestamp>"} or {"config_version": <number>}, persists the recorded document, and activates it as a new version whose rollback_of names the restored one. Versions that came from Kubernetes resources record no document; restore those at their source.

Tracing settings are initialized at process startup. Config plans, updates, and rollbacks that change global.services.observability.tracing return 409 RESTART_REQUIRED without persisting the candidate. Apply those changes through the deployment workflow and restart the Router.

Manage knowledge bases and stored data​

Knowledge-base configuration:

MethodPathUse
GET, POST/api/v1/storage/knowledge-basesList or create managed knowledge bases
GET, PUT, DELETE/api/v1/storage/knowledge-bases/{name}Read, update, or delete one knowledge base
GET/api/v1/storage/knowledge-bases/{name}/map/metadataRead generated map metadata
GET/api/v1/storage/knowledge-bases/{name}/map/data.ndjsonStream map data as NDJSON

In a router process, create, update, and delete persist a candidate and return 202 with activation_status: pending and generated_runtime_hash while the replacement generation prepares. Poll /api/v1/config/hash until active_runtime_hash matches that candidate, or stop on activation_status: failed and inspect activation. A failure already observed at mutation time returns 503 with the persisted candidate's activation detail. A second KB mutation while pending returns 409 with CONFIG_ACTIVATION_PENDING and does not overwrite it. Updates use independent asset revision paths; deletion removes the candidate config entry while retaining files needed by old and rollback generations. Old revisions require offline cleanup after their configuration references have been retired. Standalone API servers report activation_status: unknown with their normal success status because no router generation registry is present.

Router-managed storage and memory:

ResourceBase pathOperations
Long-term memory/api/v1/storage/memoriesList and delete by scope; read or delete by id
Vector stores/api/v1/storage/vector-storesCreate, list, read, update, delete, and search
Vector-store files/api/v1/storage/vector-stores/{id}/filesAttach, list, inspect, and detach files
Files/api/v1/storage/filesUpload, list, inspect, download, and delete

These routes return 503 when their required service is unavailable. File upload uses multipart form data and accepts documents (.txt, .md, .json, .csv, .html) for vector-store ingestion; upload an image (.png, .jpg, .jpeg, .gif, .webp) with purpose=vision to reference it from a Response API input_image part by file_id. Consult the live schema for limits and fields. They exist only on the management listener. /v1/files and /v1/vector_stores are not inference-listener aliases and are not registered by the Router API.

Discover and inspect plugins​

GET /api/v1/plugins projects the canonical plugin registry, including each plugin's schema URL, bindings URL, and supported operations. It includes all registered types, including plugins without a separate runtime operation. GET /api/v1/plugins/{type} selects one descriptor; GET /api/v1/plugins/{type}/bindings?recipe=<name>&decision=<name> inspects active bindings, declared enablement, reachability, and dependency availability. Availability is not a network health probe; runtime_inspected: false means that only configuration was available for inspection. Configuration links point back to the owning recipe. An empty binding list means the plugin is not configured in the selected scope.

Operation metadata distinguishes read, preview, probe, and mutation. A descriptor links to shared storage or observability resources where applicable; it does not duplicate their state under a plugin-specific resource tree.

Plugin behaviorOperation
system_prompt, request_params, header_mutation, fast_responsePOST /api/v1/plugins/{type}/preview using a typed candidate configuration or an active binding
hallucination, response_jailbreakPreview an active policy and its supplied conditions; mode: probe invokes its configured detectors
ragPreview context injection with supplied_context; mode: probe performs configured retrieval without the RAG result cache
tools, tool_selectionPreview static tool policy; semantic selection requires explicit mode: probe
context_compressionPreview retained targets and budgets using the existing compression contract
response_cache, memoryOperate their resources under /api/v1/storage
router_replay, shadow_dispatchInspect recorded outcomes under /api/v1/observability/replays

The guard, RAG, and tool previews require an explicit binding: {"recipe":"<name>","decision":"<name>"}. Their default mode is preview; they do not automatically escalate to a probe. A probe may invoke a configured local or remote classifier, embedding service, or retrieval backend. RAG previews require data.read, tool-policy previews require config.read, and guard probes require classify.invoke. These operations never execute a selected tool or generate a provider response. Returned mode, persisted, and backend_calls describe the operation: persisted: false means it does not write plugin data or conversation state; a probe may update runtime metrics or embedding memoization, and a configured external retrieval service may have its own effects. Header values in previews require secret_view.

Operate the response cache​

Response-cache endpoints are separate from inference-time cache lookup. They let operators inspect the backend, test a candidate configuration, and perform audited invalidation.

MethodPathUse
GET/api/v1/storage/response-cache/capabilitiesBackend capabilities
GET/api/v1/storage/response-cache/healthBackend health
GET/api/v1/storage/response-cache/statsRedacted statistics
POST/api/v1/storage/response-cache/testValidate and probe a candidate configuration
POST/api/v1/storage/response-cache/invalidateDry-run or invalidate a scoped partition
POST/api/v1/storage/response-cache/flushAdvance a scoped or global cache epoch

Prefer scoped invalidation and a dry run before a destructive cache mutation. Bearer roles distinguish read, invalidate, and broader cache-management permissions.

Inspect management audit​

GET /api/v1/observability/audit requires audit.read and covers audited management operations across config, recipes, data, cache, and compression. It also records how every configuration update ended, whatever its source: config.activate, config.reject and config.supersede entries carry a config object with the attempt, result, source, version and document hash, and for a rejection its stage and reason codes. Their request_id and role name the management request that caused the update, when one did; they have no method, path or status. Use an exact action filter, limit (1–1000, default 100), and after_sequence for pagination. Continue with the returned next_sequence and keep the same filter. has_more indicates more matching entries; truncated means the requested cursor predates retained data.

The hash-chained ring retains the latest 10,000 entries in this process. oldest_sequence identifies its retention boundary; the oldest entry can refer to a predecessor outside the ring. Restart resets the log and its sequence, so this endpoint is not a durable compliance archive. Audit records contain route and authorization metadata, not request bodies or plugin payloads.

Inspect context compression​

MethodPathUse
GET/api/v1/plugins/context_compression/capabilitiesRuntime capabilities
GET/api/v1/plugins/context_compression/healthRuntime health
GET/api/v1/observability/plugins/context_compression/statsRedacted statistics
POST/api/v1/plugins/context_compression/previewPreview compression without persistence
POST/api/v1/storage/context-recovery/invalidateInvalidate a trusted recovery scope

Use preview to evaluate what would be retained before enabling compression on important traffic.

Inspect replay and submit outcomes​

Router Replay is management-only. Its query endpoints and redaction model are described in Router API.

Router Learning can ingest an outcome linked to an owned replay record:

curl -sS http://localhost:8080/api/v1/observability/outcomes \
-H "Authorization: Bearer ${VSR_MGMT_TOKEN}" \
-H 'Content-Type: application/json' \
-H 'Idempotency-Key: feedback-123' \
-d '{
"replay_id": "replay_...",
"target": "model",
"verdict": "good_fit",
"score": 0.9
}'

replay_id, target, and verdict are required. The authenticated principal, not the optional source body field, determines provenance. Use a stable Idempotency-Key when a client may retry. Ingestion also requires an active Router Learning runtime and the learning.ingest permission when bearer auth is enabled.

API boundaries​

  • The management API is an operational surface, not the public inference gateway.
  • Endpoint availability can depend on compiled features and enabled services.
  • The OpenAPI document describes shape, not the behavior of a particular model, store, or external backend.
  • Keep bearer tokens out of URLs and logs. Give automation only the permissions it needs.

Complete endpoint index​

The following reference is generated from the Router's registered route catalog. Use it to scan every endpoint; use the task-oriented sections above for guidance and the running /openapi.json for exact schemas.

The former /api/v1/response-cache/* and /api/v1/context-compression/* routes are retired. Clients should rediscover operations from /api/v1 or /api/v1/plugins: cache operations now use /api/v1/storage/response-cache/*; compression capabilities, health, and preview use /api/v1/plugins/context_compression/*; statistics, recovery invalidation, and audit use the observability/storage paths above. No redirect or legacy handler changes a mutation's meaning.

system​

Health, readiness, and API contract discovery.

MethodPathDescription
GET/healthHealth check endpoint
GET/readyReadiness endpoint that turns green only after startup completes
GET/startup-statusDetailed router startup and model-download status
GET/api/v1/statusVersioned replica-local startup and configuration status; not a deployment probe
GET/api/v1Progressive API capability discovery
GET/openapi.jsonOpenAPI 3.0 specification; optionally narrowed to one path or operation
GET/docsInteractive Swagger UI documentation

config​

Validate, inspect, apply, version, and roll back Router configuration and Recipes.

MethodPathDescription
GET/api/v1/config/recipesList the default and named routing recipes with their entrypoints
POST/api/v1/config/recipes/validateValidate a recipe mutation without writing or reloading config
GET/api/v1/config/recipes/{name}Read one routing recipe and its entrypoints
PUT/api/v1/config/recipes/{name}Atomically create or replace one routing recipe; requires If-Match
DELETE/api/v1/config/recipes/{name}Delete an unreferenced named routing recipe; requires If-Match
GET/api/v1/config/schemaDiscover the canonical Router configuration contract progressively or return the complete JSON Schema
GET/api/v1/configGet the current router config as JSON (secrets redacted without secret_view)
POST/api/v1/config/validateValidate and normalize a router config without writing it
POST/api/v1/config/planPlan an exact merge or replace mutation, including hot-reload compatibility, without writing it
PATCH/api/v1/configCompare-and-swap merge of a router config update (validates, backs up, writes, triggers hot-reload)
PUT/api/v1/configCompare-and-swap replacement of the router config (validates, backs up, writes, triggers hot-reload)
POST/api/v1/config/rollbackCompare-and-swap rollback to a previous router config version
GET/api/v1/config/versionsList available router config backup versions
GET/api/v1/config/hashCompare persisted source, generated runtime, and active router config hashes

routing​

Preview routing behavior without generating an answer.

MethodPathDescription
POST/api/v1/routing/previewPreview configured signals and model selection without generating an answer. Supported native-output requests use backend render APIs to resolve per-candidate capacity; paths requiring execution remain unresolved. Learning uses read-only captured state with selection_provenance; preview_context supplies session identity and an optional preview-only sampling seed. A state-dependent or sampled result does not guarantee a later live selection. global.services.api.routing_preview controls the request deadline and concurrent worker bound.

inventory​

Inspect configured and loaded model and classifier resources.

MethodPathDescription
GET/api/v1/inventory/modelsGet information about loaded models
GET/api/v1/inventory/classifierGet classifier information and status (secrets redacted without secret_view)
GET/api/v1/inventory/embedding-modelsGet information about loaded embedding models
GET/api/v1/inventory/model-runtimeGet the model_runtime deployments: process, readiness, restarts and served model cards
GET/v1/modelsOpenAI-compatible public model and Entrypoint listing

observability​

Inspect routing replays, metrics, and management audit; submit outcome evidence.

MethodPathDescription
GET/api/v1/observability/classification-metricsGet classification metrics and statistics
POST/api/v1/observability/outcomesSubmit Router Learning outcome feedback linked to a replay record
GET/api/v1/observability/replaysList Router Replay records
GET/api/v1/observability/replays/aggregateAggregate Router Replay routing and cost metadata
GET/api/v1/observability/replays/trajectoryBuild a recipe-scoped session trajectory with each recorded routing result
GET/api/v1/observability/replays/datasetExport a shadow comparison dataset manifest built from the selected Router Replay records
GET/api/v1/observability/replays/{id}Read one Router Replay record
GET/api/v1/observability/auditPage through this Router process's bounded management mutation audit, configuration lifecycle outcomes included; filter by action and resume after a sequence
GET/api/v1/observability/plugins/context_compression/statsGet redacted context-compression statistics

storage​

Manage Router-owned knowledge bases, memories, files, vector stores, cache partitions, and context recovery.

MethodPathDescription
GET/api/v1/storage/response-cache/capabilitiesGet response-cache backend capabilities
GET/api/v1/storage/response-cache/healthCheck response-cache backend health
GET/api/v1/storage/response-cache/statsGet redacted response-cache statistics
POST/api/v1/storage/response-cache/testValidate and probe a response-cache candidate configuration
POST/api/v1/storage/response-cache/invalidateDry-run or invalidate a scoped response-cache partition
POST/api/v1/storage/response-cache/flushAdvance a scoped or global response-cache epoch
POST/api/v1/storage/context-recovery/invalidateInvalidate a trusted context-recovery request scope
GET/api/v1/storage/knowledge-basesList configured knowledge bases
POST/api/v1/storage/knowledge-basesCreate a managed knowledge base
GET/api/v1/storage/knowledge-bases/{name}Read a knowledge base
GET/api/v1/storage/knowledge-bases/{name}/map/metadataRead generated knowledge-base map metadata
GET/api/v1/storage/knowledge-bases/{name}/map/data.ndjsonStream generated knowledge-base map data as NDJSON
PUT/api/v1/storage/knowledge-bases/{name}Update a managed knowledge base
DELETE/api/v1/storage/knowledge-bases/{name}Delete a managed knowledge base
GET/api/v1/storage/memoriesList long-term memories
DELETE/api/v1/storage/memoriesDelete memories by scope
GET/api/v1/storage/memories/{id}Read one long-term memory
DELETE/api/v1/storage/memories/{id}Delete one long-term memory
POST/api/v1/storage/vector-storesCreate a vector store
GET/api/v1/storage/vector-storesList vector stores
GET/api/v1/storage/vector-stores/{id}Read a vector store
POST/api/v1/storage/vector-stores/{id}Update a vector store
DELETE/api/v1/storage/vector-stores/{id}Delete a vector store
POST/api/v1/storage/vector-stores/{id}/searchSearch a vector store
POST/api/v1/storage/vector-stores/{id}/filesAttach a file to a vector store
GET/api/v1/storage/vector-stores/{id}/filesList files attached to a vector store
DELETE/api/v1/storage/vector-stores/{id}/files/{file_id}Detach a file from a vector store
POST/api/v1/storage/filesUpload a file
GET/api/v1/storage/filesList uploaded files
GET/api/v1/storage/files/{id}Read uploaded-file metadata
DELETE/api/v1/storage/files/{id}Delete an uploaded file
GET/api/v1/storage/files/{id}/contentDownload uploaded-file content

plugins​

Discover recipe-scoped plugin bindings, dependencies, and typed behavior previews.

MethodPathDescription
GET/api/v1/plugins/context_compression/capabilitiesGet context-compression capabilities
GET/api/v1/plugins/context_compression/healthCheck context-compression runtime health
POST/api/v1/plugins/context_compression/previewPreview context compression without persistence
GET/api/v1/pluginsDiscover every registered recipe-scoped plugin, its schema, bindings, and supported operations
GET/api/v1/plugins/{type}Describe one canonical plugin type and its supported management operations
GET/api/v1/plugins/{type}/bindingsInspect active recipe and decision plugin bindings and published dependency availability; does not probe network health
POST/api/v1/plugins/system_prompt/previewPreview system instruction changes using the dispatch protocol codec; no persistence or backend calls
POST/api/v1/plugins/request_params/previewPreview the dispatch parameter policy, including blocked fields, defaults, and caps; no persistence or backend calls
POST/api/v1/plugins/header_mutation/previewPreview ordered Envoy header operations; values require secret_view; no persistence or backend calls
POST/api/v1/plugins/fast_response/previewPreview fixed assistant text before transport encoding; no persistence or backend calls
POST/api/v1/plugins/response_jailbreak/previewPreview an active binding's response-jailbreak policy; mode=probe explicitly invokes its configured classifier without generation or persistence
POST/api/v1/plugins/hallucination/previewPreview an active binding's hallucination policy and supplied fact-check/context conditions; mode=probe explicitly invokes configured detectors without generation or persistence
POST/api/v1/plugins/rag/previewPreview configured RAG context injection using supplied_context; mode=probe retrieves from the configured backend without the RAG result cache or generation
POST/api/v1/plugins/tools/previewPreview the configured tools policy; mode=probe permits semantic tool retrieval but never executes tools
POST/api/v1/plugins/tool_selection/previewPreview the configured tool_selection policy; mode=probe permits configured retrieval and embedding calls but never executes tools

diagnostics​

Inspect and invoke recipe-scoped prepared models, classifiers, embeddings, and rerankers.

MethodPathDescription
POST/api/v1/diagnostics/classify/intentClassify user queries into routing categories
POST/api/v1/diagnostics/classify/piiDetect personally identifiable information in text
POST/api/v1/diagnostics/classify/securityDetect jailbreak attempts and security threats
POST/api/v1/diagnostics/classify/fact-checkClassify if text needs fact-checking
POST/api/v1/diagnostics/classify/user-feedbackClassify user feedback type (satisfied, need_clarification, wrong_answer, want_different)
POST/api/v1/diagnostics/classify/combinedPerform combined classification (intent, PII, and security)
POST/api/v1/diagnostics/classify/batchBatch classification with configurable task_type parameter
POST/api/v1/diagnostics/embeddingsGenerate text, image, and audio embeddings
POST/api/v1/diagnostics/similarityCalculate pairwise text similarity
POST/api/v1/diagnostics/similarity/batchCalculate batch text-similarity matches
POST/api/v1/diagnostics/models/systemone/forwardForward a public native inference or discovery request through one active listener grant; requires management authorization and the original listener credentials
GET/api/v1/instanceRead the serving frontend capability mode and default native deployment
GET/api/v1/diagnostics/models/tasksList shared judgment task templates, structural model capabilities and binding provenance
GET/api/v1/diagnostics/models/systemoneList published model deployments and their native System One question capabilities
POST/api/v1/diagnostics/models/systemoneTest native System One questions against a published deployment; preserves choice, score, noul, set, span, usage and metadata; 32 MiB request, 4 MiB response, 30 second deadline
GET/api/v1/diagnostics/routes/systemoneList active native recipe entrypoints for operator diagnostics without probing models; availability describes the routing plan, not backend health
POST/api/v1/diagnostics/routes/systemoneRun a native recipe using operator classify.invoke permission; the selected algorithm owns its execution deadline and physical call budget, while signals use their own timeouts; independent of public listener grants
GET/api/v1/diagnostics/modelsList prepared model bindings in an explicitly selected recipe
POST/api/v1/diagnostics/models/labelsInspect a prepared label distribution; windowed bindings preserve their configured scan
POST/api/v1/diagnostics/models/label-scoresInspect independent label scores using the prepared operating point when configured
POST/api/v1/diagnostics/models/tokensInspect prepared token spans with original UTF-8 byte offsets and complete configured window scanning
POST/api/v1/diagnostics/models/embeddingsRun the explicitly selected prepared embedding binding at its published representation
POST/api/v1/diagnostics/models/rerankScore query-document pairs using the selected prepared relevance binding without running a RAG request; max_batch_size applies (default 100 pairs)

Testing the Decision Model​

GET /api/v1/diagnostics/models/systemone lists the published model deployments and the question types each model actually serves. POST to the same path runs native System One inference against the selected deployment. It does not change routing configuration or load another model.

{
"deployment": "your-deployed-decision-model",
"request": {
"state": "Explain how to optimize a SQL query.",
"questions": {
"task": {
"type": "choice",
"instructions": "What kind of work is requested?",
"criteria": {"coding": "Programming or databases", "other": "Other work"}
},
"difficulty": {
"type": "score",
"instructions": "How difficult is the request?",
"criteria": ["Simple", "Moderate", "Complex"]
},
"needs_facts": {"type": "noul", "instructions": "Does this require precise facts?"}
},
"options": {"return_meta": true}
}
}

The deployment ID comes from the capability list. The Router supplies its served model ID and forwards the native response, including probabilities, uncertainty, usage, per-question errors, Set answers and Span offsets where supported. Question and criteria order are preserved. The model runtime's /v1/systemone contract remains the source of truth for request fields and response semantics.

The test bypasses the Router result cache and records real model-runtime request, latency and server-phase metrics. Requests are limited to 2 MiB, responses to 4 MiB, and execution to 30 seconds. Discovery requires config.read; inference requires classify.invoke. The Dashboard's Build → System One → Decision Playground page uses its own authenticated gateway with config.read and evaluation.run.