Skip to main content
Version: Latest (unreleased)

Router API

The router data plane accepts model requests on the configured listeners. The Router serves them itself in standalone mode, the default; with --gateway extproc, Envoy serves them and calls the Router over ext_proc. In the standard local stack, the listener is http://localhost:8899; a recipe can choose a different address or port under listeners.

Use the data plane for inference. Use the management API, normally bound to 127.0.0.1:8080, for health checks, configuration, diagnostics, and replay queries. See Router management API.

Supported inference paths​

MethodPathClient formatNotes
POST/v1/chat/completionsOpenAI Chat CompletionsMain routed inference endpoint
POST/v1/responsesOpenAI ResponsesRequires the Responses service to be enabled
GET/v1/responses/{id}OpenAI ResponsesReads a stored response
DELETE/v1/responses/{id}OpenAI ResponsesDeletes a stored response
GET/v1/responses/{id}/input_itemsOpenAI ResponsesReads stored input items
POST/v1/messagesAnthropic MessagesThe router translates when the selected backend uses another protocol
POST/openai/deployments/{deployment}/chat/completionsAzure OpenAI Chat CompletionsThe deployment names the Router model; api-version is accepted and not forwarded
POST/openai/responsesAzure OpenAI ResponsesAccepts the dated api-version; the model is in the request body and the Responses service must be enabled
POST/openai/v1/responsesAzure OpenAI ResponsesThe model is in the request body and the Responses service must be enabled
POST/openai/v1/chat/completionsAzure OpenAI Chat CompletionsThe model is in the request body
GET/v1/modelsOpenAI ModelsLists models exposed by the active router configuration
POST/v1/systemone, /v1/decisionsNative System OneDirect decision models or explicitly published native recipes on a standalone listener
GET/v1/systemone/modelsNative model discoveryLists that listener's published System One models

Engine mode serves the native System One paths with recipe routing disabled. Publish native model IDs in listeners[].systemone.models; the Chat models allowlist does not grant native access. Both APIs use the listener's API keys. See the model runtime quickstart for a request.

Route a System One request​

In Router mode, publish an entrypoints item with api: systemone and grant its name in listeners[].systemone.models. The Chat default vllm-sr/auto does not automatically publish a native entrypoint. Follow the System One cascade guide to connect local deployments or remote Engine and compatible System One services.

curl -sS http://localhost:8899/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"state": "Please explain how to reset my password.",
"questions": {
"task": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"account": "Account access or authentication",
"billing": "Payments, invoices or refunds"
}
}
}
}'

Auto currently accepts explicit choice, score and noul questions, including multiple named states. The selected native response retains its answers, probabilities, usage and other fields. model stays the public entrypoint; the additional routing object identifies the execution. Set options.return_meta: true to include the selected model's runtime metadata. Calibrated cascades collect this provenance internally even when you leave response metadata disabled.

{
"recipe": "native-cascade",
"decision": "classify",
"algorithm": "cascade",
"stage": "strong",
"selected_model": "vega",
"quality": "uncalibrated",
"model_calls": 2
}

model_calls counts physical inference attempts within the selected algorithm, including its transport retries. Signal calls are outside this counter and have their own timeouts and request cancellation. It is not a count of GPU forwards. usage belongs to the returned native candidate and is not total cascade billing. The quality field names the configured acceptance method; it is not an accuracy score. Native discovery marks recipe entrypoints with routing: true and concrete models with routing: false.

Malformed or unsupported auto questions return 400. If no complete answer passes the declared acceptance rule within the call budget, the request returns 503 with systemone_unresolved; expiration of the algorithm deadline returns 504 with systemone_deadline_exceeded. Failed candidates never become empty successful answers. Direct model requests keep their existing native API contract and do not run the cascade.

A remote Engine backend must expose a concrete model. Native backend calls carry X-VSR-SystemOne-Backend: 1; a receiving Router returns 409 (systemone_nested_routing) if that request targets another recipe or remote forwarding alias. This prevents recursive cascades. The header grants no access: the receiving listener still checks its API key and model allowlist. An external compatible provider's internal execution remains outside the Router's call ledger.

Auto execution exports sr_systemone_stage_total and sr_systemone_stage_duration_seconds by configured algorithm, stage and model, plus sr_systemone_auto_requests_total and sr_systemone_auto_duration_seconds. These distinguish accepted, rejected, invalid and failed attempts; they measure serving behavior, not label-based evaluation quality. The auto duration covers the algorithm; preceding recipe signals are outside that metric.

Other /v1/* paths fail closed. In particular, /v1/files, /v1/vector_stores, and Router Replay paths are not available on a public inference listener. Router-owned file and vector-store operations use /api/v1/storage/files and /api/v1/storage/vector-stores on the management listener. Other /openai/* operations, such as embeddings and stored-response reads, return 404.

Every POST path above requires a JSON request body. A request with an empty body returns 400 from the Router and is never forwarded to a backend.

See Protocol Compatibility for the client-to-backend translation matrix, backend api_format values, and field-level portability boundaries.

Send a routed request​

Use vllm-sr/auto or an explicitly declared recipe entrypoint when you want the router to select a backend. Use a concrete model name when you want to bypass semantic model selection and target that model directly.

curl -sS http://localhost:8899/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"messages": [
{
"role": "user",
"content": "Write a Python function that merges two sorted lists."
}
]
}'

The response keeps the client protocol's shape. Its model, content, token usage, and optional router headers depend on the selected backend and recipe. See VSR routing headers for the stable observability contract.

The model names accepted by the Router come from canonical provider entries. name is the logical alias used by decisions and clients, provider_model_id is sent to the upstream provider, and providers.models[].backend_refs[] identifies the physical endpoint:

providers:
models:
- name: local-small
provider_model_id: served-model
api_format: openai
pricing:
currency: USD
prompt_per_1m: 0
completion_per_1m: 0
backend_refs:
- name: local-vllm
endpoint: model-server:8000
protocol: http
provider: vllm
weight: 1

Pricing is operator-supplied deployment metadata, not a live quote. It stays on providers.models[]; routing.modelCards only describes semantic capabilities. currency is optional and resolves to USD for accounting when omitted. When set, it must be an uppercase three-letter code. All per-million-token rates must be finite and non-negative. cached_input_per_1m and cache_write_per_1m are optional, and an explicit zero represents a free rate.

api_format declares the upstream wire contract: openai for Chat Completions, responses for the OpenAI Responses API, or anthropic for Anthropic Messages. The client may use any supported inference path; the Router translates once at the provider boundary and returns the client's original wire format.

vLLM Chat controls​

For Chat backends that implement these vLLM extensions, the Router preserves these fields through routing and request edits:

FieldAccepted values
top_kInteger: -1 or 0 disables filtering; positive values limit candidate tokens.
min_pNumber from 0 to 1.
repetition_penaltyFinite number greater than 0.
cache_saltString of 1–128 characters, excluding @, /, \, and NUL.

cache_salt selects a backend prefix-cache namespace without changing the prompt. Reuse a salt for requests that may share cached prefixes. The Router also preserves chat_template_kwargs, such as enable_thinking. These extensions are rejected when the target protocol cannot represent them.

Responses API​

curl -sS http://localhost:8899/v1/responses \
-H 'Content-Type: application/json' \
-d '{
"model": "vllm-sr/auto",
"input": "Summarize the trade-offs of retrieval-augmented generation."
}'

Creating, retrieving, and deleting Responses API objects requires its backing service and store. When Responses API support is disabled, the collection endpoint returns 404 and stored-object handling is unavailable. A configured service retains objects according to its own storage and retention settings.

Anthropic Messages​

curl -sS http://localhost:8899/v1/messages \
-H 'Content-Type: application/json' \
-H 'anthropic-version: 2023-06-01' \
-d '{
"model": "vllm-sr/auto",
"max_tokens": 256,
"messages": [
{
"role": "user",
"content": "Explain semantic routing in one paragraph."
}
]
}'

Azure OpenAI clients​

Clients built for Azure OpenAI can call the Router as if it were an Azure resource. The deployment Chat path takes the model name from the URL; the Responses and v1 Chat paths take it from the request body. In either form, A recipe entrypoint selects a route and a concrete model name targets that model directly. The Router checks the client's api-key header against the listener's api_keys when they are set, and removes the header before provider dispatch.

curl -sS 'http://localhost:8899/openai/v1/chat/completions' \
-H 'Content-Type: application/json' \
-H 'api-key: YOUR-LISTENER-KEY' \
-d '{"model":"vllm-sr/auto","messages":[{"role":"user","content":"Explain semantic routing in one paragraph."}]}'

For GitHub Copilot CLI, set COPILOT_PROVIDER_TYPE=azure, point COPILOT_PROVIDER_BASE_URL at the listener, and set COPILOT_PROVIDER_WIRE_MODEL to the Router model name. With COPILOT_PROVIDER_WIRE_API=responses, the CLI uses /openai/v1/responses, or /openai/responses when COPILOT_PROVIDER_AZURE_API_VERSION is set. The supported Chat paths are /openai/v1/chat/completions and the deployment route. The Router accepts Copilot's reasoning.summary on Responses turns. It forwards the setting to a Responses backend. With a Chat Completions or Messages backend, the turn still runs, but the summary request is dropped and reported in x-vsr-protocol-warnings.

Protocol translation is limited to fields the router supports. When a request crosses protocols, inspect x-vsr-client-protocol, x-vsr-upstream-protocol, and any x-vsr-protocol-warnings response header.

Routing errors​

When the Router cannot route a request, it answers the request itself and calls no backend. The error uses the client's protocol. In OpenAI Chat Completions and Responses errors, error.code is a stable reason code and error.message a short message. Apart from the budget errors, the message names no model, decision, or request content:

{"error":{"type":"invalid_request_error","code":"no_route","message":"no route matched the request","param":null}}
CodeStatuserror.typeMeaning
model_not_found400invalid_request_errorThe request names a model this Router does not serve.
no_route400invalid_request_errorNo decision matched, and no default model applies. A recipe entrypoint falls back to providers.defaults.model when it is configured. Looper entrypoints follow the same recipe rules; their names do not select an algorithm.
context_length_exceeded400 or 422invalid_request_errorThe request does not fit the models that could serve it: 400 from the request budget check, 422 from the models' context_window_size.
max_output_tokens_exceeded400invalid_request_errorThe requested output exceeds the configured model limit. See request budget errors.
decision_unresolved503server_errorA decision could not be evaluated because a signal it needs was unavailable, and its rules.on_unknown is fail_request. x-vsr-applied-unknown-policy names the decision.
no_eligible_model503server_errorThe selection policy rejected every candidate model of the matched decision.

The Router logs each of these failures at WARN, under the request's x-request-id, with the code and its own reason: the model the request named, the recipe and decision it reached, and the error. Anthropic Messages clients get the same status and message in Anthropic's error envelope, which has no code field.

Request budget errors​

With candidate_requirements.context: known_limits, the Router checks estimated input plus the effective output allowance against the candidates' configured limits. If all candidates fail only the budget check, it returns HTTP 400:

Error codeMeaning
context_length_exceededThe prepared input and requested output do not fit.
max_output_tokens_exceededThe requested output exceeds the configured model limit.

Missing capabilities, unknown limits, unavailable selection evidence, and mixed failures retain their selection-error behavior, no_eligible_model. Budget checks do not truncate requests by themselves; opt into context compression when appropriate.

These counts are estimates. A backend can still reject a request; its valid HTTP status and meaningful message are retained. vLLM integer codes are exposed as strings in OpenAI-compatible errors: BadRequestError with code: 400 becomes invalid_request_error with code: "400".

A streaming request rejected before generation receives the same non-2xx JSON error, not a successful SSE stream. When Replay is enabled, it records the failed status and body; Router budget rejections use terminal_reason: request_budget_exceeded.

Router Replay​

Router Replay records routing decisions and selected request lifecycle data. It is useful for debugging, evaluation, and Router Learning, but it does not change routing merely because a record is read.

Replay is disabled unless the service is enabled:

global:
services:
router_replay:
enabled: true
store_backend: memory

The in-memory backend is suitable for local inspection. Use a configured persistent backend when records must survive process restarts, and set retention appropriate to the data being captured.

Replay queries go to the management API:

curl -sS 'http://localhost:8080/api/v1/observability/replays?limit=20' \
-H "Authorization: Bearer ${VSR_MGMT_TOKEN}"
MethodPathPurpose
GET/api/v1/observability/replaysList and filter records
GET/api/v1/observability/replays/{id}Read one record
GET/api/v1/observability/replays/aggregateAggregate routing and cost metadata
GET/api/v1/observability/replays/trajectory?session_id=...&recipe=...Reconstruct one recipe's session trajectory

List and aggregate requests accept filters such as recipe, decision, model, session_id, cache_status, and search. Pagination uses limit and offset; limit is capped at 100. showDetails=true requests large body fields, so use it only when those fields are needed.

Trajectory queries use the exact recipe name. Omitting recipe is supported only when the session's records belong to one recipe; an ambiguous session returns 400. An explicit empty recipe= selects older, unscoped records. The response includes each request's route, latency, and lifecycle, including multiple requests at the same turn index.

Records, trajectory routes, and messages include conversation_id when an explicit conversation identity is available. Messages are grouped by conversation and turn, so separate conversations in one session can each start at turn zero. Insights shows their boundaries and complete IDs.

Dashboard Insights shows these routes alongside recorded signals, projections, candidate scores, and session-switch reasons. In observe mode, candidate and hold explanations describe what protection would have done; the selected model and route history still describe actual dispatch. Protection's candidate_models lists eligible models independently of their scores; an unrecorded score appears as —, while a recorded zero remains zero. Missing identity or evidence is displayed explicitly. Replay capture uses global.services.router_replay defaults and the selected decision's router_replay plugin overrides. Rejected requests without a selected decision use the global defaults. capture_personal_data: false retains routing evidence but suppresses content when personal data is detected or PII evidence is unavailable.

Configured-rate cost estimates​

Insights uses recorded token usage and configured input, cached-input, cache-write, and output rates. These are estimates, not invoices or GPU running costs; infrastructure charges and billing adjustments are excluded.

FieldMeaning
actual_costEstimated cost of the selected model for the recorded usage.
baseline_costHighest same-currency estimate in the selected recipe's model pool, using that same usage.
baseline_modelThe model used for that comparison. Equal costs use model-name order.
cost_savingsThe difference between the baseline and actual estimate.
currencyCurrency of the recorded estimate; no exchange-rate conversion is applied.

The baseline covers the recipe's models across all its decisions: model references, explicit candidate-iteration models, and route destinations. A permitted default-model fallback is included only for decisions that can use it; auxiliary planners and judges do not enlarge the pool. Other recipes, unpriced models, and different currencies are excluded. Direct requests without a recipe compare against themselves. Alternative-model eligibility and tokenization are not re-evaluated; this is a rate comparison, not another inference.

Missing usage, price, currency, or baseline stays unknown, with the reason shown in Insights. Explicitly configured free rates remain zero. Cache hits have zero additional model-inference cost; cache storage and lookup costs are excluded. Existing records keep their captured prices and baseline. Price not recorded means the historical record has no price; configuring rates today does not backfill it.

Aggregates use summary.by_currency, a sorted array with currency, total_saved, baseline_spend, actual_spend, and cost_record_count for each currency. With one currency, the flat summary fields mirror that group. With multiple currencies, flat currency is omitted and flat amounts are zero placeholders: use by_currency, not those placeholders. There is no combined cross-currency total.

cost_record_count counts complete estimates. excluded_record_count includes non-completed requests and records missing the data needed for a complete estimate; it does not imply that every excluded request lacks pricing.

When bearer authentication is enabled, replay callers need replay.read. Prompt, response, tool, and other sensitive details remain redacted unless the principal also has replay.detail. Treat replay storage as potentially sensitive even when the API normally returns a redacted view.

Replay lifecycle values describe what the recorder observed:

  • in_progress: no terminal response frame has been recorded yet.
  • completed: the response finished normally.
  • failed: routing or the upstream response failed.
  • aborted: the stream ended without a valid terminal frame, for example after a disconnect or timeout.

An HTTP 200 response header alone does not make a streaming record completed.

Which port should I use?​

TaskSurface
Send model trafficStandalone frontend, or Envoy with --gateway extproc; 8899 in the standard local stack
List public modelsGET /v1/models on the inference listener
Check health or readinessManagement API on 8080
Read or change configurationManagement API on 8080
Inspect replay recordsManagement API on 8080
Manage Router-owned files or vector storesManagement API on 8080 under /api/v1/storage/*

Do not expose the management port as a substitute for the public inference listener. Its endpoints can reveal configuration and operational data or make state-changing requests.