Skip to main content
Version: Latest (unreleased)

Protocol Compatibility Matrix

Semantic Router supports three inference wire formats on both sides of the data plane. A client request is decoded into a protocol-neutral form, routing policy selects a model, and that model's api_format selects the backend codec. The response is translated back to the client's original format.

client endpoint -> client codec -> routing -> backend codec -> model endpoint

Protocol compatibility is separate from target configuration and deployment support:

For client connection settings, virtual model limits, and a tool-loop check, start with Connect an agent harness.

Client-facing protocols​

Client APIInference endpointBufferedStreamingAvailability
OpenAI Chat CompletionsPOST /v1/chat/completionsSupportedSupportedAvailable on the public inference listener.
OpenAI ResponsesPOST /v1/responsesSupportedSupportedRequires global.services.response_api and its store to be available. Router-owned object operations are not forwarded to a model backend.
Anthropic MessagesPOST /v1/messagesSupportedSupportedSend the Anthropic request shape and an appropriate anthropic-version header. Client authentication remains deployment-specific.

The public listener also serves GET /v1/models. See the Router API for the complete method and path inventory, Responses object operations, and request examples.

System One is a separate typed API, available in either startup mode when listeners[].systemone.models publishes a concrete model. It uses POST /v1/systemone (alias /v1/decisions) and GET /v1/systemone/models. Router mode can also publish an explicit native recipe through an api: systemone entrypoint and matching listener grant. These typed requests do not use the Chat protocol-translation matrix or the Chat model allowlist. See the System One quickstart.

Backend model protocols​

Set api_format on each providers.models[] entry. It describes the wire contract implemented by that model endpoint, not the provider brand.

api_formatBackend request and response shapeDefault upstream pathNotes
openaiOpenAI Chat Completions/v1/chat/completionsDefault when api_format is omitted. backend_refs[].chat_path can override the Chat path.
responsesOpenAI Responses/v1/responsesThe backend itself must implement the Responses wire contract; enabling the Router's Responses service does not add that API to a backend.
anthropicAnthropic Messages/v1/messagesConfigure provider authentication and required version headers on the backend ref.

These fields are easy to confuse:

  • model api_format selects the request, response, error, and streaming codec;
  • backend-ref protocol selects HTTP or HTTPS transport; and
  • backend-ref provider supplies provider-specific authentication and path defaults. It does not prove that the endpoint implements an API format.

For every backend format, base_url names the complete upstream API root. Its path is retained and the protocol operation suffix is appended exactly once; the protocol's default /v1 base path is used only when the URL has no path. chat_path applies only to Chat Completions.

With --gateway extproc, the generated Envoy cluster for an HTTPS backend verifies both the server certificate chain and its DNS hostname. An HTTPS replica pool must keep one hostname because the supported Envoy runtime shares its TLS context within a cluster; use separate model aliases for different HTTPS hosts. IP-literal HTTPS targets are rejected instead of silently weakening hostname verification. A custom Envoy image used with vllm-sr serve must provide the system CA bundle at /etc/ssl/certs/ca-certificates.crt; startup validation fails rather than silently disabling verification when that trust store is unavailable.

Client-to-backend matrix​

Every client format can route to every backend format. Each cell is covered in both buffered and streaming mode.

Client protocolopenai backendresponses backendanthropic backend
OpenAI Chat CompletionsSupportedSupported through codec translationSupported through codec translation
OpenAI ResponsesSupported through codec translationSupportedSupported through codec translation
Anthropic MessagesSupported through codec translationSupported through codec translationSupported

"Supported" means the Router owns the request, response, transport-error, and streaming translation path. It does not mean every field from one protocol can be represented by every other protocol, or that every model behind an endpoint supports the requested capability.

Feature portability​

The Router checks required semantics before encoding a backend request. A feature that the selected backend format cannot represent fails explicitly instead of being silently dropped.

Semantic featureChat CompletionsResponsesMessages
Text, image input, and file inputSupportedSupportedSupported
Tools, parallel tool calls, and strict tool schemasSupportedSupportedSupported
Custom (free-form) tools and their callsSupportedSupportedNot supported
Text verbosity (low, medium, high)SupportedSupportedNot forwarded; reported as dropped
Strict JSON Schema outputSupportedSupportedSupported
Buffered and streaming responsesSupportedSupportedSupported
Reasoning content and effortSupported, except signed thinking blocksSupported, except signed thinking blocksSupported
Reasoning summary requests (reasoning.summary)Not forwarded; reported as droppedSupported; reported as dropped if the provider uses chat_template_kwargs for reasoning controlsNot forwarded; reported as dropped
JSON object mode without a schemaSupportedSupportedNot supported
Audio inputSupportedNot supportedNot supported
Hosted image-generation lifecycleNot supportedSupportedNot supported
Multiple response candidatesSupportedNot supportedNot supported
Prompt-cache directivesSupportedNot supportedSupported
Prompt cache key (prompt_cache_key)SupportedSupportedNot forwarded; reported as dropped
Reasoning token budgetSupported extensionNot supportedSupported
Seed and frequency or presence penaltiesSupportedNot supportedNot supported
top_k samplingSupported extensionNot supportedSupported for nonnegative values
min_p, repetition penalty, and cache saltSupported extensionsNot supportedNot supported
Stop sequencesSupportedNot supportedSupported
Native response or conversation state fieldsNot supportedSupportedNot supported

This table describes codec representation, not model capability. For example, an OpenAI-compatible server can accept the Chat request shape while rejecting images or tools for a particular model. Qualify the actual endpoint and model revision before adding them to a routing pool.

For an Anthropic Messages client using a Chat Completions or Responses backend, thinking.type: adaptive uses the backend model's default reasoning behavior; output_config.effort is retained. thinking.display: omitted removes reasoning from the translated response, including streaming output. Explicit thinking.type: disabled requires a configured reasoning family and an effective backend reasoning-off control. Unsupported controls fail with a typed request error. A context_management edit of clear_thinking_20251015 with keep: all has no effect and is omitted for these backends; edits that would change history are rejected. Anthropic cache_control boundaries are omitted when the selected backend uses Responses, which cannot represent them. The prompt and tool result still dispatch, and x-vsr-protocol-warnings reports a dropped diagnostic for cache_control.

Reasoning content translated out of an Anthropic backend fails when the backend attaches a reasoning signature, which its thinking responses carry by default. A Chat Completions or Responses client then receives a typed unsupported_capability failure on buffered requests; on streaming requests the failure arrives mid-stream, after the response headers and any earlier deltas, because the signature reaches the encoder only with the reasoning delta. Same-format traffic, including Messages clients reading an Anthropic backend, is unaffected.

A Responses client can still use previous_response_id with a Chat Completions or Messages backend. The Router retrieves and materializes the retained history, removes Router-owned object controls, and then encodes the stateless request in the selected backend format.

Responses custom tools use type: custom with a flattened format object. Their custom_tool_call and custom_tool_call_output items retain free-form input, tool results, and call IDs through Chat or Responses backends, including buffered and streaming responses. A Messages backend rejects them with unsupported_capability. Responses text.verbosity maps to the Chat verbosity field. Messages has no equivalent, so the Router drops this output detail hint and reports text.verbosity in x-vsr-protocol-warnings.

Configure a backend format​

The client can use any supported client-facing endpoint; api_format controls what the selected backend receives:

providers:
models:
- name: hosted/claude
provider_model_id: claude-model-id
api_format: anthropic
backend_refs:
- name: anthropic-primary
base_url: https://api.anthropic.com
provider: anthropic
api_key_env: ANTHROPIC_API_KEY
extra_headers:
anthropic-version: "2023-06-01"
weight: 100

api_format chooses only the backend codec. It does not imply Anthropic, OpenAI, or any other runtime Provider. Router-owned listeners require a physical model to declare backend_refs[].provider; metadata-only listeners: [] configurations leave transport and credentials to the external gateway. The local vllm-sr serve workflow owns backend transport through its standalone frontend or generated Envoy configuration, so it does not accept a backendless physical model; use the external-gateway deployment profile for that topology.

Test the backend directly with its native path and a minimal request first. Then send the same semantic request through the Router using the client API that your agent harness or other API client needs. A successful health check does not validate request schema, streaming, tools, or error translation.

Validation and failure behavior​

  • Public requests are decoded into the neutral contract even when client and backend formats match. Unknown or unsupported request fields fail closed.
  • Cross-protocol requests preserve shared semantics. Target-specific features that cannot be represented return a typed protocol error.
  • Anthropic tool_result.is_error: true has no equivalent in OpenAI Chat or Responses requests. The default strict policy rejects this translation with lossy_translation; a false or absent flag remains supported. Anthropic backends preserve the flag. Codec callers explicitly using LossyAllowWithDiagnostic may omit the flag with a tool_result.is_error diagnostic; the tool-result content is not rewritten.
  • The response keeps the client protocol's JSON or SSE shape. Provider transport errors and incomplete streams are translated separately from successful model responses.
  • For Anthropic-compatible providers, an absent nullable stop_sequence is interpreted as null with a bounded compatibility diagnostic. This applies to buffered Messages, message_start.message, and message_delta.delta. Explicit null needs no diagnostic; a stop_sequence stop reason still requires a non-empty matched sequence, and terminal deltas still require stop_reason.
  • When an OpenAI-compatible Chat provider names the matched stop string in choices[].stop_reason, as vLLM does, Anthropic Messages clients receive a stop_sequence stop reason with that sequence. Chat and Responses clients are unaffected.
  • x-vsr-client-protocol, x-vsr-upstream-protocol, and x-vsr-protocol-warnings expose translation details when applicable. See VSR routing headers. Diagnostics discovered after streaming headers have been sent are recorded in translation-warning metrics and structured debug logs; they cannot be added to the already-sent response header.

The repository verifies all three protocols pairwise in codec tests, at the Envoy ExtProc boundary, and in an 18-cell deployment matrix: three client formats by three backend formats by buffered or streaming mode. See the implemented codec design for the full verification and extension contract.

Codex CLI​

Codex CLI uses the Responses endpoint and sends several compatibility fields. The Router handles them as follows:

FieldRouter behavior
prompt_cache_keyForwarded to openai and responses backends. A route to an anthropic backend omits it and reports dropped, because Messages has no equivalent. A decision's request_params.blocked_params can remove it before dispatch.
include: ["reasoning.encrypted_content"]Accepted and not forwarded. The Router never relays provider-encrypted reasoning, so reasoning items carry no encrypted_content. Other include values remain unsupported.
client_metadataAccepted and not forwarded, because it carries Codex telemetry rather than model input.
text.verbosityForwarded to responses backends, mapped to verbosity for openai Chat backends, and reported as dropped for anthropic Messages backends. The only accepted values are low, medium, and high.

Each accepted but unforwarded field appears as a dropped entry in x-vsr-protocol-warnings.

Codex can also request reasoning summaries. The Router forwards reasoning.summary to a Responses backend. A Chat Completions or Messages backend cannot request a summary, so the Router accepts the turn and reports the dropped setting in x-vsr-protocol-warnings. Multi-agent namespace tools and hosted web search remain unsupported; disable those in the Codex config.toml that points at the Router. These settings were checked with Codex CLI 0.156.1:

web_search = "disabled"

[features]
multi_agent = false