Protocol Compatibility Matrix
Semantic Router supports three inference wire formats on both sides of the
data plane. A client request is decoded into a protocol-neutral form, routing
policy selects a model, and that model's api_format selects the backend codec.
The response is translated back to the client's original format.
client endpoint -> client codec -> routing -> backend codec -> model endpoint
Protocol compatibility is separate from target configuration and deployment support:
- use Backend Target Compatibility for URLs, weights, headers, discovery, and producer preservation; and
- use Deployment Support for project-maintained stacks, integrations, and hardware profiles.
For client connection settings, virtual model limits, and a tool-loop check, start with Connect an agent harness.
Client-facing protocols
| Client API | Inference endpoint | Buffered | Streaming | Availability |
|---|---|---|---|---|
| OpenAI Chat Completions | POST /v1/chat/completions | Supported | Supported | Available on the public inference listener. |
| OpenAI Responses | POST /v1/responses | Supported | Supported | Requires global.services.response_api and its store to be available. Router-owned object operations are not forwarded to a model backend. |
| Anthropic Messages | POST /v1/messages | Supported | Supported | Send the Anthropic request shape and an appropriate anthropic-version header. Client authentication remains deployment-specific. |
The public listener also serves GET /v1/models. See the
Router API for the complete method and path inventory,
Responses object operations, and request examples.
System One is a separate typed API, available in either startup mode when
listeners[].systemone.models publishes a concrete model. It uses
POST /v1/systemone (alias /v1/decisions) and
GET /v1/systemone/models. Router mode can also publish an explicit native
recipe through an api: systemone entrypoint and matching listener grant.
These typed requests do not use the Chat protocol-translation matrix or the
Chat model allowlist. See the
System One quickstart.
Backend model protocols
Set api_format on each providers.models[] entry. It describes the wire
contract implemented by that model endpoint, not the provider brand.
api_format | Backend request and response shape | Default upstream path | Notes |
|---|---|---|---|
openai | OpenAI Chat Completions | /v1/chat/completions | Default when api_format is omitted. backend_refs[].chat_path can override the Chat path. |
responses | OpenAI Responses | /v1/responses | The backend itself must implement the Responses wire contract; enabling the Router's Responses service does not add that API to a backend. |
anthropic | Anthropic Messages | /v1/messages | Configure provider authentication and required version headers on the backend ref. |
These fields are easy to confuse:
- model
api_formatselects the request, response, error, and streaming codec; - backend-ref
protocolselects HTTP or HTTPS transport; and - backend-ref
providersupplies provider-specific authentication and path defaults. It does not prove that the endpoint implements an API format.
For every backend format, base_url names the complete upstream API root. Its
path is retained and the protocol operation suffix is appended exactly once;
the protocol's default /v1 base path is used only when the URL has no path.
chat_path applies only to Chat Completions.
With --gateway extproc, the generated Envoy cluster for an HTTPS backend
verifies both the server
certificate chain and its DNS hostname. An HTTPS replica pool must keep one
hostname because the supported Envoy runtime shares its TLS context within a
cluster; use separate model aliases for different HTTPS hosts. IP-literal HTTPS
targets are rejected instead of silently weakening hostname verification. A
custom Envoy image used with vllm-sr serve must provide the system CA bundle at
/etc/ssl/certs/ca-certificates.crt; startup validation fails rather than
silently disabling verification when that trust store is unavailable.
Client-to-backend matrix
Every client format can route to every backend format. Each cell is covered in both buffered and streaming mode.
| Client protocol | openai backend | responses backend | anthropic backend |
|---|---|---|---|
| OpenAI Chat Completions | Supported | Supported through codec translation | Supported through codec translation |
| OpenAI Responses | Supported through codec translation | Supported | Supported through codec translation |
| Anthropic Messages | Supported through codec translation | Supported through codec translation | Supported |
"Supported" means the Router owns the request, response, transport-error, and streaming translation path. It does not mean every field from one protocol can be represented by every other protocol, or that every model behind an endpoint supports the requested capability.
Feature portability
The Router checks required semantics before encoding a backend request. A feature that the selected backend format cannot represent fails explicitly instead of being silently dropped.
| Semantic feature | Chat Completions | Responses | Messages |
|---|---|---|---|
| Text, image input, and file input | Supported | Supported | Supported |
| Tools, parallel tool calls, and strict tool schemas | Supported | Supported | Supported |
| Custom (free-form) tools and their calls | Supported | Supported | Not supported |
Text verbosity (low, medium, high) | Supported | Supported | Not forwarded; reported as dropped |
| Strict JSON Schema output | Supported | Supported | Supported |
| Buffered and streaming responses | Supported | Supported | Supported |
| Reasoning content and effort | Supported, except signed thinking blocks | Supported, except signed thinking blocks | Supported |
Reasoning summary requests (reasoning.summary) | Not forwarded; reported as dropped | Supported; reported as dropped if the provider uses chat_template_kwargs for reasoning controls | Not forwarded; reported as dropped |
| JSON object mode without a schema | Supported | Supported | Not supported |
| Audio input | Supported | Not supported | Not supported |
| Hosted image-generation lifecycle | Not supported | Supported | Not supported |
| Multiple response candidates | Supported | Not supported | Not supported |
| Prompt-cache directives | Supported | Not supported | Supported |
Prompt cache key (prompt_cache_key) | Supported | Supported | Not forwarded; reported as dropped |
| Reasoning token budget | Supported extension | Not supported | Supported |
| Seed and frequency or presence penalties | Supported | Not supported | Not supported |
top_k sampling | Supported extension | Not supported | Supported for nonnegative values |
min_p, repetition penalty, and cache salt | Supported extensions | Not supported | Not supported |
| Stop sequences | Supported | Not supported | Supported |
| Native response or conversation state fields | Not supported | Supported | Not supported |
This table describes codec representation, not model capability. For example, an OpenAI-compatible server can accept the Chat request shape while rejecting images or tools for a particular model. Qualify the actual endpoint and model revision before adding them to a routing pool.
For an Anthropic Messages client using a Chat Completions or Responses backend,
thinking.type: adaptive uses the backend model's default reasoning behavior;
output_config.effort is retained. thinking.display: omitted removes reasoning
from the translated response, including streaming output. Explicit
thinking.type: disabled requires a configured reasoning family and an
effective backend reasoning-off control. Unsupported controls fail with a typed
request error. A context_management edit of clear_thinking_20251015 with
keep: all has no effect and is omitted for these backends; edits that would
change history are rejected. Anthropic cache_control boundaries are omitted
when the selected backend uses Responses, which cannot represent them. The
prompt and tool result still dispatch, and x-vsr-protocol-warnings reports a
dropped diagnostic for cache_control.
Reasoning content translated out of an Anthropic backend fails when the backend
attaches a reasoning signature, which its thinking responses carry by default.
A Chat Completions or Responses client then receives a typed
unsupported_capability failure on buffered requests; on streaming requests the
failure arrives mid-stream, after the response headers and any earlier deltas,
because the signature reaches the encoder only with the reasoning delta.
Same-format traffic, including Messages clients reading an Anthropic backend, is
unaffected.
A Responses client can still use previous_response_id with a Chat
Completions or Messages backend. The Router retrieves and materializes the
retained history, removes Router-owned object controls, and then encodes the
stateless request in the selected backend format.
Responses custom tools use type: custom with a flattened format object.
Their custom_tool_call and custom_tool_call_output items retain free-form
input, tool results, and call IDs through Chat or Responses backends, including
buffered and streaming responses. A Messages backend rejects them with
unsupported_capability. Responses text.verbosity maps to the Chat
verbosity field. Messages has no equivalent, so the Router drops this output
detail hint and reports text.verbosity in x-vsr-protocol-warnings.
Configure a backend format
The client can use any supported client-facing endpoint; api_format controls
what the selected backend receives:
providers:
models:
- name: hosted/claude
provider_model_id: claude-model-id
api_format: anthropic
backend_refs:
- name: anthropic-primary
base_url: https://api.anthropic.com
provider: anthropic
api_key_env: ANTHROPIC_API_KEY
extra_headers:
anthropic-version: "2023-06-01"
weight: 100
api_format chooses only the backend codec. It does not imply Anthropic,
OpenAI, or any other runtime Provider. Router-owned listeners require a
physical model to declare backend_refs[].provider; metadata-only
listeners: [] configurations leave transport and credentials to the external
gateway. The local vllm-sr serve workflow owns backend transport through its
standalone frontend or generated Envoy configuration, so it does not accept a
backendless physical model; use the external-gateway deployment profile for
that topology.
Test the backend directly with its native path and a minimal request first. Then send the same semantic request through the Router using the client API that your agent harness or other API client needs. A successful health check does not validate request schema, streaming, tools, or error translation.
Validation and failure behavior
- Public requests are decoded into the neutral contract even when client and backend formats match. Unknown or unsupported request fields fail closed.
- Cross-protocol requests preserve shared semantics. Target-specific features that cannot be represented return a typed protocol error.
- Anthropic
tool_result.is_error: truehas no equivalent in OpenAI Chat or Responses requests. The default strict policy rejects this translation withlossy_translation; a false or absent flag remains supported. Anthropic backends preserve the flag. Codec callers explicitly usingLossyAllowWithDiagnosticmay omit the flag with atool_result.is_errordiagnostic; the tool-result content is not rewritten. - The response keeps the client protocol's JSON or SSE shape. Provider transport errors and incomplete streams are translated separately from successful model responses.
- For Anthropic-compatible providers, an absent nullable
stop_sequenceis interpreted as null with a bounded compatibility diagnostic. This applies to buffered Messages,message_start.message, andmessage_delta.delta. Explicit null needs no diagnostic; astop_sequencestop reason still requires a non-empty matched sequence, and terminal deltas still requirestop_reason. - When an OpenAI-compatible Chat provider names the matched stop string in
choices[].stop_reason, as vLLM does, Anthropic Messages clients receive astop_sequencestop reason with that sequence. Chat and Responses clients are unaffected. x-vsr-client-protocol,x-vsr-upstream-protocol, andx-vsr-protocol-warningsexpose translation details when applicable. See VSR routing headers. Diagnostics discovered after streaming headers have been sent are recorded in translation-warning metrics and structured debug logs; they cannot be added to the already-sent response header.
The repository verifies all three protocols pairwise in codec tests, at the Envoy ExtProc boundary, and in an 18-cell deployment matrix: three client formats by three backend formats by buffered or streaming mode. See the implemented codec design for the full verification and extension contract.
Codex CLI
Codex CLI uses the Responses endpoint and sends several compatibility fields. The Router handles them as follows:
| Field | Router behavior |
|---|---|
prompt_cache_key | Forwarded to openai and responses backends. A route to an anthropic backend omits it and reports dropped, because Messages has no equivalent. A decision's request_params.blocked_params can remove it before dispatch. |
include: ["reasoning.encrypted_content"] | Accepted and not forwarded. The Router never relays provider-encrypted reasoning, so reasoning items carry no encrypted_content. Other include values remain unsupported. |
client_metadata | Accepted and not forwarded, because it carries Codex telemetry rather than model input. |
text.verbosity | Forwarded to responses backends, mapped to verbosity for openai Chat backends, and reported as dropped for anthropic Messages backends. The only accepted values are low, medium, and high. |
Each accepted but unforwarded field appears as a dropped entry in
x-vsr-protocol-warnings.
Codex can also request reasoning summaries. The Router forwards
reasoning.summary to a Responses backend. A Chat Completions or Messages
backend cannot request a summary, so the Router accepts the turn and reports
the dropped setting in x-vsr-protocol-warnings. Multi-agent namespace tools
and hosted web search remain unsupported; disable those in the Codex
config.toml that points at the Router. These settings were checked with
Codex CLI 0.156.1:
web_search = "disabled"
[features]
multi_agent = false