Kubernetes Gateways
Semantic Router can run behind several Kubernetes gateways that call it through Envoy ExtProc. The routing policy stays the same; the gateway-specific resources determine when ExtProc runs, how the selected model reaches a backend, and which component owns authentication or traffic policy.
Choose an integration
| Existing data plane | Start with | What it owns |
|---|---|---|
| Agent Router (formerly Envoy AI Gateway) | Agent Router | Provider translation, provider credentials, rate limits, and Gateway API traffic policy. |
| agentgateway | agentgateway | Gateway API proxy, backend resources, and ExtProc phase policy. |
| Istio | Istio Gateway | Ingress, HTTPRoute processing, and the Envoy filter that calls Semantic Router. |
| Gateway API Inference Extension | GIE | InferencePool endpoint selection after Semantic Router chooses a model pool. |
Agent Router and agentgateway are the AI Gateway examples in this table. LiteLLM is also an AI Gateway, but this project does not have a LiteLLM guide.
Use the gateway already operated by your platform. Do not install a second gateway only to obtain semantic routing unless you have compared ownership, security policy, and upgrade requirements.
Shared contract
All integrations must agree on:
- the public model or entrypoint requested by the client;
- the model name written by Semantic Router;
- the Gateway API match or provider backend for that model; and
- the served model identity accepted by the inference endpoint.
The gateway owns client authentication and transport policy unless your deployment explicitly assigns those controls elsewhere. Semantic signals such as PII or jailbreak detection do not block traffic by themselves; a decision or plugin must enforce the intended action.
Request buffering and streaming
Start with the processing mode required by the selected integration. Change it only when request size or immediate streamed responses require it. See Streamed ExtProc for body buffering, mode override, and streaming constraints.
Verify before production
- send a direct request to each backend;
- send the same request through the gateway;
- inspect the selected-model and decision headers;
- confirm the gateway resolved the selected backend; and
- test credentials, request limits, streaming, and failure behavior.
Test a Kubernetes Gateway Deployment provides a common checklist without assuming cluster-assigned addresses.