Skip to main content
Version: Latest

Kubernetes Gateways

Semantic Router can run behind several Kubernetes gateways that call it through Envoy ExtProc. The routing policy stays the same; the gateway-specific resources determine when ExtProc runs, how the selected model reaches a backend, and which component owns authentication or traffic policy.

Choose an integration​

Existing data planeStart withWhat it owns
Agent Router (formerly Envoy AI Gateway)Agent RouterProvider translation, provider credentials, rate limits, and Gateway API traffic policy.
agentgatewayagentgatewayGateway API proxy, backend resources, and ExtProc phase policy.
IstioIstio GatewayIngress, HTTPRoute processing, and the Envoy filter that calls Semantic Router.
Gateway API Inference ExtensionGIEInferencePool endpoint selection after Semantic Router chooses a model pool.

Agent Router and agentgateway are the AI Gateway examples in this table. LiteLLM is also an AI Gateway, but this project does not have a LiteLLM guide.

Use the gateway already operated by your platform. Do not install a second gateway only to obtain semantic routing unless you have compared ownership, security policy, and upgrade requirements.

Shared contract​

All integrations must agree on:

  1. the public model or entrypoint requested by the client;
  2. the model name written by Semantic Router;
  3. the Gateway API match or provider backend for that model; and
  4. the served model identity accepted by the inference endpoint.

The gateway owns client authentication and transport policy unless your deployment explicitly assigns those controls elsewhere. Semantic signals such as PII or jailbreak detection do not block traffic by themselves; a decision or plugin must enforce the intended action.

Request buffering and streaming​

Start with the processing mode required by the selected integration. Change it only when request size or immediate streamed responses require it. See Streamed ExtProc for body buffering, mode override, and streaming constraints.

Verify before production​

  • send a direct request to each backend;
  • send the same request through the gateway;
  • inspect the selected-model and decision headers;
  • confirm the gateway resolved the selected backend; and
  • test credentials, request limits, streaming, and failure behavior.

Test a Kubernetes Gateway Deployment provides a common checklist without assuming cluster-assigned addresses.