Deploy with Agent Router
Use this topology when Agent Router (formerly Envoy AI Gateway) already owns north-south traffic and provider integration. Semantic Router can either choose a model from the request's meaning or run only as an ExtProc service for OpenAI Responses state and protocol conversion. Agent Router remains responsible for Gateway API resources, provider credentials, rate limits, and traffic policy.
For large request bodies or streamed immediate responses from Semantic Router, also see Streamed ExtProc and immediate responses. That guide shows how to switch the ExtProc filter from BUFFERED to STREAMED request bodies and how streamed Chat Completions clients receive looper or fast_response immediate responses.
Responsibility split
The deployment consists of:
- Semantic Router evaluates the selected recipe and chooses the logical model or provider alias.
- Envoy Gateway provides the Kubernetes Gateway API data plane.
- Agent Router translates provider APIs and applies gateway-owned authentication, rate limiting, and traffic policy.
- Model providers serve the selected model. This guide uses a demo backend; it does not install production inference capacity.
Provider support changes independently of Semantic Router. Use the
Agent Router provider documentation
to choose an AIServiceBackend and credential policy, then bind the provider
names to the aliases used by your Semantic Router configuration.