Skip to main content

2 posts tagged with "agentic"

View All Tags

Beyond Prompt Routing: Model Selection, Conversation State, and KV Cache

· 11 min read
Anup Sharma
AI & Distributed System @ Nutanix
Stefan Wang
Senior Software Engineer at LinkedIn

An agent has completed three turns on model A. It has read a document, stored some facts, and started working toward an answer. The fourth turn needs stronger reasoning, so the router sends it to model B.

The request succeeds. But did the conversation improve?

Model B may receive the entire transcript and still inherit an incorrect answer. Its replica may have to compute the growing prompt again. If a tool call is in flight, changing models can also change how the result is interpreted.

We benchmarked these tradeoffs through vLLM Semantic Router, Envoy, and GPU-backed vLLM servers. In the 960-call holdout, one reactive switching policy reduced strong-model use by 31.8 percentage points and mean session latency by 550 ms compared with a gate-disabled control, but failed three final checks that the control passed. A separate native-tool experiment showed the benefit of choosing a capable model before a difficult boundary.

This is the problem beyond prompt classification: choosing the right model for the next turn without compromising the rest of the task. It also determines where semantic routing belongs in an existing inference stack.

Agentic Routing on AMD ROCm

· 14 min read
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
Haichen Zhang
Sr. AI Engineer @AMD
Andy Luo
Sr. Director @AMD

Most agent systems start with a simple idea: call model: auto and let the inference layer pick the right model. That is useful, but it is not enough for long-running agents.

A coding agent can begin with architecture work, call tools, receive short tool outputs, continue with "fix that", then ask a privacy-sensitive question in the same user session. The latest message may look simple, but the route cannot be chosen from the latest message alone. The router also has to know whether this is a safe moment to switch models.

This guide shows how to deploy that pattern on AMD ROCm with vLLM Semantic Router. You will start one ROCm vLLM backend, serve the agentic routing recipe, open the dashboard, validate the OpenAI-compatible API, and use Inferoa to experience route decisions and Router Learning behavior from an agent client.

Agent session routed through router memory to model paths
Agentic routing is not only choosing a model. It is choosing when to keep one.