Beyond Prompt Routing: Model Selection, Conversation State, and KV Cache
An agent has completed three turns on model A. It has read a document, stored some facts, and started working toward an answer. The fourth turn needs stronger reasoning, so the router sends it to model B.
The request succeeds. But did the conversation improve?
Model B may receive the entire transcript and still inherit an incorrect answer. Its replica may have to compute the growing prompt again. If a tool call is in flight, changing models can also change how the result is interpreted.
We benchmarked these tradeoffs through vLLM Semantic Router, Envoy, and GPU-backed vLLM servers. In the 960-call holdout, one reactive switching policy reduced strong-model use by 31.8 percentage points and mean session latency by 550 ms—but failed three final checks that the control passed. A separate native-tool experiment showed the benefit of choosing a capable model before a difficult boundary.
This is the problem beyond prompt classification: choosing the right model for the next turn without compromising the rest of the task. It also determines where semantic routing belongs in an existing inference stack.

