Skip to main content

7 posts tagged with "vllm"

View All Tags

Beyond Prompt Routing: Model Selection, Conversation State, and KV Cache

· 11 min read
Anup Sharma
AI & Distributed System @ Nutanix
Stefan Wang
Senior Software Engineer at LinkedIn

An agent has completed three turns on model A. It has read a document, stored some facts, and started working toward an answer. The fourth turn needs stronger reasoning, so the router sends it to model B.

The request succeeds. But did the conversation improve?

Model B may receive the entire transcript and still inherit an incorrect answer. Its replica may have to compute the growing prompt again. If a tool call is in flight, changing models can also change how the result is interpreted.

We benchmarked these tradeoffs through vLLM Semantic Router, Envoy, and GPU-backed vLLM servers. In the 960-call holdout, one reactive switching policy reduced strong-model use by 31.8 percentage points and mean session latency by 550 ms compared with a gate-disabled control, but failed three final checks that the control passed. A separate native-tool experiment showed the benefit of choosing a capable model before a difficult boundary.

This is the problem beyond prompt classification: choosing the right model for the next turn without compromising the rest of the task. It also determines where semantic routing belongs in an existing inference stack.

Per-Call Model Selection: Measuring Cost and Quality Across a Multi-Agent Run

· 15 min read
Abhinav Mahajan
Software Engineer at Barclays

Every multi-agent framework asks one question before the first task arrives: which model does each agent use?

planner = Agent(role="planner", model="small-model")
summarizer = Agent(role="summarizer", model="small-model")
reviewer = Agent(role="reviewer", model="frontier-model")
writer = Agent(role="writer", model="small-model")

It looks like configuration. It is really a prediction, made once, about work nobody has seen yet.

A task never reaches a model as one request. It arrives as a stream of calls: list the changed files, summarize a module, search for a concurrency bug, write the final comment. Bind a model to an agent and every one of those calls gets the same answer, hard or trivial.

vLLM Semantic Router moves that choice onto the request. Every agent sends one model name, MoM, short for mixture of models. It is not a model but a virtual name the router resolves per call, from what the call contains, using signals and decisions written in YAML. Throughout this post, local means a small model on your own machine, free and slower, and frontier means a large hosted model, billed per token and fast.

To see what changes, we ran one four-agent crew four ways, all frontier, model per agent, model per call and all local, recording every call: which model served it, which decision chose it, what it cost, and how long the router took to decide.

Per-call routing sent 2 of 16 calls to the frontier model and cost $1.37 per 1,000 reviews. That is one fifth the cost of sending everything to the frontier model, and less than half the cost of a hand-tuned per-agent setup. It found 4 of 6 seeded bugs, matching the all-frontier arm and beating per-agent binding, which found 3. The router took 1.86 ms at p50 to make each choice.

LettuceDetect v2 in Semantic Router: Generative Hallucination Detection as a vLLM Endpoint

· 9 min read
Ádám Kovács
Co-founder @ KR Labs · LettuceDetect
Bowei He
Postdoctoral Researcher @ MBZUAI · McGill
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
Huamin Chen
AI @Microsoft

Semantic Router can now verify grounded responses with a generative span detector served by vLLM. The new endpoint detector backend runs LettuceDetect v2 against every fact-checkable answer: unsupported spans are located to the character, typed against a hallucination taxonomy, and explained — in one call, before the response reaches the user.

The models come out of a joint paper between KR Labs and the Semantic Router team, Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents (arXiv:2607.00895). This post walks through the paper — the benchmark, the taxonomy, the models, and what they score — and then through the integration that puts the detector into the serving stack.

LettuceDetect v2 flagging contract hallucinations through Semantic Router

Adding Cursor-Style Auto Model Selection to OpenCode with vLLM Semantic Router

· 11 min read
Anup Sharma
AI & Distributed System @ Nutanix
Aayush Saini
SDE, Data and AI @ Red Hat
Shivji Kumar Jha
Staff Engineer (Data & AI) @ Nutanix

OpenCode with vLLM Semantic Router: open provider interface, AgentGateway integration layer, and semantic routing hub

The Feature Everyone Wants and Almost Nobody Has​

Cursor's Auto mode is deceptively simple: the developer types, and the IDE chooses whether a prompt deserves a frontier model or something faster and cheaper. It is easy to stop noticing — until moving to an open tool where every request starts with a model dropdown.

Giving AgentGateway a Semantic Brain with vLLM Semantic Router

· 10 min read
Aayush Saini
SDE, Data and AI @ Red Hat
Anup Sharma
AI & Distributed System @ Nutanix

vLLM Agent Architecture Workflow: Custom Semantic Routing with AgentGateway and Semantic Router

Agent systems that span multiple models — a local endpoint for coding, a frontier cloud model for deep reasoning, and a fast general-purpose model for everyday tasks — all face the same routing question: how should each request be directed to the right backend?

Many deployments start with a lightweight Python proxy or keyword matcher in front of the gateway. That approach works at small scale, but misroutes grow quickly as traffic, languages, and task types diversify. This post shows how vLLM Semantic Router running as an Envoy ExtProc sidecar inside AgentGateway replaces that pattern with semantic, config-driven routing.

Agentic Routing on AMD ROCm

· 14 min read
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
Haichen Zhang
Sr. AI Engineer @AMD
Andy Luo
Sr. Director @AMD

Most agent systems start with a simple idea: call model: auto and let the inference layer pick the right model. That is useful, but it is not enough for long-running agents.

A coding agent can begin with architecture work, call tools, receive short tool outputs, continue with "fix that", then ask a privacy-sensitive question in the same user session. The latest message may look simple, but the route cannot be chosen from the latest message alone. The router also has to know whether this is a safe moment to switch models.

This guide shows how to deploy that pattern on AMD ROCm with vLLM Semantic Router. You will start one ROCm vLLM backend, serve the agentic routing recipe, open the dashboard, validate the OpenAI-compatible API, and use Inferoa to experience route decisions and Router Learning behavior from an agent client.

Agent session routed through router memory to model paths
Agentic routing is not only choosing a model. It is choosing when to keep one.

Deploying vLLM Semantic Router on AMD Developer Cloud

· 12 min read
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
Haichen Zhang
Sr. AI Engineer @AMD
Andy Luo
Sr. Director @AMD

AMD Developer Cloud and vLLM Semantic Router overview

Running vLLM Semantic Router on AMD Developer Cloud is not just about bringing up one more inference endpoint. It is about turning it into a routed multi-tier system that can classify requests, choose a semantic lane, and make replay and Insights immediately useful.

This post walks through the practical path: start the ROCm backend on an AMD Developer Cloud instance, install vLLM-SR, import the reference profile, and validate the deployment end to end.