Skip to main content

3 posts tagged with "agents"

View All Tags

Per-Call Model Selection: Measuring Cost and Quality Across a Multi-Agent Run

· 15 min read
Abhinav Mahajan
Software Engineer at Barclays

Every multi-agent framework asks one question before the first task arrives: which model does each agent use?

planner = Agent(role="planner", model="small-model")
summarizer = Agent(role="summarizer", model="small-model")
reviewer = Agent(role="reviewer", model="frontier-model")
writer = Agent(role="writer", model="small-model")

It looks like configuration. It is really a prediction, made once, about work nobody has seen yet.

A task never reaches a model as one request. It arrives as a stream of calls: list the changed files, summarize a module, search for a concurrency bug, write the final comment. Bind a model to an agent and every one of those calls gets the same answer, hard or trivial.

vLLM Semantic Router moves that choice onto the request. Every agent sends one model name, MoM, short for mixture of models. It is not a model but a virtual name the router resolves per call, from what the call contains, using signals and decisions written in YAML. Throughout this post, local means a small model on your own machine, free and slower, and frontier means a large hosted model, billed per token and fast.

To see what changes, we ran one four-agent crew four ways, all frontier, model per agent, model per call and all local, recording every call: which model served it, which decision chose it, what it cost, and how long the router took to decide.

Per-call routing sent 2 of 16 calls to the frontier model and cost $1.37 per 1,000 reviews. That is one fifth the cost of sending everything to the frontier model, and less than half the cost of a hand-tuned per-agent setup. It found 4 of 6 seeded bugs, matching the all-frontier arm and beating per-agent binding, which found 3. The router took 1.86 ms at p50 to make each choice.

Giving AgentGateway a Semantic Brain with vLLM Semantic Router

· 10 min read
Aayush Saini
SDE, Data and AI @ Red Hat
Anup Sharma
AI & Distributed System @ Nutanix

vLLM Agent Architecture Workflow: Custom Semantic Routing with AgentGateway and Semantic Router

Agent systems that span multiple models — a local endpoint for coding, a frontier cloud model for deep reasoning, and a fast general-purpose model for everyday tasks — all face the same routing question: how should each request be directed to the right backend?

Many deployments start with a lightweight Python proxy or keyword matcher in front of the gateway. That approach works at small scale, but misroutes grow quickly as traffic, languages, and task types diversify. This post shows how vLLM Semantic Router running as an Envoy ExtProc sidecar inside AgentGateway replaces that pattern with semantic, config-driven routing.

Semantic Tool Selection: Building Smarter AI Agents with Context-Aware Routing

· 11 min read
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
Huamin Chen
AI @Microsoft

Anthropic recently published an insightful blog post on code execution with MCP, highlighting a critical challenge in modern AI systems: as agents connect to more tools, loading all tool definitions upfront becomes increasingly inefficient. Their solution—using code execution to load tools on-demand—demonstrates how established software engineering patterns can dramatically improve agent efficiency.

This resonates deeply with our experience building the vLLM Semantic Router. We've observed the same problem from a different angle: when AI agents have access to hundreds or thousands of tools, how do they know which tools are relevant for a given task?

Our solution: semantic tool selection—using semantic similarity to automatically select the most relevant tools for each user query before the request even reaches the LLM.

tools