Skip to main content

One post tagged with "observability"

View All Tags

Per-Call Model Selection: Measuring Cost and Quality Across a Multi-Agent Run

· 15 min read
Abhinav Mahajan
Software Engineer at Barclays

Every multi-agent framework asks one question before the first task arrives: which model does each agent use?

planner = Agent(role="planner", model="small-model")
summarizer = Agent(role="summarizer", model="small-model")
reviewer = Agent(role="reviewer", model="frontier-model")
writer = Agent(role="writer", model="small-model")

It looks like configuration. It is really a prediction, made once, about work nobody has seen yet.

A task never reaches a model as one request. It arrives as a stream of calls: list the changed files, summarize a module, search for a concurrency bug, write the final comment. Bind a model to an agent and every one of those calls gets the same answer, hard or trivial.

vLLM Semantic Router moves that choice onto the request. Every agent sends one model name, MoM, short for mixture of models. It is not a model but a virtual name the router resolves per call, from what the call contains, using signals and decisions written in YAML. Throughout this post, local means a small model on your own machine, free and slower, and frontier means a large hosted model, billed per token and fast.

To see what changes, we ran one four-agent crew four ways, all frontier, model per agent, model per call and all local, recording every call: which model served it, which decision chose it, what it cost, and how long the router took to decide.

Per-call routing sent 2 of 16 calls to the frontier model and cost $1.37 per 1,000 reviews. That is one fifth the cost of sending everything to the frontier model, and less than half the cost of a hand-tuned per-agent setup. It found 4 of 6 seeded bugs, matching the all-frontier arm and beating per-agent binding, which found 3. The router took 1.86 ms at p50 to make each choice.