跳到主要内容

6 篇博文 含有标签「routing」

查看所有标签

Decision Models Need a Router, Too

· 阅读需 9 分钟
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
System One Auto intelligently selects from a decision-model pool. An Auto routing hub highlights a selected path among Decision 2.0, Cloudflare Clef, Perplexity Decider, Laya, Fastino GLiDE and TypeSafe Jev. Route language. Route decisions.

One API for a growing decision-model ecosystem. Provider logos illustrate the ecosystem; the experiment below measures Kai → Vega.

Save your largest decision model for the requests that need it.

vLLM Semantic Router already routes across open and closed LLMs. Decision models are becoming a model pool of their own. Open families such as Decision 2.0, Clef, Decider and Laya sit alongside hosted APIs such as Jev and GLiDE. They classify intent, score candidates and check conditions—but which model should handle each request?

Meet System One Auto. Connect self-hosted decision models or compatible hosted APIs behind vllm-sr/auto. Configure how to select a model, when to escalate through a cascade, and whether an LLM judge should review the final candidates. Your application keeps one native API; you choose the model pool and routing policy.

Beyond Prompt Routing: Model Selection, Conversation State, and KV Cache

· 阅读需 11 分钟
Anup Sharma
AI & Distributed System @ Nutanix
Stefan Wang
Senior Software Engineer at LinkedIn

An agent has completed three turns on model A. It has read a document, stored some facts, and started working toward an answer. The fourth turn needs stronger reasoning, so the router sends it to model B.

The request succeeds. But did the conversation improve?

Model B may receive the entire transcript and still inherit an incorrect answer. Its replica may have to compute the growing prompt again. If a tool call is in flight, changing models can also change how the result is interpreted.

We benchmarked these tradeoffs through vLLM Semantic Router, Envoy, and GPU-backed vLLM servers. In the 960-call holdout, one reactive switching policy reduced strong-model use by 31.8 percentage points and mean session latency by 550 ms compared with a gate-disabled control, but failed three final checks that the control passed. A separate native-tool experiment showed the benefit of choosing a capable model before a difficult boundary.

This is the problem beyond prompt classification: choosing the right model for the next turn without compromising the rest of the task. It also determines where semantic routing belongs in an existing inference stack.

Per-Call Model Selection: Measuring Cost and Quality Across a Multi-Agent Run

· 阅读需 15 分钟
Abhinav Mahajan
Software Engineer at Barclays

Every multi-agent framework asks one question before the first task arrives: which model does each agent use?

planner = Agent(role="planner", model="small-model")
summarizer = Agent(role="summarizer", model="small-model")
reviewer = Agent(role="reviewer", model="frontier-model")
writer = Agent(role="writer", model="small-model")

It looks like configuration. It is really a prediction, made once, about work nobody has seen yet.

A task never reaches a model as one request. It arrives as a stream of calls: list the changed files, summarize a module, search for a concurrency bug, write the final comment. Bind a model to an agent and every one of those calls gets the same answer, hard or trivial.

vLLM Semantic Router moves that choice onto the request. Every agent sends one model name, MoM, short for mixture of models. It is not a model but a virtual name the router resolves per call, from what the call contains, using signals and decisions written in YAML. Throughout this post, local means a small model on your own machine, free and slower, and frontier means a large hosted model, billed per token and fast.

To see what changes, we ran one four-agent crew four ways, all frontier, model per agent, model per call and all local, recording every call: which model served it, which decision chose it, what it cost, and how long the router took to decide.

Per-call routing sent 2 of 16 calls to the frontier model and cost $1.37 per 1,000 reviews. That is one fifth the cost of sending everything to the frontier model, and less than half the cost of a hand-tuned per-agent setup. It found 4 of 6 seeded bugs, matching the all-frontier arm and beating per-agent binding, which found 3. The router took 1.86 ms at p50 to make each choice.

Adding Cursor-Style Auto Model Selection to OpenCode with vLLM Semantic Router

· 阅读需 11 分钟
Anup Sharma
AI & Distributed System @ Nutanix
Aayush Saini
SDE, Data and AI @ Red Hat
Shivji Kumar Jha
Staff Engineer (Data & AI) @ Nutanix

OpenCode with vLLM Semantic Router: open provider interface, AgentGateway integration layer, and semantic routing hub

The Feature Everyone Wants and Almost Nobody Has​

Cursor's Auto mode is deceptively simple: the developer types, and the IDE chooses whether a prompt deserves a frontier model or something faster and cheaper. It is easy to stop noticing — until moving to an open tool where every request starts with a model dropdown.

Giving AgentGateway a Semantic Brain with vLLM Semantic Router

· 阅读需 10 分钟
Aayush Saini
SDE, Data and AI @ Red Hat
Anup Sharma
AI & Distributed System @ Nutanix

vLLM Agent Architecture Workflow: Custom Semantic Routing with AgentGateway and Semantic Router

Agent systems that span multiple models — a local endpoint for coding, a frontier cloud model for deep reasoning, and a fast general-purpose model for everyday tasks — all face the same routing question: how should each request be directed to the right backend?

Many deployments start with a lightweight Python proxy or keyword matcher in front of the gateway. That approach works at small scale, but misroutes grow quickly as traffic, languages, and task types diversify. This post shows how vLLM Semantic Router running as an Envoy ExtProc sidecar inside AgentGateway replaces that pattern with semantic, config-driven routing.

Agentic Routing on AMD ROCm

· 阅读需 14 分钟
Xunzhuo Liu
Intelligent Routing @vLLM | Open Source & AI @AMD
Haichen Zhang
Sr. AI Engineer @AMD
Andy Luo
Sr. Director @AMD

Most agent systems start with a simple idea: call model: auto and let the inference layer pick the right model. That is useful, but it is not enough for long-running agents.

A coding agent can begin with architecture work, call tools, receive short tool outputs, continue with "fix that", then ask a privacy-sensitive question in the same user session. The latest message may look simple, but the route cannot be chosen from the latest message alone. The router also has to know whether this is a safe moment to switch models.

This guide shows how to deploy that pattern on AMD ROCm with vLLM Semantic Router. You will start one ROCm vLLM backend, serve the agentic routing recipe, open the dashboard, validate the OpenAI-compatible API, and use Inferoa to experience route decisions and Router Learning behavior from an agent client.

Agent session routed through router memory to model paths
Agentic routing is not only choosing a model. It is choosing when to keep one.