Skip to main content
Blog

vLLM Semantic Router v0.4 Hermes: Many Models, One Improving System

An application should not need a new integration every time the best model changes. Hermes gives it one stable model name, while vLLM Semantic Router chooses from a qualified pool, delivers the answer, and measures whether that choice helped.

Themis (v0.3) made routing stateful and inspectable. Hermes turns that foundation into a model system that can improve as models and workloads change.

982 commits · 130 contributors · 107 first-time contributors since Themis. This is a milestone built by a growing community across models, runtime, evaluation, and the experience of putting it all to work.

vLLM Semantic Router v0.4 Hermes: one model name, a system built to improve

Hermes: many models, one improving system.

Hermes in three moves​

  • Open the model foundation. Vela helps understand a request; Decision explores how software can make the next choice from evidence and criteria.
  • Give applications one name. A virtual model maps to an isolated Mixture-of-Models (MoM) recipe. Behind it, the Router can select one backend, escalate when needed, or run bounded model collaboration.
  • Measure the whole route. The request must survive protocol translation, long sessions, and backend failures. Then sr-bench can compare that routed experience with individual models on the same work.

Two open model launches​

Vela 1.0: understand the work before choosing a model​

A router needs to recognize what a request is asking for. Vela launched with 14 open models; Shield has since joined (a contribution from KR-Labs), bringing the family to 15. Each has a specific task in understanding a request, finding context, or checking what comes back.

Every request finds its way. Meet Vela in 87 seconds.

ModelTaskWhat it does
Vela-1.0-Encoder-307M-DomainClassificationIdentifies the request's subject for specialist routing.
Vela-1.0-Encoder-307M-FeedbackClassificationRecognizes corrections, clarification, and other feedback.
Vela-1.0-Encoder-307M-ModalityClassificationDetects whether the request calls for text, an image, or both.
Vela-1.0-Encoder-307M-FactCheckClassificationFlags requests that need external factual knowledge or retrieval.
Vela-1.0-Encoder-307M-GuardClassificationSpots prompt injection and jailbreak attempts.
Vela-1.0-Encoder-307M-SafetyClassificationDetects unsafe content.
Vela-1.0-Encoder-307M-HazardMulti-label ClassificationScores distinct content-risk categories.
Vela-1.0-Encoder-307M-ShieldMulti-task ClassificationScores request and response safety signals in one encoder.
Vela-1.0-Encoder-307M-PIIToken ClassificationFinds sensitive spans for privacy-aware handling.
Vela-1.0-Encoder-307M-HaluToken ClassificationMarks answer spans unsupported by supplied evidence.
Vela-1.0-Encoder-307M-EmbeddingSemantic SimilarityProduces multilingual vectors for retrieval and clustering.
Vela-1.0-Encoder-307M-RerankerRerankingPuts the most relevant retrieved passages first.
Vela-1.0-Omni-NanoMultimodal EmbeddingMakes compact embeddings for text, images, and audio.
Vela-1.0-Omni-MiniMultimodal EmbeddingAdds a larger shared space and instructed text search.
Vela-1.0-Encoder-307MMasked Language ModelingProvides a multilingual base for specialized fine-tuning.

Hermes brings Vela into the running Router too: selected text models serve routing signals, Halu appears in a reference configuration, and Omni has an optional native path for multimodal embeddings.

Explore the 15 models → · Try Vela Studio → · Read the launch story →

Decision 1.0: make the choice itself open​

A system also needs to decide what happens next. Decision 1.0 releases six open-weight models that score the choices an application provides. They can choose an option, judge a condition, or score an ordered rubric, returning distributions that software can use and inspect.

Meet the six models and the decisions they make possible.

ModelWhere it fitsInput budget
Decision-1.0-Kai-0.6BCompact general-purpose routing, conditions, and action selection1,024 tokens
Decision-1.0-Lex-0.6BOperational workflows such as service, invoices, incidents, and agent traces1,024 tokens
Decision-1.0-Eos-0.8BThe smallest hybrid decoder for decisions with longer evidence16,384 tokens
Decision-1.0-Sol-2BMore capacity for decisions over extended context16,384 tokens
Decision-1.0-Nox-4BCombining conditions, applying rules, and choosing actions16,384 tokens
Decision-1.0-Lux-9BThe largest model and strongest overall result in the released suite16,384 tokens

That opens a broader direction for routing, policy, and agent actions. The models and their interface are available now; native Decision integration in vLLM Semantic Router is the next step.

Nox-4B chooses actions step by step in a robot-arm simulation; the on-screen number is model latency for that step.

On the published 54-task evaluation, Lux-9B leads the displayed open-model references overall. The matrix breaks that result into decisions, composition, reading, inference, and transfer, so you can see where each model is strongest.

Decision model ranking on the published selected task suite: Lux-9B scores 76.94 overall, Nox-4B scores 73.09, and the hosted Jev reference scores 81.05.

The rank view: how Decision models compare with the published references overall.

Capability matrix for Decision models and reference models across decisions, composition, reading, inference, transfer, and overall score.

The matrix view: choose a model for the kind of decisions your application makes.

Explore the six models → · Try Decision Studio → · Read the launch story →

The Tetris demo puts response time in motion, showing Kai-0.6B alongside Lux-9B and Jev Upstream as the boards change.

A quick look at decision speed in the Tetris demo.

A stronger route, end to end​

Themis made the Signal → Projection → Decision → Algorithm → Model path inspectable. Hermes makes the evidence, choices, and model execution along that path more dependable. Route plugins add behavior around the chosen path.

A request moves from signal and projection through decision and algorithm to a selected model, with a route plugin branching from the decision

The core routing path, with route plugins around the chosen decision.

Signals: see more of the request and the answer​

Themis could already detect domains, conversation shape, context demand, and safety risks. Hermes gives those signals more useful inputs and clearer failure states:

  • Recognize what arrived. Input modality detects the presence of text, image, or audio parts without an inference call. Bounded metadata rules can use an application's routing hint, while authorization still comes from trusted identity.
  • Bring model evidence into policy. The reusable classifier signal accepts local or external classifiers and their label scores; complexity can use an external scorer. Vela Safety and Hazard distinguish broad content risk from specific categories instead of treating every safety concern as a prompt attack.
  • Check long content deliberately. Guard and PII can scan admitted text in bounded, overlapping windows rather than relying on its first slice; hallucination detection can scan a long answer too. A failed window does not become a clean result, and the configured document budget still limits what can be scanned.
  • Observe what came back. Response jailbreak and hallucination are response-stage signals recorded in Replay. Hallucination checks an answer against supplied grounding context; without that evidence it cannot verify the answer. A route can attach a plugin to act on the result. For a stream, the observation is made after completion, so it cannot retract bytes already sent.

The failure path matters as much as a positive detection. A classifier error remains unknown. A weighted projection that depends on that signal withholds its score instead of turning the missing value into zero and producing a misleading low-risk band.

Decisions: when several routes match, why does one win?​

The Router's Decision stage is the policy engine running today; the Decision 1.0 model family above is a separate launch. Themis already had rules, tiers, and priority and confidence strategies. Hermes makes their difficult cases explicit:

StepWhat the Router does
1. MatchEvaluate AND, OR, and NOT rules. If the final result remains unknown, on_unknown can skip the route, match it, or stop the request.
2. Set the boundaryA lower tier takes precedence. Within the selected tier, conditional matches rank ahead of an unconditional catch-all.
3. Rank the matchesrouting.strategy uses priority or confidence within that tier. Confidence applies only when every non-catch-all match has one comparable measured score of the same kind; otherwise the pool falls back to priority.
4. Explain the resultPreview and Replay show the strategy, tier, whether the evidence was comparable, and the reason the winner beat the next route.

A keyword hit is a policy gate, not 100% confidence. A domain probability is not interchangeable with an embedding similarity. Keeping those numbers separate makes a close decision easier to trust and to improve.

For prompt attacks, a safety route can override a model pinned by the client and send the request to an eligible guarded path.

Algorithms: choose one model or coordinate several​

After a decision matches, its algorithm sees only that decision's candidates. Hermes now filters models by context, protocol, and declared task capabilities before scoring; a route requiring a minimum pool fails rather than quietly running a smaller panel or cascade.

PathHermes highlight
Select oneHybrid can be tuned per decision. The new, experimental prompt selector lets a helper model choose from the eligible candidate names, with a bounded fallback if that call fails.
Work across modelsNew, experimental Fusion runs a panel and judge with a usable-response quorum and configurable fallback. New, experimental Workflows assign bounded roles and can resume tool-driven work.
Refine existing pathsConfidence records per-attempt results, timing, and token usage. ReMoM adds round timeouts and an optional successful-response quorum.

These are different choices: a selector returns one qualified model; a Looper can run a bounded cascade or fan-out of model calls. The recipe determines when that extra work is justified.

Plugins: shape the chosen route​

Themis already had Memory, RAG, Replay, and caching. Hermes adds route-local behaviors and makes existing ones safer:

PluginWhat changed
Context CompressionNew opt-in handling for long history and tool output, with a defined transformation order. It protects instructions and tool-call pairs, and preserves the current user turn unless truncation is explicitly enabled.
Shadow DispatchA new bounded, sampled copy of a single-model request goes to a candidate without affecting the live answer. Replay records completed or failed shadow calls; answer quality still needs separate evaluation.
Response CacheExact and semantic reuse now checks request compatibility, including history, tools, format, model, and recipe. Negation checks and a failed optional verifier turn a questionable hit into a miss.
Router MemoryResponse-side extraction can persist useful facts asynchronously and report whether the write succeeded, failed, or was dropped.
Tool SelectionExisting semantic tool filtering now batches and caches tool embeddings instead of repeatedly encoding identical tool schemas.

Together, these changes give an operator a practical loop: route a real request, inspect why it took that path, observe a candidate through Shadow Dispatch, and use sr-bench to test whether changing the recipe improves quality, cost, and latency.

Seven Workgroups, one routed experience​

The seven Workgroups advance this system from model research to the request path and the tools around it:

WorkgroupThemis foundationHermes highlight
Developer Experience & EcosystemOne configuration across CLI and DashboardA guided path from setup to an explainable live request
Enterprise & EnvironmentInspectable operationsStronger management controls and clearer deployment support
Router Models & Inference RuntimeModel-backed signalsShared model resources and bounded inference paths
MoM & RoutingDecisions select backendsStable virtual names, five MoM recipes, and bounded Micro-Agent collaboration
Agentic & ContextSession-aware routingSafer context changes, memory outcomes, and optional switch evidence
Data Plane & NetworkingMultiple protocol pathsShared codecs, image routing, and bounded recovery
Evaluation & QualityReplay explains a choiceFrozen-task MoM comparisons and clearer evidence boundaries

Developer Experience & Ecosystem: get from setup to a route you can explain​

The first useful result is a request whose path makes sense.

A developer moves from setup through a published recipe to an explainable routed request

From setup to a route you can explain.

  • Start with guidance. The agent installation flow discovers the running Router, checks a configuration, previews its decision, and probes a real backend.
  • See the whole product. Model Hub helps discover providers and model cards; Dashboard connects models, recipes, entrypoints, and Playground.
  • Use the local runtime that fits. The installer now offers Podman alongside its existing local path.

Preview tells you which route the policy would choose. Probe verifies that a backend can actually deliver an answer. That distinction makes the first deployment easier to trust.

The OpenCode Auto Mode example shows why the stable name matters: a coding assistant can keep one model setting while the Router chooses from local and cloud backends.

Enterprise & Environment: make change safer to operate​

A model system needs a clear boundary around who can change it and where it runs.

Management controls and deployment choices around a routing recipe

Management controls and deployment choices around a recipe.

  • Protect management paths. Route-level permission checks, identity handling, and sensitive-field redaction strengthen the management API; Dashboard adds CSRF and origin checks.
  • Isolate local stacks. Managed data services receive per-stack credentials.
  • Show the supported path. The deployment matrix makes maintained, supported, and experimental surfaces visible.

Production deployments still need their own access and backend choices configured.

Router Models & Inference Runtime: put model intelligence on the request path​

Open checkpoints become useful when the Router can run and share them predictably.

Router Models turn request content into typed evidence before backend selection

Router Models turn request content into typed evidence.

  • Use Vela as evidence. The Router includes Vela task models, Halu in a reference configuration, and an optional Omni path for multimodal embeddings.
  • Check grounded answers. An optional LettuceDetect v2 endpoint can mark unsupported spans in answers that have grounding evidence.
  • Share what is expensive. A resource inventory lets recipes reuse physical Router Models while keeping their consumers visible.
  • Bound the work. Configurable inference concurrency and queues help keep small-model work from overwhelming the request path. OpenVINO, Apple Metal, and RISC-V CPU paths also gained targeted support.

MoM & Routing: make a model pool feel like one product​

An application names the service it wants; a recipe decides how to provide it.

A virtual model recipe selects or coordinates qualified backends for its objective

A virtual model recipe selects or coordinates backends.

A public entrypoint resolves to an isolated recipe. Operators connect their own eligible backends, then choose one of five MoM V1 objectives:

RecipeObjective
BlendBalance everyday quality, latency, and cost
LiteKeep routine service economical, with targeted escalation
FlashFavor responsive conversation and streaming
UltraKeep stronger answers and review paths available
VaultApply a privacy-oriented policy with operator-approved backends

Picture a support request behind one virtual model. A routine turn can go straight to a responsive backend. A recipe for uncertain answers can escalate through Confidence. A harder task can ask several models to work together and still return one response to the client.

That collaboration lives in the Router's bounded Micro-Agent runtime. Confidence escalates, Ratings returns one identifiable answer per successful candidate, ReMoM explores multiple reasoning rounds, Fusion judges a panel, and Workflows assigns limited roles. Each uses the limits that fit its shape, including thresholds, concurrency caps, timeouts, and failure rules.

Hermes makes Fusion's analysis mode, minimum valid responses, and fallback behavior recipe-controlled. Decision ranking and Preview explain why a path won when several could match. Our MoM vision sets out where this model identity can go next.

The stable name belongs to the application. The changing pool, tradeoffs, and evidence remain visible to the operator.

Agentic & Context: preserve continuity across a long session​

A good route for one turn can be a bad switch in the middle of an agent's work.

An agent session retains important context while memory and switching remain bounded

Session continuity with bounded memory and switching.

Themis introduced Session-Aware Agentic Routing to keep tool loops and provider continuations safe. Think of a coding assistant returning from a tool call with a growing conversation. Hermes gives context transformations a defined order, protecting important instructions and tool exchanges while trimming less useful material.

Response-side Router Memory can extract useful facts and report whether an asynchronous write succeeded, failed, or was dropped. An optional recent-outcome gate considers a short window of session results and cooldowns before switching models; it is off by default. The operator can see what happened to memory and why the session stayed or moved.

Data Plane & Networking: deliver the model choice​

Selection matters only if the chosen backend can answer in the client's language.

A selected route crosses protocol translation, backend dispatch, and streaming delivery

From selection through protocol translation to delivery.

  • Translate consistently. A shared codec carries Chat Completions, Responses, and Anthropic Messages through buffered and streaming paths.
  • Route by capability. Image-generation requests can reach an image-capable backend.
  • Recover within limits. Opt-in cross-model fallback can try another eligible backend after an upstream error, but stops once a stream is committed or a tool side effect may have happened.

The AgentGateway integration story shows how a gateway can consume the Router's decision through Envoy ExtProc while the application keeps its usual API.

Evaluation & Quality: prove a better route on the same work​

A MoM earns its place by doing better on representative tasks, with cost and latency in view.

Frozen tasks compare a routed MoM with individual models on quality, cost, and latency

Compare a MoM with individual models on the same frozen tasks.

  • Compare fairly. sr-bench 1.0 runs a MoM and individual models on the same frozen tasks through the endpoints users call.
  • Protect the holdout. Disjoint tasks and history exclusions keep earlier evaluation work out of a reserved holdout; Replay can export reproducible shadow data for later analysis.
  • Read each kind of evidence. The Open Intelligence Index and Arena collect published model results. sr-bench measures a particular deployed MoM.

That gives a team a practical loop: change the recipe, run the same work, and keep the quality, cost, and latency together.

Workgroups are part of the release​

The technology grew with a new way to build it. Seven public Workgroups now have charters and clear Epic ownership. Weekly lists show work that contributors can pick up, and each technical direction has people responsible for turning ideas into reviewable changes.

Seven Workgroups give contributors a clear path to build vLLM Semantic Router together

Seven Workgroups give contributors a clear path in.

A contributor can:

  1. Find a direction in a Workgroup charter and choose an available issue.
  2. Claim and ship a focused change with its Workgroup and reviewers.
  3. Grow into membership through substantive contributions and a reviewed roster update.

The Workgroups make the project easier to join and the product easier to improve across boundaries.

Start with one model name​

  1. Start the Router with the installation guide and connect your own model endpoints.
  2. Publish a recipe such as Flash or Blend under a virtual model name. Use Preview to see which decision and backend it would select.
  3. Send a real request, then measure it. Probe delivery through the backend and use sr-bench to compare your routed policy with individual models.

Looking ahead: v0.5 Ariadne​

In Greek mythology, Ariadne's thread marks a way through the labyrinth. For v0.5, the thread is the route a real task takes across models, tools, and long sessions. We want to keep that route coherent and measure whether it led to a better result.

Ariadne's thread connects models, agent sessions, and evaluation

Ariadne: follow the route across models, sessions, and evaluation.

The seven Workgroups are taking that question into the next release from different directions:

WorkgroupDirection for v0.5 Ariadne
Developer Experience & EcosystemHelp a developer move from a first routed request to a recipe change they can inspect and evaluate.
Enterprise & EnvironmentMake model and policy changes safer to operate through versioned activation and rollback.
Router Models & Inference RuntimeImprove Vela and Decision model quality and reduce the cost of Decision inference on the request path.
MoM & RoutingMove toward portable, versioned Mixture-of-Models that carry their pool, recipe, and evaluation evidence together, so the next iteration can be compared and shared.
Agentic & ContextOptimize context for long-running agent work while retaining the instructions, tool state, and session history that the task needs.
Data Plane & NetworkingBring /v1/messages, /v1/chat/completions, and /v1/responses to GA, with complete end-to-end coverage of their supported streaming, tool, error, and fallback paths.
Evaluation & QualityGrow SR Bench 2.0 to evaluate complete agent workflows, multi-turn sessions, and robustness under changing inputs and backend failures, measuring task success and continuity alongside quality, latency, and cost.

The aim is to follow a request all the way through and know whether each change made the result better.

Growing together​

A release this large belongs to the people who built it. We are glad to welcome 11 Committers and 3 Maintainers to vLLM Semantic Router.

Growing Together: Committers Guan-Ming Chiu, Abhinav Mahajan, Park Soobin, Anup Sharma, Alex Jia, Ádám Kovács, Theo Hsiung, Wilson Wu, Yincheng Ren, Stefan Wang, and Binbin Zhang; Maintainers Jinyu Chou, Kun-Tai Wu, and Aayush Saini

The next wave is already here: 61 builders across 83 Workgroup applications in this snapshot. Each is a person choosing where to contribute. Find your Workgroup.

September 26, 2026 Workgroup applications: 83 application records from 61 builders across all seven groups

Workgroup applications on September 26, 2026. Open the image to see the people behind them.

Acknowledgments​

Our new colleagues bring perspectives from National Taiwan University and New York University, and from Quid, Barclays, Rebellions, Nutanix, KR Labs, Delta Electronics, Red Hat, DaoCloud, Meta, and the PyTorch community. We thank the people in these schools, teams, and communities who make open work possible.

As in Themis, we also thank collaborators at MBZUAI, McGill University, and Mila, and the wider vLLM, Hugging Face, and open-source communities building model serving and AI infrastructure with us.

Thank you for the code, reviews, ideas, and patience that made Hermes possible. Find your Workgroup and help build what comes next.

Contributors to vLLM Semantic Router throughout the project's history

See the live contributor wall →