跳到主要内容
版本:最新版(未发布)

ReMoM 推理编排

概览​

remom 在有界轮次中运行多个候选模型,并将它们的响应合成一个答案。

通过 entrypoints 将公开模型名映射到 recipe 来暴露 remom。公开名字没有内置分发逻辑:所选 recipe 评估自己的 signals 和 decisions,由 algorithm.type=remom 启动算法。若入口只应运行 remom 策略,请将这些策略放在独立 recipe 中。

灵感来源:PaCoRe — 扩展为支持模型混合。

主要优势​

  • 多轮并行推理,带宽调度可配置。
  • 从多个模型响应中智能合成。
  • 模型分配策略:weighted、equal、round_robin 或 first_only。
  • 用压缩策略管理跨轮 token 预算。
  • 可选的法定人数和轮次超时控制,避免等待提供商长尾。
  • 可自定义合成模板。

算法原理​

ReMoM 编排多轮并行模型调用:

  1. 第 1 轮:按 breadth_schedule[0] 在候选模型上发起并行调用。
  2. 压缩:可选地压缩中间响应(full 或 last_n_tokens)。
  3. 第 2 轮:按 breadth_schedule[1] 发起调用,并把压缩后的响应作为上下文。
  4. 最终合成:最后一次调用将所有中间结果合成连贯答案。

ReMoM 后端子请求是非流式的,以便每轮收到完整输出。仅在最终合成之后,才向流式客户端发出响应。

带宽调度控制每轮调用次数。例如 [32, 4] 表示第 1 轮 32 次调用、第 2 轮 4 次调用,然后是 1 次最终合成调用。

执行流程​

模型分配策略​

策略说明
weighted按 modelRefs 中的模型权重按比例分配调用
equal在所有候选模型间均分调用
round_robin按配置顺序轮询候选模型
first_only全部调用发给第一个声明的模型

解决什么问题?​

有些任务更适合并行探索再合成,而不是一次性选出单个模型。remom 为路由器提供一种受带宽控制的方式,探索多条推理路径并合并成一个最终答案。

何时使用​

  • 一条路由应在多轮中协调多个模型。
  • 需要可配置的带宽调度,而不是一步升级。
  • 中间响应应显式包含或排除。
  • 多轮推理再合成比单次调用效果更好。

已知限制​

  • token 消耗高:每轮都会生成多个响应。
  • 合成质量取决于合成模板和模型能力。
  • 轮次串行执行,延迟更长。
  • 需要仔细调优 breadth_schedule,以平衡质量与成本。

配置​

将公开名字映射到下方的默认路由。若需隔离策略,请将 routing 块放入命名 recipe,并修改入口的 recipe 引用。

entrypoints:
- model_names: [vllm-sr/remom]
recipe: default

配置一条 ReMoM 决策:

routing:
decisions:
- name: reasoning_panel
description: Combine a bounded reasoning panel into one answer.
priority: 100
output_contract: Preserve any explicit output format exactly.
output_contract_spec:
type: reference_selection
reference:
source: candidate_responses
id_format: index
extract:
mode: exact
sources: [content]
postprocess:
- type: dereference_selected_reference
modelRefs:
- model: qwen3-32b
- model: deepseek-worker
algorithm:
type: remom
remom:
breadth_schedule: [3, 2]
model_distribution: weighted

output_contract 是决策范围的提示词文本。把它用于应同时作用于 ReMoM、Fusion 和 Flow 的基准或应用格式要求,而不是把任务特定提示词硬编码进算法。output_contract_spec 是路由器可执行的类型化契约,用于后处理与归一化;把运行时行为放在这里,而不是编码成提示词启发式。提取默认精确匹配 content;仅当决策明确允许更宽的解析器时,才使用 extract.sources 或 extract.mode: json_object。

最小算法配置:

algorithm:
type: remom
remom:
breadth_schedule: [3, 2] # Parallel calls before final synthesis
model_distribution: weighted # weighted, equal, round_robin, or first_only
temperature: 0.7 # Temperature for model calls
include_reasoning: false # Include reasoning in synthesis
compaction_strategy: full # full or last_n_tokens
compaction_tokens: 1000 # Tokens to keep for last_n_tokens
synthesis_template: "" # Custom synthesis template (optional)
max_concurrent: 3 # Max concurrent calls per round
max_completion_tokens: 1024 # Completion limit for each subrequest
round_timeout_seconds: 120 # Optional round-level wait cap
min_successful_responses: 2 # Optional early-success quorum
shuffle_seed: 42 # Seed for response shuffling
include_intermediate_responses: false # Include intermediate responses in output
max_responses_per_round: null # Limit responses per round
on_error: skip # skip or fail

参数​

参数类型默认值说明
breadth_schedulelist[int]必填最终合成调用之前的并行调用次数(例如 [3, 2])
model_distributionstringweighted策略:weighted、equal、round_robin、first_only
temperaturefloat1.0模型调用的温度
include_reasoningboolfalse在合成提示中包含推理内容
compaction_strategystringfull策略:full 或 last_n_tokens
compaction_tokensint1000last_n_tokens 压缩时保留的 token 数
synthesis_templatestring—自定义合成提示模板
max_concurrentint—每轮最大并发模型调用数
max_completion_tokensint请求默认值应用到每个 ReMoM 子请求的最大补全 token 数
round_timeout_secondsint—在 on_error: skip 时,一轮最多等待的秒数,超时后使用部分响应
min_successful_responsesint—并行轮次在达到该成功响应数后即可返回
shuffle_seedint42响应打乱的随机种子
include_intermediate_responsesbooltrue在输出中包含中间响应
max_responses_per_roundint—每轮最多保留的响应数
on_errorstringskip失败时的行为:skip 或 fail

每轮会与分配到的模型共享请求派生文本和中间文本,合成模型会收到收集到的结果。投产前请限制带宽、补全 token、并发和超时。完整示例见: config/fragments/algorithm/looper/remom.yaml。