Skip to main content
Version: Latest (unreleased)

ReMoM (Reasoning for Mixture of Models)

Overview​

remom runs several candidate models across bounded rounds and synthesizes their responses into one answer.

Expose remom through an ordinary entrypoints mapping to a recipe. The public name has no built-in dispatch behavior: the selected recipe evaluates its signals and decisions, and algorithm.type=remom activates the algorithm. Use a dedicated recipe when this entrypoint should run only remom policies.

Inspired by: PaCoRe — extended to support mixture of models.

Key Advantages​

  • Multi-round parallel reasoning with configurable breadth schedule.
  • Intelligent synthesis from multiple model responses.
  • Model distribution strategies: weighted, equal, round_robin, or first_only.
  • Compaction strategy to manage token budgets across rounds.
  • Optional quorum and round timeout controls to avoid waiting on provider long tails.
  • Customizable synthesis templates.

Algorithm Principle​

ReMoM orchestrates multiple rounds of parallel model calls:

  1. Round 1: Launch breadth_schedule[0] parallel calls across candidate models.
  2. Compaction: Optionally compact intermediate responses (full or last_n_tokens).
  3. Round 2: Launch breadth_schedule[1] calls, feeding compacted responses as context.
  4. Final Synthesis: One final call synthesizes all intermediate results into a coherent answer.

ReMoM backend subrequests are non-streaming so each round receives complete outputs. A streaming client response is emitted only after final synthesis.

The breadth schedule controls how many calls happen per round. For example [32, 4] means 32 calls in round 1, 4 in round 2, then 1 final synthesis call.

Execution Flow​

Model Distribution Strategies​

StrategyDescription
weightedDistribute calls proportional to model weights in modelRefs
equalDistribute calls equally across all candidate models
round_robinCycle through candidate models in configured order
first_onlyAll calls go to the first declared model

What Problem Does It Solve?​

Some tasks benefit from parallel exploration and later synthesis rather than one-shot selection of a single model. remom gives the router a breadth-controlled way to explore multiple reasoning paths and merge them into one final answer.

When to Use​

  • One route should coordinate multiple models over several passes.
  • You need a configurable breadth schedule instead of one-step escalation.
  • Intermediate responses should be included or excluded explicitly.
  • Multi-round reasoning with synthesis produces better answers than single-shot.

Known Limitations​

  • High token consumption: each round generates multiple responses.
  • Synthesis quality depends on the synthesis template and model capability.
  • Longer latency due to sequential round execution.
  • Requires careful tuning of breadth_schedule to balance quality vs. cost.

Configuration​

Map the public name to the recipe shown below. Move the routing block into a named recipe to isolate it from default routing.

entrypoints:
- model_names: [vllm-sr/remom]
recipe: default

Configure a ReMoM decision:

routing:
decisions:
- name: reasoning_panel
description: Combine a bounded reasoning panel into one answer.
priority: 100
output_contract: Preserve any explicit output format exactly.
output_contract_spec:
type: reference_selection
reference:
source: candidate_responses
id_format: index
extract:
mode: exact
sources: [content]
postprocess:
- type: dereference_selected_reference
modelRefs:
- model: qwen3-32b
- model: deepseek-worker
algorithm:
type: remom
remom:
breadth_schedule: [3, 2]
model_distribution: weighted

output_contract is decision-scoped prompt text. Use it for benchmark or application format requirements that should apply across ReMoM, Fusion, and Flow instead of hard-coding task-specific prompts into an algorithm. output_contract_spec is the typed router-executable contract for post-processing and normalization; keep runtime behavior there instead of encoding it as prompt-text heuristics. Extraction defaults to exact content matching; use extract.sources or extract.mode: json_object only when the decision explicitly permits a wider parser.

Minimal algorithm configuration:

algorithm:
type: remom
remom:
breadth_schedule: [3, 2] # Parallel calls before final synthesis
model_distribution: weighted # weighted, equal, round_robin, or first_only
temperature: 0.7 # Temperature for model calls
include_reasoning: false # Include reasoning in synthesis
compaction_strategy: full # full or last_n_tokens
compaction_tokens: 1000 # Tokens to keep for last_n_tokens
synthesis_template: "" # Custom synthesis template (optional)
max_concurrent: 3 # Max concurrent calls per round
max_completion_tokens: 1024 # Completion limit for each subrequest
round_timeout_seconds: 120 # Optional round-level wait cap
min_successful_responses: 2 # Optional early-success quorum
shuffle_seed: 42 # Seed for response shuffling
include_intermediate_responses: false # Include intermediate responses in output
max_responses_per_round: null # Limit responses per round
on_error: skip # skip or fail

Parameters​

ParameterTypeDefaultDescription
breadth_schedulelist[int]requiredParallel calls before the final synthesis call (e.g., [3, 2])
model_distributionstringweightedStrategy: weighted, equal, round_robin, first_only
temperaturefloat1.0Temperature for model calls
include_reasoningboolfalseInclude reasoning content in synthesis prompts
compaction_strategystringfullStrategy: full or last_n_tokens
compaction_tokensint1000Tokens to keep for last_n_tokens compaction
synthesis_templatestring—Custom synthesis prompt template
max_concurrentint—Maximum concurrent model calls per round
max_completion_tokensintrequest defaultMaximum completion tokens applied to every ReMoM subrequest
round_timeout_secondsint—Maximum seconds to wait for a round before using partial responses when on_error: skip
min_successful_responsesint—Return from a parallel round after this many successful responses
shuffle_seedint42Random seed for response shuffling
include_intermediate_responsesbooltrueInclude intermediate responses in output
max_responses_per_roundint—Maximum responses to keep per round
on_errorstringskipBehavior on failure: skip or fail

Each round shares request-derived and intermediate text with its assigned models, and the synthesis model receives the collected results. Bound breadth, completion tokens, concurrency, and timeouts before production use. See a complete example: config/fragments/algorithm/looper/remom.yaml.

Optional trace transport​

Router-generated reasoning_mom_responses evidence is preserved in OpenAI Chat Completions JSON and SSE responses, including tool-call responses where the algorithm supports them. The trace is optional: if adding it would exceed the complete response limit or an individual SSE frame limit, the router omits the whole trace while serving the valid answer. It does not truncate trace JSON or answer text.

OpenAI Responses and Anthropic Messages responses omit this Chat-specific extension. Both an unsupported target protocol and a size-based omission emit a x-vsr-protocol-warnings response header with action dropped, field reasoning_mom_responses, and reason router_extension_unsupported_protocol or router_extension_size_limit. The warning also accompanies immediate Looper responses. Required answer content remains subject to the normal protocol limits; an answer that cannot fit is rejected.

Provider fields do not acquire router provenance by using the same name. The router validates provider content separately from its internal trace data.