Skip to main content
Version: Latest (unreleased)

sr-bench 1.0

sr-bench compares a Mixture of Models (MoM) entrypoint with single models on the same frozen tasks. The CLI and Dashboard → Evaluation share one durable service, run IDs, results and reports. Use a small development slice to tune routing, then evaluate a disjoint holdout before making a quality or savings claim.

Choose a scope​

Each number below is a whole task count per target, not a model-call count. Coding and agent tasks can make several calls; judges and user simulators add separately reported costs.

Benchmark IDCapabilitySmokeQuick/devStandard/holdout
mmlu-proKnowledge across 14 subjects145002,000
gpqa-diamondScientific reasoning440158
hleHLE text-only reasoning440200
livecodebenchProgramming, cumulative v6 tasks230150
scicodeScientific programming, whole main problems1320
terminal-bench-2.1Sandboxed terminal tasks1315
simpleqa-verifiedFactual correctness5100500
arc-agi-2Public evaluation puzzles, exact output grids21280
tau3τ³ text interaction, three domains31260
Total367403,183

Run vllm-sr benchmark catalog for the installed adapter identities. Select a capability slice when that answers the current question. A 500-question MMLU-Pro slice is a development comparison, not the complete 12,032-question upstream benchmark. A run without all nine benchmarks has no complete sr-bench score.

Preparation uses pinned sources, stable task IDs, content hashes, a recorded seed and proportional stratification. Smoke is a subset of quick; standard is disjoint from quick. Related units such as SciCode subproblems stay together. Public tasks are not contamination-free. Previously seen GPQA labels require a retest disclosure even when the local split is called holdout.

Next​

  1. Connect the shared worker: set up the service that the CLI and Dashboard share.
  2. Prepare reusable tasks and targets: freeze datasets, reserve named evaluation history and register targets.
  3. Plan and run: freeze a manifest and run it within limits.
  4. Iterate with preview, replay and live evaluation: tune routing on dev tasks, then evaluate a holdout.
  5. Read the results: interpret reports, comparisons and Dashboard views, and recover unfinished work.