Benchmarking
Use sr-bench 1.0 to compare a real MoM entrypoint with single models on reusable frozen tasks. The CLI and Dashboard share quality, cost, token, latency and wall-time evidence. Begin with a bounded development slice and use a disjoint holdout for acceptance.
| Question | Surface |
|---|---|
| Does Balance match the best single model while reducing subject cost? | sr-bench 1.0 |
| How do I inspect routing and improve a recipe before live evaluation? | Tune and verify a recipe |
| How do I publish measured model evidence into routing configuration? | Custom evaluations |
| Did a code change regress allocations or component latency? | Component microbenchmarks |
| Which cache store or inference binding works better here? | Backend comparisons |
Component benchmarks and specialized scripts under bench/ remain developer
diagnostics. They do not produce a complete sr-bench score or replace a paired
live model comparison.
Component microbenchmarks
The perf/ package contains Go benchmarks for classification, decision
evaluation, response-cache operations, ExtProc processing, and Looper-family
paths. They do not need a running Router. Classification and cache benchmarks
serve the catalog's pinned Vela models through the model runtime, which
downloads them on first start; install the runtime once:
make model-runtime-install
make perf-bench-quick
Useful targets:
make perf-benchruns the full component set.make perf-bench-classification,make perf-bench-decision,make perf-bench-cache, andmake perf-bench-loopernarrow the run.make perf-checkrecords benchmark output and fails when a gated allocation or byte baseline regresses beyond its configured threshold.make perf-comparecompares an existingreports/bench-output.txtwithout failing on the result.make perf-profile-cpuandmake perf-profile-memproduce pprof data.
The regression gate uses allocs/op and B/op for pass/fail. ns/op is
reported as advisory because it varies with the runner. Performance CI is
selected for changes owned by the performance domain and is also available in
manual and nightly workflows; it is not run for every documentation or product
change.
See the repository's
perf/README.md
for baseline and profiling details.
Backend comparisons
These targets start or build their own dependencies. Run them on the hardware and container runtime you intend to evaluate:
# Response-cache stores
make benchmark-cache-comparison
make benchmark-hybrid-vs-milvus
make benchmark-redis
make benchmark-valkey
Model latency and throughput are measured per model with the model runtime's
benchmarks; their records are under
src/model-runtime/docs/records.
Do not interpret a store or model comparison as an end-to-end routing result. Network placement, warmup, dataset shape, model files, and host contention can change the outcome.
Run status and exit codes
vllm-sr benchmark run and vllm-sr benchmark preview poll the run until it reaches a terminal state — completed, failed, cancelled, or interrupted — print the report, and then exit:
0— the run completed.2— the run reached a terminal state other than completed.
preview inspects routing decisions without producing quality scores; run produces the full quality report. Both use the same status and exit-code contract.
Exit code 2 conventionally signals a usage error, so automation that only checks for a non-zero exit misreads a failed run as CLI misuse. Check for 2 explicitly when a script needs to tell the two apart.
--detach opts out of this contract: the command returns immediately with the durable run ID before the run reaches a terminal state, so there is no exit code to interpret. Use vllm-sr benchmark runs to inspect detached runs and vllm-sr benchmark cancel to stop one.
Reporting results
For any number intended to guide a deployment or public claim, record:
- repository commit and complete Router configuration
- model, dataset, and dependency revisions
- hardware, driver, runtime, and backend topology
- exact command, warmup, concurrency, and sample count
- failures and excluded samples
- raw artifacts and the aggregation method
Compare alternatives on the same workload and environment. A QPS, latency, accuracy, cost, or savings number without this context is a local observation, not an expected property of Semantic Router.
Model-selection evaluation used during training is documented separately in Model Performance Evaluation.