ModernBERT-base-32k Performance Benchmark Results
This tutorial provides benchmark results and performance tuning guidance for ModernBERT-base-32k integration. Use these results to provision hardware and adjust workload expectations for your deployment.
Overview
ModernBERT-base-32k extends the context window from 512 tokens (BERT-base) to 32,768 tokens, enabling processing of long documents and conversations. This guide presents empirical benchmark results from comprehensive testing across different context lengths and concurrency levels.
Test Environment:
- GPU: NVIDIA L4 (23GB VRAM)
- Flash Attention 2: Enabled
- Model:
llm-semantic-router/modernbert-base-32k - Test Tool:
candle-binding/examples/benchmark_concurrent.rs