From ModernBERT-base to 0.48 NanoBEIR nDCG@10 in 25 Minutes on One RTX 3090

tomaarsen · x · 2026-08-18

Training late-interaction models in v6.0 works like other model types: 4 new losses (including cached & distillation variants), 5 new evaluators, and a Trainer taking the same arguments you already know. From a bare ModernBERT-base, NanoBEIR mean nDCG@10 improves from 0.1338 to 0.4831 in 25 minutes on a single RTX 3090. Sentence Transformers also doesn't need its own late-interaction index: these indexes store whatever encodedocument returned — Qdrant, Weaviate, Vespa, LanceDB, VectorChord & Milvus index multi-vectors natively, and LightOn's fast-plaid is a pip install away.

Related event: Sentence Transformers v6.0 Ships Late-Interaction Multi-Vector Models(27 posts)→

Original post →

More from coding & agent

coding & agent channel →