From ModernBERT-base to 0.48 NanoBEIR nDCG@10 in 25 Minutes on One RTX 3090
tomaarsen · x · 2026-08-18
Training late-interaction models in v6.0 works like other model types: 4 new losses (including cached & distillation variants), 5 new evaluators, and a Trainer taking the same arguments you already know. From a bare ModernBERT-base, NanoBEIR mean nDCG@10 improves from 0.1338 to 0.4831 in 25 minutes on a single RTX 3090. Sentence Transformers also doesn't need its own late-interaction index: these indexes store whatever encodedocument returned — Qdrant, Weaviate, Vespa, LanceDB, VectorChord & Milvus index multi-vectors natively, and LightOn's fast-plaid is a pip install away.
Related event: Sentence Transformers v6.0 Ships Late-Interaction Multi-Vector Models(27 posts)→
More from coding & agent
- Swarm's First Commercial Project: Autonomous Drones for Solar Parks in Greece — bittingthembits · 2026-08-18
- DeepSeek open sources agent runtime with multi-vendor orchestration — abhishek__AI · 2026-08-18
- Gemini Launches Managed Agents with Linux Sandboxes — _philschmid · 2026-08-18
- ZenML open-sources Kitaru: replay production agent traces against different models or prompts — htahir1 · 2026-08-18
- Opinion: Generated code is safer than package manager code — rickasaurus · 2026-08-18
- Went from writing 95% of my code to under 1% in less than a year — cross__entropy · 2026-08-18