mLateOn tops multilingual ColBERT retrieval with 115M active params, rivals 8B dense models
IgorCarron · x · 2026-08-20
LightOn's mLateOn (115M active / 312M total params) leads multilingual ColBERT-style retrieval on HAKARI-Bench: 65.52 Overall Macro, 8.33 points ahead of runner-up pplx-embed-v1-late-0.6b while using about a quarter of its active parameters. It combines a multilingual mmBERT encoder with ColBERT-style token matching, supports up to 8,192 tokens, and scores 63.33 on MNanoBEIR — between Qwen3-Embedding-8B (62.87) and Nemotron-3-Embed-8B (64.28).
Related event: LightOn's mLateOn Tops Multilingual ColBERT Retrieval Benchmark(2 posts)→
More from Research
- Analysis: Debunking 'Copycat' Claims on Chinese Labs & Deep Dive into Scaling Law — GaryMarcus · 2026-08-20
- Sept NVIDIA FLARE Day to feature FedUMM: Federated Learning for Unified Multimodal Models — jindong_wang92 · 2026-08-20
- Harvard and MIT release lecture on estimation with AI-generated data — JeremyNguyenPhD · 2026-08-20
- Papers with Code adds paper visualizations powered by the Excalidraw MCP — NielsRogge · 2026-08-20
- Epoch AI Releases Interactive Explorer for Cybersecurity Vulnerability Trends — scaling01 · 2026-08-20
- Recirculation Mechanism Verified: +17% GSM8K with Zero Weight Changes — savvyRL · 2026-08-20