LightOn Open-Sources Multilingual Retrieval Models Topping BEIR & MTEB Leaderboards

antoine_chaffin · x · 2026-07-30

LightOn has released mDenseOn and mLateOn, two fully open-source 307M-parameter multilingual retrieval models. These models achieve state-of-the-art results across English general-domain retrieval (BEIR), long-document retrieval (MLDR), multilingual retrieval (MIRACL), and code retrieval (MTEB Code).

Built upon a validated English data recipe, the pipeline was extended to eight additional languages via translate-train, creating a massive 2.8B-pair corpus. Most notably, the late-interaction model (mLateOn) demonstrates strong zero-shot generalization, remaining competitive in languages and scripts entirely absent from retrieval training (such as Cyrillic, Japanese, Korean, and Chinese). This effectively removes the main limitation of translate-train pipelines with dense models. Models, datasets, and training code are fully open-sourced.

Related event: LightOn Releases Fully Open-Source Multilingual Retrieval Models mDenseOn and mLateOn(11 posts)→

Original post →

More from Research

Research channel →