LightOn Fully Open-Sources Multilingual Retrieval Models and 2.8B Training Pairs

The LightOn team has released mDenseOn and mLateOn, 307M-parameter multilingual retrieval models. This is a highly comprehensive open-source release, including not only model weights but also 2.8 billion training pairs, 16.3 million fine-tuning samples, and the training code. The models achieve SOTA within their tier on benchmarks like BEIR and MLDR, aiming to bridge the reproducibility gap caused by closed-source data.

Confirmed

Why it matters

2026-07-30 ~ 2026-07-30 · 11 related posts

Primary sources

8 near-duplicate retellings: antoine_chaffin · antoine_chaffin · tomaarsen · antoine_chaffin · IgorCarron · antoine_chaffin · antoine_chaffin · antoine_chaffin