LightOn Releases Multilingual Retrieval Models and 2.8B Training Pairs

IgorCarron · x · 2026-07-30

LightOn has introduced two new multilingual retrieval models, mDenseOn and mLateOn, focusing on multilingual, long-context, and code retrieval.

The highlight of this release is its full openness. Alongside the models, the team has open-sourced a massive training corpus comprising 2.8 billion curated query-document pairs across 9 languages, including French, Italian, German, and Arabic. The developers claim these models achieve frontier performance in multilingual retrieval.

Related event: LightOn Releases Open-Source Multilingual Retrieval Models with 2.8B Training Pairs(10 posts)→

Original post →

More from Models

Models channel →