LightOn Releases 2.8B Pair Multilingual Retrieval Dataset and Two SOTA Models

antoine_chaffin · x · 2026-07-30

LightOn announced the release of their largest retrieval work to date: a massive training corpus containing 2.8 billion curated query-document pairs across 9 languages, including French, German, and Arabic. Alongside the dataset, they also released two new SOTA models trained on it: mLateOn and mDenseOn.

Related event: LightOn Releases 307M Open-Source Multilingual Retrieval Models with 2.8B Training Pairs(13 posts)→

Original post →

More from Models

Models channel →