LightOn Releases 2.8B Pair Multilingual Retrieval Dataset and Two SOTA Models
antoine_chaffin · x · 2026-07-30
LightOn announced the release of their largest retrieval work to date: a massive training corpus containing 2.8 billion curated query-document pairs across 9 languages, including French, German, and Arabic. Alongside the dataset, they also released two new SOTA models trained on it: mLateOn and mDenseOn.
More from Models
- PolyAI Launches Dialog-RSN-1 Voice Model, Beats GPT and Gemini in Enterprise Tests — matthen2 · 2026-07-31
- Rumor: xAI to Release Grok 4.6 Next Week — mark_k · 2026-07-31
- Measuring Intelligence Per Watt Reveals Lack of Frontier Pricing Moat — ajratner · 2026-07-31
- CrisperWhisper2.0 Trends on HF: Verbatim Transcription & Word Timestamps — nyralabs · 2026-07-31
- OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching — AWS ML Blog · 2026-07-31
- Developer Test: Treating Claude Opus as an Unsteerable Mad Agent Works Better — brandon_galang · 2026-07-30